docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips stretches of speech" finding, including other Parakeet weights. Investigation only; nothing deployed. Against ground truth (official SCOTUS transcript, Gutenberg #38916) the drops are real: production v3 loses 140 / 66 clean words per transcript on the two public files and ~50 on each private one (Whisper-referenced, Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause: the v2/v3 0.6B weights collapse deep inside long full-attention windows; the encoder output is degraded, the audio alone transcribes fine, and 1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy variants, max_symbols, beam), slice length, local attention, loudness, resampling and a noise floor do not fix it. Controls: A-vs-A, silence positive control (>=15 words 36/36), null control, bootstrap floor. Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds speech but no word came out (-80 to -90 % lost words on all four recordings, lower WER, no invented text, +10 MiB) and adds an explicit PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed. Reviewed at high effort, all findings fixed; built and tested as scriberr:local-blackwell-a353078-dropout2, not deployed. scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks both Parakeet scripts (seam-check --standard for the short-audio one).
This commit is contained in:
@@ -181,7 +181,7 @@ _As of 2026-09-30 ~0120 PT._
|
||||
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
||||
- **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`.
|
||||
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout1`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
|
||||
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
|
||||
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
|
||||
|
||||
Reference in New Issue
Block a user