docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)

Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
This commit is contained in:
vh
2026-09-30 15:48:02 -07:00
parent 1189adbf18
commit 38015a1977
7 changed files with 1162 additions and 17 deletions
+13
View File
@@ -8,6 +8,8 @@ distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `proposed/0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **proposed, not applied** (the rebuild script reads only this directory, not `proposed/`); built and tested as `scriberr:local-blackwell-a353078-dropout2`, not deployed. Why and how: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
@@ -104,6 +106,17 @@ each other; the default is the simplest of them.
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc.
### 0002 (proposed) — gap retry and an explicit model path
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
the Go adapter record it as `ModelUsed`. To adopt: move it up into this directory and
rebuild. To test-build it on top of the carried set under its own suffix:
`scripts/scriberr-rebuild --suffix <name> --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
### Upgrading upstream
```bash