Prime's ask (via the coordinator): investigate the "Parakeet skips stretches of speech" finding, including other Parakeet weights. Investigation only; nothing deployed. Against ground truth (official SCOTUS transcript, Gutenberg #38916) the drops are real: production v3 loses 140 / 66 clean words per transcript on the two public files and ~50 on each private one (Whisper-referenced, Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause: the v2/v3 0.6B weights collapse deep inside long full-attention windows; the encoder output is degraded, the audio alone transcribes fine, and 1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy variants, max_symbols, beam), slice length, local attention, loudness, resampling and a noise floor do not fix it. Controls: A-vs-A, silence positive control (>=15 words 36/36), null control, bootstrap floor. Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds speech but no word came out (-80 to -90 % lost words on all four recordings, lower WER, no invented text, +10 MiB) and adds an explicit PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed. Reviewed at high effort, all findings fixed; built and tested as scriberr:local-blackwell-a353078-dropout2, not deployed. scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks both Parakeet scripts (seam-check --standard for the short-audio one).
390 lines
25 KiB
Markdown
390 lines
25 KiB
Markdown
# Parakeet dropout investigation (2026-09-30)
|
||
|
||
Prime's ask, via the coordinator, 2026-09-30: investigate the "Parakeet skips
|
||
stretches of speech" finding from the slicer bench
|
||
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md`), and include a different
|
||
Parakeet weight. **Investigation only:** nothing here changed the live Scriberr
|
||
container or its `.env`; deploying anything is Prime's call.
|
||
|
||
**Privacy:** two of the four recordings are Prime's. Their audio, transcripts and
|
||
diarization stay in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700).
|
||
This document carries metrics only.
|
||
|
||
## Verdict
|
||
|
||
- **The drops are real.** Against ground truth, today's Parakeet (v3, shipped
|
||
slicer) loses whole stretches of clean speech: 140 words per transcript on a
|
||
clean audiobook, 66 on a court argument, and about 50 on each of Prime's two
|
||
recordings (Whisper-referenced, Canary-confirmed). The old "reference" was
|
||
mostly right; on SCOTUS it drops speech too.
|
||
- **Root cause:** the Granary-era 0.6B weights (v3 and v2) sometimes stop
|
||
producing words for tens of seconds deep inside a long full-attention window.
|
||
The encoder output there is degraded (a fresh decoder recovers only 22 to 51 %),
|
||
the same audio transcribed alone is fine, and older Parakeets with the **same
|
||
TDT decoder** do not do it. It is chaotic in cut placement and not explained by
|
||
loudness, SNR, speaking rate or language. Decoding settings, beam search,
|
||
shorter slices, loudness normalisation and resampling do not fix it.
|
||
- **Fix that works:** re-transcribe any ≥ 3 s stretch where the audio holds
|
||
speech but the model produced no words. That cuts the loss 80 to 90 % on all
|
||
four recordings (to 19, 13, 7 and 5 words per transcript) with no invented
|
||
text, at +10 MiB and no measurable time. It is patch 0002 (proposed, built and
|
||
tested, **not deployed**), which also adds an explicit `PARAKEET_MODEL_PATH`.
|
||
- **Weights:** v2 loses nothing on Prime's recordings or SCOTUS but collapses on
|
||
read speech; `parakeet-unified-en-0.6b` (newest) is the best-balanced
|
||
punctuating Parakeet but needs NeMo 3.0.0 and a licence change. The 1.1B and
|
||
CTC Parakeets never collapse but produce no punctuation.
|
||
- **Separate, smaller issue:** overlapping speakers (SCOTUS interjections) are
|
||
lost by every single-stream model. That is not this bug.
|
||
|
||
## 1. Is the instrument right? Ground truth, not a model reference
|
||
|
||
The first finding measured losses against the whole-file local-attention
|
||
transcript. That reference could itself be inserting text, so the public
|
||
recordings were re-scored against real ground truth.
|
||
|
||
| file | ground truth | provenance |
|
||
|---|---|---|
|
||
| scotus (first 30 min of No. 22-451) | the official argument transcript, `supremecourt.gov/oral_arguments/argument_transcripts/2023/22-451_114p.pdf` (sha256 `4feb7786…78a2`) | cover pages, page and line numbers, running headers, time stamps, argument headings and speaker labels removed; cut where the audio ends by alignment (5,627 of 14,764 transcript words) |
|
||
| wilde (LibriVox section 1) | Project Gutenberg #38916, *The Trial of Oscar Wilde, from the Shorthand Reports* (the LibriVox page's own "online text" link), sha256 `2271f271…b3a9` | section 1 is the book's Preface, read by the narrator alone ("It is wrong for us…" to "…came too late."), plus LibriVox's standard spoken intro and outro |
|
||
|
||
**Normalisation:** Whisper's English normaliser (transformers 4.53.3, the version
|
||
in Scriberr's env, no spelling map), applied word by word to both sides so every
|
||
hypothesis token keeps its timestamp. Fillers are dropped, numbers become digits
|
||
and contractions are expanded.
|
||
|
||
**Alignment and the dropout detector:** difflib opcodes, then exact Levenshtein
|
||
inside each mismatch. A **dropout** is a stretch between solid matches (≥ 3
|
||
tokens) holding ≥ 10 reference words where the hypothesis emitted fewer than half
|
||
as many. An **insertion run** is the mirror image. Matching islands shorter than
|
||
3 tokens count as part of the stretch, so one spurious "the" cannot split a
|
||
skipped paragraph in two.
|
||
|
||
**Timed ground truth:** every ground-truth token gets the median start time of
|
||
the transcripts that matched it (14 transcripts); the 143 (scotus) and 61
|
||
(wilde) tokens that no transcript matched are interpolated. There are no time
|
||
reversals over 1 s.
|
||
|
||
### Result: the drops are real, measured against ground truth
|
||
|
||
The same outputs the first finding used (the shipped slicer and upstream's fixed
|
||
cutter, 3 placements each), now scored against ground truth:
|
||
|
||
| file | system | WER | dropouts | words dropped |
|
||
|---|---|---|---|---|
|
||
| wilde | whole-file local attention (the old stand-in reference) | 2.12 % | 0 | 0 |
|
||
| wilde | upstream fixed 120 s cutter | 2.57 % | 0 | 0 |
|
||
| wilde | shipped slicer, at 120 / 110 / 100 s | 2.15 / **13.02** / 2.40 % | 0 / 3 / 0 | 0 / **394** / 0 |
|
||
| scotus | whole-file local attention | 6.01 % | 6 | 146 |
|
||
| scotus | upstream fixed 120 s cutter | 4.54 % | 4 | 46 |
|
||
| scotus | shipped slicer, at 120 / 110 / 100 s | 5.42 / 5.89 / 5.30 % | 4 / 5 / 4 | 94 / 137 / 111 |
|
||
|
||
On Wilde, a clean single-narrator audiobook, the stand-in reference was right
|
||
and the chunked runs genuinely lost whole paragraphs; which paragraphs depends
|
||
only on where the cuts fall. On SCOTUS the stand-in reference **also** drops
|
||
speech (146 words), so the first finding's "reference insertions" there were
|
||
really reference deletions.
|
||
|
||
### Adjudicating Prime's recordings without ground truth
|
||
|
||
Two independent models give a second opinion on every disputed stretch (≥ 10
|
||
words one system has and another lacks): **Whisper large-v3** (openai, pinned
|
||
`06f233fe`, transformers sequential long-form decoding; its word times are
|
||
spread evenly within segments, so it gets a ±3 s window) and **Canary-1b-v2**
|
||
(the copy inside Scriberr's env, sha-matched to HF `d4557063`; cut in pauses,
|
||
no overlap, ±1 s window). "Speech is real" needs both to have ≥ 50 % of the
|
||
disputed words; "no speech" needs both under 20 %; anything else is ambiguous.
|
||
|
||
**The adjudicator is calibrated on the public files first**, where ground truth
|
||
gives the right answer. It decided 129 of 140 disputed stretches and **got all
|
||
129 right**; it abstained on 11, mostly crosstalk that the official transcript
|
||
renders differently. Whisper large-v3 on its own scores 1.96 % (wilde) and 3.93 %
|
||
(scotus) WER against ground truth with **zero** dropouts, so for scoring Prime's
|
||
files it serves as the reference, and a dropout counts only if Canary
|
||
independently has the words. On the public files that metric reproduces the
|
||
ground-truth clean-speech dropout counts within a few percent (1,060 vs 1,116;
|
||
287 vs 310; 240 vs 246; 5,450 vs 5,513; 515 vs 526 words).
|
||
|
||
### Controls and floor
|
||
|
||
- **A-vs-A.** Production v3 (full attention) is byte-identical across three
|
||
separate processes at every placement tested. **Local attention is not:** at
|
||
the same SCOTUS placement three runs dropped 68, 169 and 100 words (WER 4.73 to
|
||
6.32 %). Its numbers below carry that run-to-run noise.
|
||
- **Positive control.** Copies of both public files with 30, 15, 5 and 4 s of
|
||
audio replaced by digital silence (the ground truth still holds those words):
|
||
the 30, 15 and 5 s stretches (15 to 73 words) were reported in all 36 runs; the
|
||
4 s stretch (11 to 14 words) in 11 of 12.
|
||
- **Null control.** In those same runs, chunks that touch no silenced stretch get
|
||
byte-identical input; with full attention their words were identical and their
|
||
dropouts equal in all 6 runs, so the instrument manufactures nothing. (With
|
||
local attention they differ; that is the non-determinism above.)
|
||
- **A "should-not-matter" perturbation** (−0.5 dB gain) left SCOTUS identical and
|
||
Wilde within its placement spread (1,027 vs 1,116 words). Parakeet normalises
|
||
each mel feature per chunk, so a constant gain is nearly a no-op by design.
|
||
- **Sensitivity floor:** a dropout of ≥ 15 words is caught every time (36/36);
|
||
10 to 14 words, 11/12; shorter losses are not counted as dropouts at all (they
|
||
still count in WER). Because the failure is chaotic in cut placement, every
|
||
configuration is run at 8 placements (4 for the private files and the 1.1B
|
||
diagnostics); totals over 8 placements that differ by less than about a
|
||
quarter are not a difference.
|
||
|
||
## 2. What the real drops look like
|
||
|
||
Production v3 (the shipped slicer: 120 s chunks, 4 s overlap), 8 cut placements
|
||
per public file, each dropout compared with 20 random windows of the same length
|
||
from the same file (clean speech; SCOTUS crosstalk split out below):
|
||
|
||
| | Wilde (one narrator) | SCOTUS (argument) |
|
||
|---|---|---|
|
||
| dropouts / words | 17 / 1,116 | 29 / 730 |
|
||
| start, seconds into its chunk | median 47, **never before 16** | median 53, never before 13 |
|
||
| runs to the end of its chunk | 35 % | 10 % |
|
||
| level vs file speech level (drops / controls) | −1.0 / −1.35 dB | −0.1 / −2.0 dB |
|
||
| local SNR (drops / controls) | 42 / 41 dB | 29 / 27 dB |
|
||
| speaking rate (drops / controls) | 2.4 / 2.5 words/s | 4.0 / 3.2 words/s |
|
||
| pause just before (drops / controls) | 0.62 / 0.37 s | 0.02 / 0.02 s |
|
||
| overlapped speech, Sortformer (drops / controls) | 0 / 0 | **16 % / 0 %** of the window |
|
||
| non-English words within ±10 s | 0 | 0 |
|
||
|
||
Two different things are being counted:
|
||
|
||
- **A. Long-context collapse.** Clean speech lost in stretches of tens of seconds
|
||
(up to 60 s), beginning at a sentence boundary well into a long chunk, often
|
||
running to the chunk's end. It is **not explained by the audio**: level, SNR
|
||
and rate match the controls, there is one speaker, and nothing is non-English.
|
||
Which stretches go is chaotic: it moves with the cut placement, and on Wilde
|
||
no word is lost by more than 75 % of placements.
|
||
- **B. Crosstalk.** On SCOTUS, a justice's interjection over counsel ("Well,
|
||
wait a minute…") is lost at the same few spots in almost every run. A
|
||
single-stream model transcribes one voice; the official transcript records
|
||
both. That is a limit of single-stream ASR and of the transcript convention,
|
||
not the failure this investigation is about, so it is **separated out**
|
||
(diarized overlap ≥ 10 % of the stretch) in every comparison below.
|
||
|
||
Every comparison below counts **clean-speech dropped words per transcript**, mean
|
||
with a 95 % bootstrap interval over cut placements.
|
||
|
||
**Language ID is not the cause.** v3's only non-ASCII output is legitimate French
|
||
names in the text (Mallarmé, Comédie), never near a drop, and English-only v2
|
||
collapses too. Prime's recordings are English (function-word share 0.35 to 0.39,
|
||
the same as the public English files at 0.39 to 0.44; no non-ASCII letters).
|
||
|
||
## 3. Mechanism
|
||
|
||
- **The encoder output is degraded, not just the decoder.** On the three chunks
|
||
that went fully quiet (0 words after the drop began), a fresh decoder state
|
||
started on the same full-attention encoder output recovered only 22 to 51 % of
|
||
the ground-truth words there. The same audio encoded on its own recovered 72 to
|
||
102 %, and local attention over the full chunk 97 to 102 % (a fourth, partial
|
||
drop: 98 %, 57 %, 98 %; a control chunk: all ≈ production). That is n = 4 hand-picked
|
||
chunks, so it points the way; the unbiased tests below carry the weight.
|
||
- **It belongs to the weights, not to TDT decoding.** On the same slices,
|
||
`parakeet-tdt-1.1b` (TDT), `parakeet-rnnt-1.1b`, `parakeet-ctc-1.1b`,
|
||
`parakeet-ctc-0.6b` and both heads (TDT and CTC) of `parakeet-tdt_ctc-1.1b` all
|
||
lose **0** words on Wilde. v3 and v2, the Granary-era 0.6B TDT models, collapse.
|
||
The new `parakeet-unified-en-0.6b` (RNN-T) collapses once in 8.
|
||
- **It is a knife edge.** Deterministic for a given input (v3 at full attention is
|
||
byte-identical across processes), but the float-level non-determinism of the
|
||
local-attention kernel is enough to flip whether a stretch is transcribed
|
||
(68 / 169 / 100 words at one placement).
|
||
- **Recording style matters, differently per weight.** Wilde and p2 are gated
|
||
recordings (pauses near −64 and −74 dBFS); SCOTUS and p1 never go quiet (room
|
||
tone at −40 and −36 dBFS). v2 collapses badly on Wilde but never on the other
|
||
three; a −50 dBFS noise floor more than halves v3's and v2's Wilde losses but
|
||
makes SCOTUS worse. Not a clean lever.
|
||
|
||
## 4. Hypotheses tested (unbiased: full files at 8 placements; clean-speech words per transcript)
|
||
|
||
| hypothesis | test | Wilde | SCOTUS | Prime p1 / p2 | verdict |
|
||
|---|---|---|---|---|---|
|
||
| — | **production v3** | 140 [48–240] | 66 [41–88] | 50 [40–63] / 51 [25–82] | baseline |
|
||
| (a) decoding | CUDA graphs on | identical output | identical | | no effect |
|
||
| (a) | greedy (per-frame) instead of greedy_batch | identical output | identical | | no effect |
|
||
| (a) | max_symbols 20 | identical at its 3 placements | identical | | no effect |
|
||
| (a) | TDT beam 4 (text only; NeMo 2.5.3 cannot give timestamps with TDT beam), all dropouts, vs greedy under the same slicing | 194 vs 136 | 488 vs 57 | | **worse**, 3–13× slower, unusable in Scriberr |
|
||
| (b) context | 60 s slices | 88 [41–138] | 61 [41–85] | 31 / 31 | within the floor |
|
||
| (b) | 30 s slices | 165 [95–236] | 77 [49–105] | 32 / 13 | worse on Wilde |
|
||
| (b) | local attention ±64 frames in 120 s slices | 34 [21–47] | 97 [72–121] | | trades one file for the other |
|
||
| (b) | local attention ±128 | 39 [24–51] | 128 [92–168] | 0 / 3 | trades; **non-deterministic** |
|
||
| (b) | local attention ±256 | 31 [15–47] | 80 [43–120] | 60 / 64 | no help on Prime's files |
|
||
| (c) preprocessing | −0.5 dB gain (null) | 128 | identical | 54 / 30 | no effect (the null) |
|
||
| (c) | loudness normalisation to −20 dBFS, peak-safe | **498** [382–613] | 66 | | **worse** on quiet audio |
|
||
| (c) | soxr resampling instead of ffmpeg (all dropouts; production on the same measure: 140 / 91) | 206 | 92 | | within the floor |
|
||
| (c) | −50 dBFS noise floor | 59 [30–87] | 87 [74–100] | | helps Wilde, hurts SCOTUS |
|
||
| (d) weights | see § 5 | | | | the lever |
|
||
| new | **re-transcribe speech gaps ≥ 3 s** (patch 0002) | **19 [0–38]** | **13 [3–25]** | **7 [0–22] / 5 [0–15]** | **fixes most of it** |
|
||
|
||
WER against ground truth moves the same way: production v3 4.83 % / 5.33 % median
|
||
(Wilde / SCOTUS) against 2.40 % / 4.47 % with the gap retry.
|
||
|
||
## 5. Candidates
|
||
|
||
Clean-speech **dropped words per transcript** (mean, 95 % bootstrap interval over
|
||
placements; public: against ground truth, 8 placements; Prime's: against Whisper,
|
||
confirmed by Canary, 4 placements) and **median WER** (public: against ground
|
||
truth; Prime's: against Whisper). **Memory**: per-process peak on the 35-min file
|
||
(p1), nvidia-smi every 0.2 s, only the run's own container PIDs, n = 3, zero
|
||
spread in every case. **Speed**: median wall time per job for the 35.3-min file
|
||
including model load, n = 3 (spread ±5 s). "CLI" rows ran the actual patched
|
||
Scriberr scripts under Scriberr's invocation; "lab" rows ran the lab harness
|
||
because Scriberr's env cannot run them as-is.
|
||
|
||
| candidate | Wilde words / WER | SCOTUS words / WER | p1 words / WER | p2 words / WER | peak MiB | job time | old GPU 1 budget (5,496) | GPU 3 (+Blender 270 MiB) | punctuation |
|
||
|---|---|---|---|---|---|---|---|---|---|
|
||
| **v3, production** (0001) | 140 [48–240] / 4.83 % | 66 [41–88] / 5.33 % | 50 / 2.96 % | 51 / 2.93 % | 5,496 (CLI) | 53 s | fits (0 spare) | fits | yes |
|
||
| **v3 + gap retry** (0001+0002) | **19** [0–38] / **2.40 %** | **13** [3–25] / **4.47 %** | **7** / 2.52 % | **5** / 2.38 % | 5,506 (CLI) | 48 s | **10 MiB over** | fits | yes |
|
||
| v2 | 689 [490–971] / 17.6 % | **0** / 4.55 % | **0** / 2.22 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
|
||
| v2 + gap retry | 151 [116–192] / 5.84 % | **0** / 4.58 % | **0** / 2.25 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
|
||
| parakeet-unified-en-0.6b (NeMo 3.0.0) | 31 [0–92] / **2.21 %** | 22 [14–32] / 4.92 % | **0** / **2.19 %** | **0** / **1.83 %** | 5,438 (lab) | 34 s | fits | fits | yes |
|
||
| parakeet-tdt-1.1b | **0** / 2.31 % | 3 / 6.78 % | **0** / 3.30 % | **0** / 2.65 % | 8,884 (lab) | 46 s | no | fits | **no** |
|
||
| parakeet-ctc-0.6b | **0** / 2.46 % | 3 / 6.41 % | – | – | 5,360 (lab) | 31 s | fits | fits | **no** |
|
||
| canary-1b-v2 (30 s pause cuts) | 11 [4–21] / 9.64 % ⚠ | 5 / 5.25 % | **0** / 3.88 % | **0** / 2.84 % | 10,504 (lab) | 112 s | no | fits | yes |
|
||
| *Whisper large-v3 (reference, 1 run)* | *0 / 1.96 %* | *0 / 3.93 %* | – | – | – | – | – | – | *yes* |
|
||
|
||
⚠ Canary's Wilde WER is its **hallucination loops**: 6 insertion runs across 8
|
||
placements, one of 417 words. It barely drops speech but invents it.
|
||
|
||
Other 1.1B diagnostics (no punctuation, so not candidates): `parakeet-rnnt-1.1b`
|
||
0 / 7, `parakeet-ctc-1.1b` 0 / 3, `parakeet-tdt_ctc-1.1b` TDT head 0 / 6 and CTC
|
||
head 0 / 0 (Wilde / SCOTUS words per transcript, 4 placements).
|
||
|
||
Provenance of every weight (all pulled revision-pinned, sha256 equal to the HF
|
||
LFS oid, in `/tank/aimodels/huggingface/hub/`):
|
||
|
||
| repo | revision | licence (read at the raw card) | file sha256 |
|
||
|---|---|---|---|
|
||
| nvidia/parakeet-tdt-0.6b-v3 (in Scriberr's env) | `541d1f99` | CC-BY-4.0 | `3cbdc858…` |
|
||
| nvidia/parakeet-tdt-0.6b-v2 | `ae9ad070` | CC-BY-4.0 | `d99e3995…` |
|
||
| nvidia/parakeet-unified-en-0.6b (2026-04-07, newest Parakeet) | `fe53cd88` | **NVIDIA Open Model License** | `ec23ed91…` |
|
||
| nvidia/parakeet-tdt-1.1b | `53276c64` | CC-BY-4.0 | `9c563d52…` |
|
||
| nvidia/parakeet-rnnt-1.1b | `2acc4c61` | CC-BY-4.0 | `535896f0…` |
|
||
| nvidia/parakeet-ctc-1.1b | `20e63a0f` | CC-BY-4.0 | `8e91253d…` |
|
||
| nvidia/parakeet-tdt_ctc-1.1b | `675e7868` | CC-BY-4.0 | `4e7ccfdd…` |
|
||
| nvidia/parakeet-ctc-0.6b | `ad09ba1c` | CC-BY-4.0 | `bc01f3f8…` |
|
||
| nvidia/canary-1b-v2 (in Scriberr's env) | `d4557063` | CC-BY-4.0 | `ae5ef1bf…` |
|
||
| openai/whisper-large-v3 (adjudicator only) | `06f233fe` | Apache-2.0 | `a8e94b85…` |
|
||
|
||
All repo ids were verified with an authenticated HF API call before any pull (no
|
||
phantoms). `parakeet-unified-en-0.6b` needs **NeMo 3.0.0**: its card says 2.7.3,
|
||
but released 2.7.3 lacks its encoder argument (`att_chunk_context_size`), and its
|
||
`.nemo` ships without a `validation_ds` config that `transcribe()` reads (a
|
||
two-line shim). It ran in a throwaway env, not Scriberr's (NeMo 2.5.3).
|
||
|
||
## 6. The fix that works on any weight: re-transcribe speech gaps
|
||
|
||
The mechanism says the lost audio is fine on its own; only its long-window
|
||
context breaks. So after stitching, the buffered script looks for stretches of
|
||
**≥ 3 s with no word where at least half the 25 ms frames sit within 12 dB of the
|
||
recording's typical speech level**, re-transcribes each on its own (pieces of at
|
||
most 60 s, 0.5 s of padding) and splices in the words that land inside it.
|
||
Output that triggers no retry is byte-identical to today's.
|
||
|
||
- **Effect:** v3's clean-speech losses fall 80 to 90 % on all four recordings
|
||
(140 → 19, 66 → 13, 50 → 7, 51 → 5 words per transcript) and WER falls with
|
||
them (Wilde 4.83 → 2.40 %, SCOTUS 5.33 → 4.47 %). It adds **no insertion runs**,
|
||
so it is not making text up.
|
||
- **Trigger:** 3 s beat 6 s on Prime's files (p1 7 vs 27 words per transcript)
|
||
and matched it on the public ones; 6 retries per 35-minute transcript were
|
||
typical.
|
||
- **Cost:** peak 5,506 MiB vs 5,496 (the retry's shorter chunk shifts the
|
||
allocator by 10 MiB); job time unchanged within the ±5 s run-to-run spread.
|
||
- **Residual:** stretches where the model still emits a few words (no clean gap),
|
||
and retries that come back short. It does not rescue v2 on the audiobook
|
||
(689 → 151).
|
||
- **Validity:** the production implementation (after code review) reproduces
|
||
the measured lab version at every shared placement: identical dropped words
|
||
on all 16 file-placement pairs tried and WER equal or up to 0.02 points lower
|
||
(it now removes the odd retried word that repeated its neighbour). It is
|
||
deterministic across processes, and the JSON seam checks pass for both
|
||
scripts with the rebuilt segments.
|
||
- **Hardening from review:** a failing retry piece is skipped and a failing retry
|
||
keeps the first pass (never loses a finished transcript); a piece that looks
|
||
like a hallucination loop (mostly one repeated token, or > 7 words/s) is not
|
||
spliced in; retried copies of the words at a gap's edge are dropped; audio goes
|
||
through a per-run temp directory; frame levels use bounded memory (the first
|
||
version needed ~1.7 GB of RAM for 35 min); frame indexing is correct at any
|
||
sample rate. **Residual risk:** the gap detector is an energy test, so a loud
|
||
non-speech stretch (a music bed) is retried; the loop guard is the only check
|
||
on what comes back. None of the four recordings has music.
|
||
|
||
## 7. Recommendation
|
||
|
||
| claim | strength | basis | reversibility |
|
||
|---|---|---|---|
|
||
| The drops are real speech lost by Parakeet, not a reference artefact | **insist** | measured against ground truth on 2 files, 129/129 adjudications correct on the calibration | n/a |
|
||
| Ship patch 0002 (gap retry on by default, v3 weights) | **strongly recommend** | measured: 80–90 % fewer lost words on all 4 recordings, lower WER, no new insertions, n = 4–8 placements each | reversible (image rollback) |
|
||
| Keep v3 rather than switch to v2 | **lean** | measured: v2 is perfect on Prime's two files and SCOTUS but loses 151 words per transcript on read speech even with the retry, and is English-only; the right answer depends on what Prime transcribes | reversible (one env var) |
|
||
| Evaluate `parakeet-unified-en-0.6b` as the next weight, as its own project | **lean** | measured: the best-balanced punctuating Parakeet (0 / 0 on Prime's files, lowest WER there); but it needs NeMo 3.0.0 in Scriberr's env (Canary and Sortformer too), a packaging shim, and a licence change from CC-BY-4.0 to the NVIDIA Open Model License | costly to reverse (env rebuild) |
|
||
| Do not use local attention, beam search, shorter slices, loudness normalisation or a noise floor | **recommend against** | measured: each worse or mixed, local attention also non-deterministic | reversible |
|
||
| Do not switch to Canary-1b-v2 | **recommend against** | measured: hallucination loops (up to 417 invented words) and twice the memory | reversible |
|
||
| Do not use the 1.1B or CTC Parakeets | **recommend against** | read at source: no punctuation or casing, which Scriberr's segmentation needs | reversible |
|
||
|
||
## 8. The integration patch
|
||
|
||
`stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch`
|
||
(on top of 0001, upstream `a353078`; sha256 `e3098bf2…7e62`). It is in
|
||
`proposed/`, so `scripts/scriberr-rebuild` does **not** apply it until it is moved
|
||
up a directory. What it does:
|
||
|
||
1. **Gap retry** (§ 6) in both scripts, `--retry-gaps SECS` (default 3, 0 off).
|
||
Go never passes the flag, so the default is the behaviour; the JSON gains
|
||
`retried_gaps`.
|
||
2. **`PARAKEET_MODEL_PATH`**: an absolute path (or one relative to the env) to
|
||
the `.nemo` both scripts load; default unchanged. The JSON `model` field now
|
||
reports the file actually loaded, and the Go adapter records it as
|
||
`ModelUsed` instead of the hardcoded "parakeet-tdt-0.6b-v3". **No model is
|
||
ever swapped in under v3's filename.**
|
||
3. Tests: 18 new unit tests (39 in total), and upstream's own standard and
|
||
buffered tests pass in the built image. Reviewed at high effort; all ten
|
||
findings fixed (§ 6).
|
||
|
||
**Validated as a build** (`scripts/scriberr-rebuild --suffix dropout2 --budget
|
||
5600 --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`): all
|
||
stages PASS, including the embed of both scripts, 39 unit tests, the JSON seam
|
||
for the long- and the short-audio script, and memory (5,506 MiB). The image
|
||
`scriberr:local-blackwell-a353078-dropout2` exists on fv-ml1 and is **not
|
||
deployed**.
|
||
|
||
**To ship it (Prime's call):** either deploy the already-built
|
||
`scriberr:local-blackwell-a353078-dropout2` per `patches/README.md` § Deploy, or
|
||
move the patch up into `stacks/scriberr/patches/` first (so it becomes part of
|
||
the carried set) and rebuild under a new suffix with `--budget 5600` (see the
|
||
note on the budget).
|
||
|
||
**To also switch weights (only if Prime chooses v2 or, later, another weight):**
|
||
in `stacks/scriberr/compose.yaml`, mount the shared model cache read-only and
|
||
name the pinned file. The revision is visible in the path and recorded in every
|
||
transcript's metadata:
|
||
|
||
```yaml
|
||
volumes:
|
||
- /tank/aimodels/huggingface:/models:ro
|
||
environment:
|
||
- PARAKEET_MODEL_PATH=/models/hub/models--nvidia--parakeet-tdt-0.6b-v2/snapshots/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/parakeet-tdt-0.6b-v2.nemo
|
||
```
|
||
|
||
**How it survives upgrades:** `scriberr-rebuild` re-applies 0001 and 0002 to any
|
||
pinned upstream sha and stops on a conflict; the model choice lives in our
|
||
compose file, not in the image or the env directory, so an upgrade cannot
|
||
silently change it. If upstream ever grows its own model selection, 0002 is
|
||
dropped in favour of it.
|
||
|
||
**Budget note:** the rebuild script's default memory budget is still GPU 1's old
|
||
5,496 MiB. With Scriberr on GPU 3 that number no longer protects anything, but
|
||
the patched build peaks at 5,506, so either pass `--budget` or retire the old
|
||
default (a one-line change; left for Prime or the coordinator).
|
||
|
||
## 9. Harness, data and reproduction
|
||
|
||
All code is in fv-ml1 `/tank/spikes/scriberr-slicer/code/dropout/`: `lab.py`
|
||
(the production pipeline with every knob), `gtscore.py`, `boot.py`,
|
||
`adjudicate.py`, `adj_score.py`, `characterize.py`, `probe.py`, `fit.sh`,
|
||
`fitlab.sh`, plus the run scripts and configs. Metrics are in `…/metrics/`. The
|
||
public audio, ground truth and transcripts are in `…/public/` and `…/gt/`;
|
||
Prime's are in `…/private/` (mode 700). Throwaway envs: `…/envs/nemo300`
|
||
(NeMo 3.0.0, for the unified model).
|
||
|
||
Every run was a transient `--rm` container on GPU 3, with Scriberr's env and the
|
||
model cache mounted read-only. GPU 1 and the live Scriberr container were not
|
||
touched.
|