Files
esh-pfi-infrastructure/docs/pfi/parakeet-dropout-investigation-2026-09-30.md
T
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00

25 KiB
Raw Blame History

Parakeet dropout investigation (2026-09-30)

Prime's ask, via the coordinator, 2026-09-30: investigate the "Parakeet skips stretches of speech" finding from the slicer bench (docs/pfi/scriberr-slicer-bench-2026-09-30.md), and include a different Parakeet weight. Investigation only: nothing here changed the live Scriberr container or its .env; deploying anything is Prime's call.

Privacy: two of the four recordings are Prime's. Their audio, transcripts and diarization stay in fv-ml1 /tank/spikes/scriberr-slicer/private/ (mode 700). This document carries metrics only.

Verdict

  • The drops are real. Against ground truth, today's Parakeet (v3, shipped slicer) loses whole stretches of clean speech: 140 words per transcript on a clean audiobook, 66 on a court argument, and about 50 on each of Prime's two recordings (Whisper-referenced, Canary-confirmed). The old "reference" was mostly right; on SCOTUS it drops speech too.
  • Root cause: the Granary-era 0.6B weights (v3 and v2) sometimes stop producing words for tens of seconds deep inside a long full-attention window. The encoder output there is degraded (a fresh decoder recovers only 22 to 51 %), the same audio transcribed alone is fine, and older Parakeets with the same TDT decoder do not do it. It is chaotic in cut placement and not explained by loudness, SNR, speaking rate or language. Decoding settings, beam search, shorter slices, loudness normalisation and resampling do not fix it.
  • Fix that works: re-transcribe any ≥ 3 s stretch where the audio holds speech but the model produced no words. That cuts the loss 80 to 90 % on all four recordings (to 19, 13, 7 and 5 words per transcript) with no invented text, at +10 MiB and no measurable time. It is patch 0002 (proposed, built and tested, not deployed), which also adds an explicit PARAKEET_MODEL_PATH.
  • Weights: v2 loses nothing on Prime's recordings or SCOTUS but collapses on read speech; parakeet-unified-en-0.6b (newest) is the best-balanced punctuating Parakeet but needs NeMo 3.0.0 and a licence change. The 1.1B and CTC Parakeets never collapse but produce no punctuation.
  • Separate, smaller issue: overlapping speakers (SCOTUS interjections) are lost by every single-stream model. That is not this bug.

1. Is the instrument right? Ground truth, not a model reference

The first finding measured losses against the whole-file local-attention transcript. That reference could itself be inserting text, so the public recordings were re-scored against real ground truth.

file ground truth provenance
scotus (first 30 min of No. 22-451) the official argument transcript, supremecourt.gov/oral_arguments/argument_transcripts/2023/22-451_114p.pdf (sha256 4feb7786…78a2) cover pages, page and line numbers, running headers, time stamps, argument headings and speaker labels removed; cut where the audio ends by alignment (5,627 of 14,764 transcript words)
wilde (LibriVox section 1) Project Gutenberg #38916, The Trial of Oscar Wilde, from the Shorthand Reports (the LibriVox page's own "online text" link), sha256 2271f271…b3a9 section 1 is the book's Preface, read by the narrator alone ("It is wrong for us…" to "…came too late."), plus LibriVox's standard spoken intro and outro

Normalisation: Whisper's English normaliser (transformers 4.53.3, the version in Scriberr's env, no spelling map), applied word by word to both sides so every hypothesis token keeps its timestamp. Fillers are dropped, numbers become digits and contractions are expanded.

Alignment and the dropout detector: difflib opcodes, then exact Levenshtein inside each mismatch. A dropout is a stretch between solid matches (≥ 3 tokens) holding ≥ 10 reference words where the hypothesis emitted fewer than half as many. An insertion run is the mirror image. Matching islands shorter than 3 tokens count as part of the stretch, so one spurious "the" cannot split a skipped paragraph in two.

Timed ground truth: every ground-truth token gets the median start time of the transcripts that matched it (14 transcripts); the 143 (scotus) and 61 (wilde) tokens that no transcript matched are interpolated. There are no time reversals over 1 s.

Result: the drops are real, measured against ground truth

The same outputs the first finding used (the shipped slicer and upstream's fixed cutter, 3 placements each), now scored against ground truth:

file system WER dropouts words dropped
wilde whole-file local attention (the old stand-in reference) 2.12 % 0 0
wilde upstream fixed 120 s cutter 2.57 % 0 0
wilde shipped slicer, at 120 / 110 / 100 s 2.15 / 13.02 / 2.40 % 0 / 3 / 0 0 / 394 / 0
scotus whole-file local attention 6.01 % 6 146
scotus upstream fixed 120 s cutter 4.54 % 4 46
scotus shipped slicer, at 120 / 110 / 100 s 5.42 / 5.89 / 5.30 % 4 / 5 / 4 94 / 137 / 111

On Wilde, a clean single-narrator audiobook, the stand-in reference was right and the chunked runs genuinely lost whole paragraphs; which paragraphs depends only on where the cuts fall. On SCOTUS the stand-in reference also drops speech (146 words), so the first finding's "reference insertions" there were really reference deletions.

Adjudicating Prime's recordings without ground truth

Two independent models give a second opinion on every disputed stretch (≥ 10 words one system has and another lacks): Whisper large-v3 (openai, pinned 06f233fe, transformers sequential long-form decoding; its word times are spread evenly within segments, so it gets a ±3 s window) and Canary-1b-v2 (the copy inside Scriberr's env, sha-matched to HF d4557063; cut in pauses, no overlap, ±1 s window). "Speech is real" needs both to have ≥ 50 % of the disputed words; "no speech" needs both under 20 %; anything else is ambiguous.

The adjudicator is calibrated on the public files first, where ground truth gives the right answer. It decided 129 of 140 disputed stretches and got all 129 right; it abstained on 11, mostly crosstalk that the official transcript renders differently. Whisper large-v3 on its own scores 1.96 % (wilde) and 3.93 % (scotus) WER against ground truth with zero dropouts, so for scoring Prime's files it serves as the reference, and a dropout counts only if Canary independently has the words. On the public files that metric reproduces the ground-truth clean-speech dropout counts within a few percent (1,060 vs 1,116; 287 vs 310; 240 vs 246; 5,450 vs 5,513; 515 vs 526 words).

Controls and floor

  • A-vs-A. Production v3 (full attention) is byte-identical across three separate processes at every placement tested. Local attention is not: at the same SCOTUS placement three runs dropped 68, 169 and 100 words (WER 4.73 to 6.32 %). Its numbers below carry that run-to-run noise.
  • Positive control. Copies of both public files with 30, 15, 5 and 4 s of audio replaced by digital silence (the ground truth still holds those words): the 30, 15 and 5 s stretches (15 to 73 words) were reported in all 36 runs; the 4 s stretch (11 to 14 words) in 11 of 12.
  • Null control. In those same runs, chunks that touch no silenced stretch get byte-identical input; with full attention their words were identical and their dropouts equal in all 6 runs, so the instrument manufactures nothing. (With local attention they differ; that is the non-determinism above.)
  • A "should-not-matter" perturbation (−0.5 dB gain) left SCOTUS identical and Wilde within its placement spread (1,027 vs 1,116 words). Parakeet normalises each mel feature per chunk, so a constant gain is nearly a no-op by design.
  • Sensitivity floor: a dropout of ≥ 15 words is caught every time (36/36); 10 to 14 words, 11/12; shorter losses are not counted as dropouts at all (they still count in WER). Because the failure is chaotic in cut placement, every configuration is run at 8 placements (4 for the private files and the 1.1B diagnostics); totals over 8 placements that differ by less than about a quarter are not a difference.

2. What the real drops look like

Production v3 (the shipped slicer: 120 s chunks, 4 s overlap), 8 cut placements per public file, each dropout compared with 20 random windows of the same length from the same file (clean speech; SCOTUS crosstalk split out below):

Wilde (one narrator) SCOTUS (argument)
dropouts / words 17 / 1,116 29 / 730
start, seconds into its chunk median 47, never before 16 median 53, never before 13
runs to the end of its chunk 35 % 10 %
level vs file speech level (drops / controls) −1.0 / −1.35 dB −0.1 / −2.0 dB
local SNR (drops / controls) 42 / 41 dB 29 / 27 dB
speaking rate (drops / controls) 2.4 / 2.5 words/s 4.0 / 3.2 words/s
pause just before (drops / controls) 0.62 / 0.37 s 0.02 / 0.02 s
overlapped speech, Sortformer (drops / controls) 0 / 0 16 % / 0 % of the window
non-English words within ±10 s 0 0

Two different things are being counted:

  • A. Long-context collapse. Clean speech lost in stretches of tens of seconds (up to 60 s), beginning at a sentence boundary well into a long chunk, often running to the chunk's end. It is not explained by the audio: level, SNR and rate match the controls, there is one speaker, and nothing is non-English. Which stretches go is chaotic: it moves with the cut placement, and on Wilde no word is lost by more than 75 % of placements.
  • B. Crosstalk. On SCOTUS, a justice's interjection over counsel ("Well, wait a minute…") is lost at the same few spots in almost every run. A single-stream model transcribes one voice; the official transcript records both. That is a limit of single-stream ASR and of the transcript convention, not the failure this investigation is about, so it is separated out (diarized overlap ≥ 10 % of the stretch) in every comparison below.

Every comparison below counts clean-speech dropped words per transcript, mean with a 95 % bootstrap interval over cut placements.

Language ID is not the cause. v3's only non-ASCII output is legitimate French names in the text (Mallarmé, Comédie), never near a drop, and English-only v2 collapses too. Prime's recordings are English (function-word share 0.35 to 0.39, the same as the public English files at 0.39 to 0.44; no non-ASCII letters).

3. Mechanism

  • The encoder output is degraded, not just the decoder. On the three chunks that went fully quiet (0 words after the drop began), a fresh decoder state started on the same full-attention encoder output recovered only 22 to 51 % of the ground-truth words there. The same audio encoded on its own recovered 72 to 102 %, and local attention over the full chunk 97 to 102 % (a fourth, partial drop: 98 %, 57 %, 98 %; a control chunk: all ≈ production). That is n = 4 hand-picked chunks, so it points the way; the unbiased tests below carry the weight.
  • It belongs to the weights, not to TDT decoding. On the same slices, parakeet-tdt-1.1b (TDT), parakeet-rnnt-1.1b, parakeet-ctc-1.1b, parakeet-ctc-0.6b and both heads (TDT and CTC) of parakeet-tdt_ctc-1.1b all lose 0 words on Wilde. v3 and v2, the Granary-era 0.6B TDT models, collapse. The new parakeet-unified-en-0.6b (RNN-T) collapses once in 8.
  • It is a knife edge. Deterministic for a given input (v3 at full attention is byte-identical across processes), but the float-level non-determinism of the local-attention kernel is enough to flip whether a stretch is transcribed (68 / 169 / 100 words at one placement).
  • Recording style matters, differently per weight. Wilde and p2 are gated recordings (pauses near −64 and −74 dBFS); SCOTUS and p1 never go quiet (room tone at −40 and −36 dBFS). v2 collapses badly on Wilde but never on the other three; a −50 dBFS noise floor more than halves v3's and v2's Wilde losses but makes SCOTUS worse. Not a clean lever.

4. Hypotheses tested (unbiased: full files at 8 placements; clean-speech words per transcript)

hypothesis test Wilde SCOTUS Prime p1 / p2 verdict
— production v3 140 [48–240] 66 [41–88] 50 [40–63] / 51 [25–82] baseline
(a) decoding CUDA graphs on identical output identical no effect
(a) greedy (per-frame) instead of greedy_batch identical output identical no effect
(a) max_symbols 20 identical at its 3 placements identical no effect
(a) TDT beam 4 (text only; NeMo 2.5.3 cannot give timestamps with TDT beam), all dropouts, vs greedy under the same slicing 194 vs 136 488 vs 57 worse, 3–13× slower, unusable in Scriberr
(b) context 60 s slices 88 [41–138] 61 [41–85] 31 / 31 within the floor
(b) 30 s slices 165 [95–236] 77 [49–105] 32 / 13 worse on Wilde
(b) local attention ±64 frames in 120 s slices 34 [21–47] 97 [72–121] trades one file for the other
(b) local attention ±128 39 [24–51] 128 [92–168] 0 / 3 trades; non-deterministic
(b) local attention ±256 31 [15–47] 80 [43–120] 60 / 64 no help on Prime's files
(c) preprocessing −0.5 dB gain (null) 128 identical 54 / 30 no effect (the null)
(c) loudness normalisation to −20 dBFS, peak-safe 498 [382–613] 66 worse on quiet audio
(c) soxr resampling instead of ffmpeg (all dropouts; production on the same measure: 140 / 91) 206 92 within the floor
(c) −50 dBFS noise floor 59 [30–87] 87 [74–100] helps Wilde, hurts SCOTUS
(d) weights see § 5 the lever
new re-transcribe speech gaps ≥ 3 s (patch 0002) 19 [0–38] 13 [3–25] 7 [0–22] / 5 [0–15] fixes most of it

WER against ground truth moves the same way: production v3 4.83 % / 5.33 % median (Wilde / SCOTUS) against 2.40 % / 4.47 % with the gap retry.

5. Candidates

Clean-speech dropped words per transcript (mean, 95 % bootstrap interval over placements; public: against ground truth, 8 placements; Prime's: against Whisper, confirmed by Canary, 4 placements) and median WER (public: against ground truth; Prime's: against Whisper). Memory: per-process peak on the 35-min file (p1), nvidia-smi every 0.2 s, only the run's own container PIDs, n = 3, zero spread in every case. Speed: median wall time per job for the 35.3-min file including model load, n = 3 (spread ±5 s). "CLI" rows ran the actual patched Scriberr scripts under Scriberr's invocation; "lab" rows ran the lab harness because Scriberr's env cannot run them as-is.

candidate Wilde words / WER SCOTUS words / WER p1 words / WER p2 words / WER peak MiB job time old GPU 1 budget (5,496) GPU 3 (+Blender 270 MiB) punctuation
v3, production (0001) 140 [48–240] / 4.83 % 66 [41–88] / 5.33 % 50 / 2.96 % 51 / 2.93 % 5,496 (CLI) 53 s fits (0 spare) fits yes
v3 + gap retry (0001+0002) 19 [0–38] / 2.40 % 13 [3–25] / 4.47 % 7 / 2.52 % 5 / 2.38 % 5,506 (CLI) 48 s 10 MiB over fits yes
v2 689 [490–971] / 17.6 % 0 / 4.55 % 0 / 2.22 % 0 / 2.09 % 5,438 (CLI) 49 s fits fits yes
v2 + gap retry 151 [116–192] / 5.84 % 0 / 4.58 % 0 / 2.25 % 0 / 2.09 % 5,438 (CLI) 49 s fits fits yes
parakeet-unified-en-0.6b (NeMo 3.0.0) 31 [0–92] / 2.21 % 22 [14–32] / 4.92 % 0 / 2.19 % 0 / 1.83 % 5,438 (lab) 34 s fits fits yes
parakeet-tdt-1.1b 0 / 2.31 % 3 / 6.78 % 0 / 3.30 % 0 / 2.65 % 8,884 (lab) 46 s no fits no
parakeet-ctc-0.6b 0 / 2.46 % 3 / 6.41 % – – 5,360 (lab) 31 s fits fits no
canary-1b-v2 (30 s pause cuts) 11 [4–21] / 9.64 % ⚠ 5 / 5.25 % 0 / 3.88 % 0 / 2.84 % 10,504 (lab) 112 s no fits yes
Whisper large-v3 (reference, 1 run) 0 / 1.96 % 0 / 3.93 % – – – – – – yes

⚠ Canary's Wilde WER is its hallucination loops: 6 insertion runs across 8 placements, one of 417 words. It barely drops speech but invents it.

Other 1.1B diagnostics (no punctuation, so not candidates): parakeet-rnnt-1.1b 0 / 7, parakeet-ctc-1.1b 0 / 3, parakeet-tdt_ctc-1.1b TDT head 0 / 6 and CTC head 0 / 0 (Wilde / SCOTUS words per transcript, 4 placements).

Provenance of every weight (all pulled revision-pinned, sha256 equal to the HF LFS oid, in /tank/aimodels/huggingface/hub/):

repo revision licence (read at the raw card) file sha256
nvidia/parakeet-tdt-0.6b-v3 (in Scriberr's env) 541d1f99 CC-BY-4.0 3cbdc858…
nvidia/parakeet-tdt-0.6b-v2 ae9ad070 CC-BY-4.0 d99e3995…
nvidia/parakeet-unified-en-0.6b (2026-04-07, newest Parakeet) fe53cd88 NVIDIA Open Model License ec23ed91…
nvidia/parakeet-tdt-1.1b 53276c64 CC-BY-4.0 9c563d52…
nvidia/parakeet-rnnt-1.1b 2acc4c61 CC-BY-4.0 535896f0…
nvidia/parakeet-ctc-1.1b 20e63a0f CC-BY-4.0 8e91253d…
nvidia/parakeet-tdt_ctc-1.1b 675e7868 CC-BY-4.0 4e7ccfdd…
nvidia/parakeet-ctc-0.6b ad09ba1c CC-BY-4.0 bc01f3f8…
nvidia/canary-1b-v2 (in Scriberr's env) d4557063 CC-BY-4.0 ae5ef1bf…
openai/whisper-large-v3 (adjudicator only) 06f233fe Apache-2.0 a8e94b85…

All repo ids were verified with an authenticated HF API call before any pull (no phantoms). parakeet-unified-en-0.6b needs NeMo 3.0.0: its card says 2.7.3, but released 2.7.3 lacks its encoder argument (att_chunk_context_size), and its .nemo ships without a validation_ds config that transcribe() reads (a two-line shim). It ran in a throwaway env, not Scriberr's (NeMo 2.5.3).

6. The fix that works on any weight: re-transcribe speech gaps

The mechanism says the lost audio is fine on its own; only its long-window context breaks. So after stitching, the buffered script looks for stretches of ≥ 3 s with no word where at least half the 25 ms frames sit within 12 dB of the recording's typical speech level, re-transcribes each on its own (pieces of at most 60 s, 0.5 s of padding) and splices in the words that land inside it. Output that triggers no retry is byte-identical to today's.

  • Effect: v3's clean-speech losses fall 80 to 90 % on all four recordings (140 → 19, 66 → 13, 50 → 7, 51 → 5 words per transcript) and WER falls with them (Wilde 4.83 → 2.40 %, SCOTUS 5.33 → 4.47 %). It adds no insertion runs, so it is not making text up.
  • Trigger: 3 s beat 6 s on Prime's files (p1 7 vs 27 words per transcript) and matched it on the public ones; 6 retries per 35-minute transcript were typical.
  • Cost: peak 5,506 MiB vs 5,496 (the retry's shorter chunk shifts the allocator by 10 MiB); job time unchanged within the ±5 s run-to-run spread.
  • Residual: stretches where the model still emits a few words (no clean gap), and retries that come back short. It does not rescue v2 on the audiobook (689 → 151).
  • Validity: the production implementation (after code review) reproduces the measured lab version at every shared placement: identical dropped words on all 16 file-placement pairs tried and WER equal or up to 0.02 points lower (it now removes the odd retried word that repeated its neighbour). It is deterministic across processes, and the JSON seam checks pass for both scripts with the rebuilt segments.
  • Hardening from review: a failing retry piece is skipped and a failing retry keeps the first pass (never loses a finished transcript); a piece that looks like a hallucination loop (mostly one repeated token, or > 7 words/s) is not spliced in; retried copies of the words at a gap's edge are dropped; audio goes through a per-run temp directory; frame levels use bounded memory (the first version needed ~1.7 GB of RAM for 35 min); frame indexing is correct at any sample rate. Residual risk: the gap detector is an energy test, so a loud non-speech stretch (a music bed) is retried; the loop guard is the only check on what comes back. None of the four recordings has music.

7. Recommendation

claim strength basis reversibility
The drops are real speech lost by Parakeet, not a reference artefact insist measured against ground truth on 2 files, 129/129 adjudications correct on the calibration n/a
Ship patch 0002 (gap retry on by default, v3 weights) strongly recommend measured: 80–90 % fewer lost words on all 4 recordings, lower WER, no new insertions, n = 4–8 placements each reversible (image rollback)
Keep v3 rather than switch to v2 lean measured: v2 is perfect on Prime's two files and SCOTUS but loses 151 words per transcript on read speech even with the retry, and is English-only; the right answer depends on what Prime transcribes reversible (one env var)
Evaluate parakeet-unified-en-0.6b as the next weight, as its own project lean measured: the best-balanced punctuating Parakeet (0 / 0 on Prime's files, lowest WER there); but it needs NeMo 3.0.0 in Scriberr's env (Canary and Sortformer too), a packaging shim, and a licence change from CC-BY-4.0 to the NVIDIA Open Model License costly to reverse (env rebuild)
Do not use local attention, beam search, shorter slices, loudness normalisation or a noise floor recommend against measured: each worse or mixed, local attention also non-deterministic reversible
Do not switch to Canary-1b-v2 recommend against measured: hallucination loops (up to 417 invented words) and twice the memory reversible
Do not use the 1.1B or CTC Parakeets recommend against read at source: no punctuation or casing, which Scriberr's segmentation needs reversible

8. The integration patch

stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch (on top of 0001, upstream a353078; sha256 e3098bf2…7e62). It is in proposed/, so scripts/scriberr-rebuild does not apply it until it is moved up a directory. What it does:

  1. Gap retry (§ 6) in both scripts, --retry-gaps SECS (default 3, 0 off). Go never passes the flag, so the default is the behaviour; the JSON gains retried_gaps.
  2. PARAKEET_MODEL_PATH: an absolute path (or one relative to the env) to the .nemo both scripts load; default unchanged. The JSON model field now reports the file actually loaded, and the Go adapter records it as ModelUsed instead of the hardcoded "parakeet-tdt-0.6b-v3". No model is ever swapped in under v3's filename.
  3. Tests: 18 new unit tests (39 in total), and upstream's own standard and buffered tests pass in the built image. Reviewed at high effort; all ten findings fixed (§ 6).

Validated as a build (scripts/scriberr-rebuild --suffix dropout2 --budget 5600 --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed): all stages PASS, including the embed of both scripts, 39 unit tests, the JSON seam for the long- and the short-audio script, and memory (5,506 MiB). The image scriberr:local-blackwell-a353078-dropout2 exists on fv-ml1 and is not deployed.

To ship it (Prime's call): either deploy the already-built scriberr:local-blackwell-a353078-dropout2 per patches/README.md § Deploy, or move the patch up into stacks/scriberr/patches/ first (so it becomes part of the carried set) and rebuild under a new suffix with --budget 5600 (see the note on the budget).

To also switch weights (only if Prime chooses v2 or, later, another weight): in stacks/scriberr/compose.yaml, mount the shared model cache read-only and name the pinned file. The revision is visible in the path and recorded in every transcript's metadata:

    volumes:
      - /tank/aimodels/huggingface:/models:ro
    environment:
      - PARAKEET_MODEL_PATH=/models/hub/models--nvidia--parakeet-tdt-0.6b-v2/snapshots/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/parakeet-tdt-0.6b-v2.nemo

How it survives upgrades: scriberr-rebuild re-applies 0001 and 0002 to any pinned upstream sha and stops on a conflict; the model choice lives in our compose file, not in the image or the env directory, so an upgrade cannot silently change it. If upstream ever grows its own model selection, 0002 is dropped in favour of it.

Budget note: the rebuild script's default memory budget is still GPU 1's old 5,496 MiB. With Scriberr on GPU 3 that number no longer protects anything, but the patched build peaks at 5,506, so either pass --budget or retire the old default (a one-line change; left for Prime or the coordinator).

9. Harness, data and reproduction

All code is in fv-ml1 /tank/spikes/scriberr-slicer/code/dropout/: lab.py (the production pipeline with every knob), gtscore.py, boot.py, adjudicate.py, adj_score.py, characterize.py, probe.py, fit.sh, fitlab.sh, plus the run scripts and configs. Metrics are in …/metrics/. The public audio, ground truth and transcripts are in …/public/ and …/gt/; Prime's are in …/private/ (mode 700). Throwaway envs: …/envs/nemo300 (NeMo 3.0.0, for the unified model).

Every run was a transient --rm container on GPU 3, with Scriberr's env and the model cache mounted read-only. GPU 1 and the live Scriberr container were not touched.