diff --git a/docs/pfi/parakeet-dropout-investigation-2026-09-30.md b/docs/pfi/parakeet-dropout-investigation-2026-09-30.md new file mode 100644 index 0000000..4165aec --- /dev/null +++ b/docs/pfi/parakeet-dropout-investigation-2026-09-30.md @@ -0,0 +1,389 @@ +# Parakeet dropout investigation (2026-09-30) + +Prime's ask, via the coordinator, 2026-09-30: investigate the "Parakeet skips +stretches of speech" finding from the slicer bench +(`docs/pfi/scriberr-slicer-bench-2026-09-30.md`), and include a different +Parakeet weight. **Investigation only:** nothing here changed the live Scriberr +container or its `.env`; deploying anything is Prime's call. + +**Privacy:** two of the four recordings are Prime's. Their audio, transcripts and +diarization stay in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700). +This document carries metrics only. + +## Verdict + +- **The drops are real.** Against ground truth, today's Parakeet (v3, shipped + slicer) loses whole stretches of clean speech: 140 words per transcript on a + clean audiobook, 66 on a court argument, and about 50 on each of Prime's two + recordings (Whisper-referenced, Canary-confirmed). The old "reference" was + mostly right; on SCOTUS it drops speech too. +- **Root cause:** the Granary-era 0.6B weights (v3 and v2) sometimes stop + producing words for tens of seconds deep inside a long full-attention window. + The encoder output there is degraded (a fresh decoder recovers only 22 to 51 %), + the same audio transcribed alone is fine, and older Parakeets with the **same + TDT decoder** do not do it. It is chaotic in cut placement and not explained by + loudness, SNR, speaking rate or language. Decoding settings, beam search, + shorter slices, loudness normalisation and resampling do not fix it. +- **Fix that works:** re-transcribe any ≥ 3 s stretch where the audio holds + speech but the model produced no words. That cuts the loss 80 to 90 % on all + four recordings (to 19, 13, 7 and 5 words per transcript) with no invented + text, at +10 MiB and no measurable time. It is patch 0002 (proposed, built and + tested, **not deployed**), which also adds an explicit `PARAKEET_MODEL_PATH`. +- **Weights:** v2 loses nothing on Prime's recordings or SCOTUS but collapses on + read speech; `parakeet-unified-en-0.6b` (newest) is the best-balanced + punctuating Parakeet but needs NeMo 3.0.0 and a licence change. The 1.1B and + CTC Parakeets never collapse but produce no punctuation. +- **Separate, smaller issue:** overlapping speakers (SCOTUS interjections) are + lost by every single-stream model. That is not this bug. + +## 1. Is the instrument right? Ground truth, not a model reference + +The first finding measured losses against the whole-file local-attention +transcript. That reference could itself be inserting text, so the public +recordings were re-scored against real ground truth. + +| file | ground truth | provenance | +|---|---|---| +| scotus (first 30 min of No. 22-451) | the official argument transcript, `supremecourt.gov/oral_arguments/argument_transcripts/2023/22-451_114p.pdf` (sha256 `4feb7786…78a2`) | cover pages, page and line numbers, running headers, time stamps, argument headings and speaker labels removed; cut where the audio ends by alignment (5,627 of 14,764 transcript words) | +| wilde (LibriVox section 1) | Project Gutenberg #38916, *The Trial of Oscar Wilde, from the Shorthand Reports* (the LibriVox page's own "online text" link), sha256 `2271f271…b3a9` | section 1 is the book's Preface, read by the narrator alone ("It is wrong for us…" to "…came too late."), plus LibriVox's standard spoken intro and outro | + +**Normalisation:** Whisper's English normaliser (transformers 4.53.3, the version +in Scriberr's env, no spelling map), applied word by word to both sides so every +hypothesis token keeps its timestamp. Fillers are dropped, numbers become digits +and contractions are expanded. + +**Alignment and the dropout detector:** difflib opcodes, then exact Levenshtein +inside each mismatch. A **dropout** is a stretch between solid matches (≥ 3 +tokens) holding ≥ 10 reference words where the hypothesis emitted fewer than half +as many. An **insertion run** is the mirror image. Matching islands shorter than +3 tokens count as part of the stretch, so one spurious "the" cannot split a +skipped paragraph in two. + +**Timed ground truth:** every ground-truth token gets the median start time of +the transcripts that matched it (14 transcripts); the 143 (scotus) and 61 +(wilde) tokens that no transcript matched are interpolated. There are no time +reversals over 1 s. + +### Result: the drops are real, measured against ground truth + +The same outputs the first finding used (the shipped slicer and upstream's fixed +cutter, 3 placements each), now scored against ground truth: + +| file | system | WER | dropouts | words dropped | +|---|---|---|---|---| +| wilde | whole-file local attention (the old stand-in reference) | 2.12 % | 0 | 0 | +| wilde | upstream fixed 120 s cutter | 2.57 % | 0 | 0 | +| wilde | shipped slicer, at 120 / 110 / 100 s | 2.15 / **13.02** / 2.40 % | 0 / 3 / 0 | 0 / **394** / 0 | +| scotus | whole-file local attention | 6.01 % | 6 | 146 | +| scotus | upstream fixed 120 s cutter | 4.54 % | 4 | 46 | +| scotus | shipped slicer, at 120 / 110 / 100 s | 5.42 / 5.89 / 5.30 % | 4 / 5 / 4 | 94 / 137 / 111 | + +On Wilde, a clean single-narrator audiobook, the stand-in reference was right +and the chunked runs genuinely lost whole paragraphs; which paragraphs depends +only on where the cuts fall. On SCOTUS the stand-in reference **also** drops +speech (146 words), so the first finding's "reference insertions" there were +really reference deletions. + +### Adjudicating Prime's recordings without ground truth + +Two independent models give a second opinion on every disputed stretch (≥ 10 +words one system has and another lacks): **Whisper large-v3** (openai, pinned +`06f233fe`, transformers sequential long-form decoding; its word times are +spread evenly within segments, so it gets a ±3 s window) and **Canary-1b-v2** +(the copy inside Scriberr's env, sha-matched to HF `d4557063`; cut in pauses, +no overlap, ±1 s window). "Speech is real" needs both to have ≥ 50 % of the +disputed words; "no speech" needs both under 20 %; anything else is ambiguous. + +**The adjudicator is calibrated on the public files first**, where ground truth +gives the right answer. It decided 129 of 140 disputed stretches and **got all +129 right**; it abstained on 11, mostly crosstalk that the official transcript +renders differently. Whisper large-v3 on its own scores 1.96 % (wilde) and 3.93 % +(scotus) WER against ground truth with **zero** dropouts, so for scoring Prime's +files it serves as the reference, and a dropout counts only if Canary +independently has the words. On the public files that metric reproduces the +ground-truth clean-speech dropout counts within a few percent (1,060 vs 1,116; +287 vs 310; 240 vs 246; 5,450 vs 5,513; 515 vs 526 words). + +### Controls and floor + +- **A-vs-A.** Production v3 (full attention) is byte-identical across three + separate processes at every placement tested. **Local attention is not:** at + the same SCOTUS placement three runs dropped 68, 169 and 100 words (WER 4.73 to + 6.32 %). Its numbers below carry that run-to-run noise. +- **Positive control.** Copies of both public files with 30, 15, 5 and 4 s of + audio replaced by digital silence (the ground truth still holds those words): + the 30, 15 and 5 s stretches (15 to 73 words) were reported in all 36 runs; the + 4 s stretch (11 to 14 words) in 11 of 12. +- **Null control.** In those same runs, chunks that touch no silenced stretch get + byte-identical input; with full attention their words were identical and their + dropouts equal in all 6 runs, so the instrument manufactures nothing. (With + local attention they differ; that is the non-determinism above.) +- **A "should-not-matter" perturbation** (−0.5 dB gain) left SCOTUS identical and + Wilde within its placement spread (1,027 vs 1,116 words). Parakeet normalises + each mel feature per chunk, so a constant gain is nearly a no-op by design. +- **Sensitivity floor:** a dropout of ≥ 15 words is caught every time (36/36); + 10 to 14 words, 11/12; shorter losses are not counted as dropouts at all (they + still count in WER). Because the failure is chaotic in cut placement, every + configuration is run at 8 placements (4 for the private files and the 1.1B + diagnostics); totals over 8 placements that differ by less than about a + quarter are not a difference. + +## 2. What the real drops look like + +Production v3 (the shipped slicer: 120 s chunks, 4 s overlap), 8 cut placements +per public file, each dropout compared with 20 random windows of the same length +from the same file (clean speech; SCOTUS crosstalk split out below): + +| | Wilde (one narrator) | SCOTUS (argument) | +|---|---|---| +| dropouts / words | 17 / 1,116 | 29 / 730 | +| start, seconds into its chunk | median 47, **never before 16** | median 53, never before 13 | +| runs to the end of its chunk | 35 % | 10 % | +| level vs file speech level (drops / controls) | −1.0 / −1.35 dB | −0.1 / −2.0 dB | +| local SNR (drops / controls) | 42 / 41 dB | 29 / 27 dB | +| speaking rate (drops / controls) | 2.4 / 2.5 words/s | 4.0 / 3.2 words/s | +| pause just before (drops / controls) | 0.62 / 0.37 s | 0.02 / 0.02 s | +| overlapped speech, Sortformer (drops / controls) | 0 / 0 | **16 % / 0 %** of the window | +| non-English words within ±10 s | 0 | 0 | + +Two different things are being counted: + +- **A. Long-context collapse.** Clean speech lost in stretches of tens of seconds + (up to 60 s), beginning at a sentence boundary well into a long chunk, often + running to the chunk's end. It is **not explained by the audio**: level, SNR + and rate match the controls, there is one speaker, and nothing is non-English. + Which stretches go is chaotic: it moves with the cut placement, and on Wilde + no word is lost by more than 75 % of placements. +- **B. Crosstalk.** On SCOTUS, a justice's interjection over counsel ("Well, + wait a minute…") is lost at the same few spots in almost every run. A + single-stream model transcribes one voice; the official transcript records + both. That is a limit of single-stream ASR and of the transcript convention, + not the failure this investigation is about, so it is **separated out** + (diarized overlap ≥ 10 % of the stretch) in every comparison below. + +Every comparison below counts **clean-speech dropped words per transcript**, mean +with a 95 % bootstrap interval over cut placements. + +**Language ID is not the cause.** v3's only non-ASCII output is legitimate French +names in the text (Mallarmé, Comédie), never near a drop, and English-only v2 +collapses too. Prime's recordings are English (function-word share 0.35 to 0.39, +the same as the public English files at 0.39 to 0.44; no non-ASCII letters). + +## 3. Mechanism + +- **The encoder output is degraded, not just the decoder.** On the three chunks + that went fully quiet (0 words after the drop began), a fresh decoder state + started on the same full-attention encoder output recovered only 22 to 51 % of + the ground-truth words there. The same audio encoded on its own recovered 72 to + 102 %, and local attention over the full chunk 97 to 102 % (a fourth, partial + drop: 98 %, 57 %, 98 %; a control chunk: all ≈ production). That is n = 4 hand-picked + chunks, so it points the way; the unbiased tests below carry the weight. +- **It belongs to the weights, not to TDT decoding.** On the same slices, + `parakeet-tdt-1.1b` (TDT), `parakeet-rnnt-1.1b`, `parakeet-ctc-1.1b`, + `parakeet-ctc-0.6b` and both heads (TDT and CTC) of `parakeet-tdt_ctc-1.1b` all + lose **0** words on Wilde. v3 and v2, the Granary-era 0.6B TDT models, collapse. + The new `parakeet-unified-en-0.6b` (RNN-T) collapses once in 8. +- **It is a knife edge.** Deterministic for a given input (v3 at full attention is + byte-identical across processes), but the float-level non-determinism of the + local-attention kernel is enough to flip whether a stretch is transcribed + (68 / 169 / 100 words at one placement). +- **Recording style matters, differently per weight.** Wilde and p2 are gated + recordings (pauses near −64 and −74 dBFS); SCOTUS and p1 never go quiet (room + tone at −40 and −36 dBFS). v2 collapses badly on Wilde but never on the other + three; a −50 dBFS noise floor more than halves v3's and v2's Wilde losses but + makes SCOTUS worse. Not a clean lever. + +## 4. Hypotheses tested (unbiased: full files at 8 placements; clean-speech words per transcript) + +| hypothesis | test | Wilde | SCOTUS | Prime p1 / p2 | verdict | +|---|---|---|---|---|---| +| — | **production v3** | 140 [48–240] | 66 [41–88] | 50 [40–63] / 51 [25–82] | baseline | +| (a) decoding | CUDA graphs on | identical output | identical | | no effect | +| (a) | greedy (per-frame) instead of greedy_batch | identical output | identical | | no effect | +| (a) | max_symbols 20 | identical at its 3 placements | identical | | no effect | +| (a) | TDT beam 4 (text only; NeMo 2.5.3 cannot give timestamps with TDT beam), all dropouts, vs greedy under the same slicing | 194 vs 136 | 488 vs 57 | | **worse**, 3–13× slower, unusable in Scriberr | +| (b) context | 60 s slices | 88 [41–138] | 61 [41–85] | 31 / 31 | within the floor | +| (b) | 30 s slices | 165 [95–236] | 77 [49–105] | 32 / 13 | worse on Wilde | +| (b) | local attention ±64 frames in 120 s slices | 34 [21–47] | 97 [72–121] | | trades one file for the other | +| (b) | local attention ±128 | 39 [24–51] | 128 [92–168] | 0 / 3 | trades; **non-deterministic** | +| (b) | local attention ±256 | 31 [15–47] | 80 [43–120] | 60 / 64 | no help on Prime's files | +| (c) preprocessing | −0.5 dB gain (null) | 128 | identical | 54 / 30 | no effect (the null) | +| (c) | loudness normalisation to −20 dBFS, peak-safe | **498** [382–613] | 66 | | **worse** on quiet audio | +| (c) | soxr resampling instead of ffmpeg (all dropouts; production on the same measure: 140 / 91) | 206 | 92 | | within the floor | +| (c) | −50 dBFS noise floor | 59 [30–87] | 87 [74–100] | | helps Wilde, hurts SCOTUS | +| (d) weights | see § 5 | | | | the lever | +| new | **re-transcribe speech gaps ≥ 3 s** (patch 0002) | **19 [0–38]** | **13 [3–25]** | **7 [0–22] / 5 [0–15]** | **fixes most of it** | + +WER against ground truth moves the same way: production v3 4.83 % / 5.33 % median +(Wilde / SCOTUS) against 2.40 % / 4.47 % with the gap retry. + +## 5. Candidates + +Clean-speech **dropped words per transcript** (mean, 95 % bootstrap interval over +placements; public: against ground truth, 8 placements; Prime's: against Whisper, +confirmed by Canary, 4 placements) and **median WER** (public: against ground +truth; Prime's: against Whisper). **Memory**: per-process peak on the 35-min file +(p1), nvidia-smi every 0.2 s, only the run's own container PIDs, n = 3, zero +spread in every case. **Speed**: median wall time per job for the 35.3-min file +including model load, n = 3 (spread ±5 s). "CLI" rows ran the actual patched +Scriberr scripts under Scriberr's invocation; "lab" rows ran the lab harness +because Scriberr's env cannot run them as-is. + +| candidate | Wilde words / WER | SCOTUS words / WER | p1 words / WER | p2 words / WER | peak MiB | job time | old GPU 1 budget (5,496) | GPU 3 (+Blender 270 MiB) | punctuation | +|---|---|---|---|---|---|---|---|---|---| +| **v3, production** (0001) | 140 [48–240] / 4.83 % | 66 [41–88] / 5.33 % | 50 / 2.96 % | 51 / 2.93 % | 5,496 (CLI) | 53 s | fits (0 spare) | fits | yes | +| **v3 + gap retry** (0001+0002) | **19** [0–38] / **2.40 %** | **13** [3–25] / **4.47 %** | **7** / 2.52 % | **5** / 2.38 % | 5,506 (CLI) | 48 s | **10 MiB over** | fits | yes | +| v2 | 689 [490–971] / 17.6 % | **0** / 4.55 % | **0** / 2.22 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes | +| v2 + gap retry | 151 [116–192] / 5.84 % | **0** / 4.58 % | **0** / 2.25 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes | +| parakeet-unified-en-0.6b (NeMo 3.0.0) | 31 [0–92] / **2.21 %** | 22 [14–32] / 4.92 % | **0** / **2.19 %** | **0** / **1.83 %** | 5,438 (lab) | 34 s | fits | fits | yes | +| parakeet-tdt-1.1b | **0** / 2.31 % | 3 / 6.78 % | **0** / 3.30 % | **0** / 2.65 % | 8,884 (lab) | 46 s | no | fits | **no** | +| parakeet-ctc-0.6b | **0** / 2.46 % | 3 / 6.41 % | – | – | 5,360 (lab) | 31 s | fits | fits | **no** | +| canary-1b-v2 (30 s pause cuts) | 11 [4–21] / 9.64 % ⚠ | 5 / 5.25 % | **0** / 3.88 % | **0** / 2.84 % | 10,504 (lab) | 112 s | no | fits | yes | +| *Whisper large-v3 (reference, 1 run)* | *0 / 1.96 %* | *0 / 3.93 %* | – | – | – | – | – | – | *yes* | + +⚠ Canary's Wilde WER is its **hallucination loops**: 6 insertion runs across 8 +placements, one of 417 words. It barely drops speech but invents it. + +Other 1.1B diagnostics (no punctuation, so not candidates): `parakeet-rnnt-1.1b` +0 / 7, `parakeet-ctc-1.1b` 0 / 3, `parakeet-tdt_ctc-1.1b` TDT head 0 / 6 and CTC +head 0 / 0 (Wilde / SCOTUS words per transcript, 4 placements). + +Provenance of every weight (all pulled revision-pinned, sha256 equal to the HF +LFS oid, in `/tank/aimodels/huggingface/hub/`): + +| repo | revision | licence (read at the raw card) | file sha256 | +|---|---|---|---| +| nvidia/parakeet-tdt-0.6b-v3 (in Scriberr's env) | `541d1f99` | CC-BY-4.0 | `3cbdc858…` | +| nvidia/parakeet-tdt-0.6b-v2 | `ae9ad070` | CC-BY-4.0 | `d99e3995…` | +| nvidia/parakeet-unified-en-0.6b (2026-04-07, newest Parakeet) | `fe53cd88` | **NVIDIA Open Model License** | `ec23ed91…` | +| nvidia/parakeet-tdt-1.1b | `53276c64` | CC-BY-4.0 | `9c563d52…` | +| nvidia/parakeet-rnnt-1.1b | `2acc4c61` | CC-BY-4.0 | `535896f0…` | +| nvidia/parakeet-ctc-1.1b | `20e63a0f` | CC-BY-4.0 | `8e91253d…` | +| nvidia/parakeet-tdt_ctc-1.1b | `675e7868` | CC-BY-4.0 | `4e7ccfdd…` | +| nvidia/parakeet-ctc-0.6b | `ad09ba1c` | CC-BY-4.0 | `bc01f3f8…` | +| nvidia/canary-1b-v2 (in Scriberr's env) | `d4557063` | CC-BY-4.0 | `ae5ef1bf…` | +| openai/whisper-large-v3 (adjudicator only) | `06f233fe` | Apache-2.0 | `a8e94b85…` | + +All repo ids were verified with an authenticated HF API call before any pull (no +phantoms). `parakeet-unified-en-0.6b` needs **NeMo 3.0.0**: its card says 2.7.3, +but released 2.7.3 lacks its encoder argument (`att_chunk_context_size`), and its +`.nemo` ships without a `validation_ds` config that `transcribe()` reads (a +two-line shim). It ran in a throwaway env, not Scriberr's (NeMo 2.5.3). + +## 6. The fix that works on any weight: re-transcribe speech gaps + +The mechanism says the lost audio is fine on its own; only its long-window +context breaks. So after stitching, the buffered script looks for stretches of +**≥ 3 s with no word where at least half the 25 ms frames sit within 12 dB of the +recording's typical speech level**, re-transcribes each on its own (pieces of at +most 60 s, 0.5 s of padding) and splices in the words that land inside it. +Output that triggers no retry is byte-identical to today's. + +- **Effect:** v3's clean-speech losses fall 80 to 90 % on all four recordings + (140 → 19, 66 → 13, 50 → 7, 51 → 5 words per transcript) and WER falls with + them (Wilde 4.83 → 2.40 %, SCOTUS 5.33 → 4.47 %). It adds **no insertion runs**, + so it is not making text up. +- **Trigger:** 3 s beat 6 s on Prime's files (p1 7 vs 27 words per transcript) + and matched it on the public ones; 6 retries per 35-minute transcript were + typical. +- **Cost:** peak 5,506 MiB vs 5,496 (the retry's shorter chunk shifts the + allocator by 10 MiB); job time unchanged within the ±5 s run-to-run spread. +- **Residual:** stretches where the model still emits a few words (no clean gap), + and retries that come back short. It does not rescue v2 on the audiobook + (689 → 151). +- **Validity:** the production implementation (after code review) reproduces + the measured lab version at every shared placement: identical dropped words + on all 16 file-placement pairs tried and WER equal or up to 0.02 points lower + (it now removes the odd retried word that repeated its neighbour). It is + deterministic across processes, and the JSON seam checks pass for both + scripts with the rebuilt segments. +- **Hardening from review:** a failing retry piece is skipped and a failing retry + keeps the first pass (never loses a finished transcript); a piece that looks + like a hallucination loop (mostly one repeated token, or > 7 words/s) is not + spliced in; retried copies of the words at a gap's edge are dropped; audio goes + through a per-run temp directory; frame levels use bounded memory (the first + version needed ~1.7 GB of RAM for 35 min); frame indexing is correct at any + sample rate. **Residual risk:** the gap detector is an energy test, so a loud + non-speech stretch (a music bed) is retried; the loop guard is the only check + on what comes back. None of the four recordings has music. + +## 7. Recommendation + +| claim | strength | basis | reversibility | +|---|---|---|---| +| The drops are real speech lost by Parakeet, not a reference artefact | **insist** | measured against ground truth on 2 files, 129/129 adjudications correct on the calibration | n/a | +| Ship patch 0002 (gap retry on by default, v3 weights) | **strongly recommend** | measured: 80–90 % fewer lost words on all 4 recordings, lower WER, no new insertions, n = 4–8 placements each | reversible (image rollback) | +| Keep v3 rather than switch to v2 | **lean** | measured: v2 is perfect on Prime's two files and SCOTUS but loses 151 words per transcript on read speech even with the retry, and is English-only; the right answer depends on what Prime transcribes | reversible (one env var) | +| Evaluate `parakeet-unified-en-0.6b` as the next weight, as its own project | **lean** | measured: the best-balanced punctuating Parakeet (0 / 0 on Prime's files, lowest WER there); but it needs NeMo 3.0.0 in Scriberr's env (Canary and Sortformer too), a packaging shim, and a licence change from CC-BY-4.0 to the NVIDIA Open Model License | costly to reverse (env rebuild) | +| Do not use local attention, beam search, shorter slices, loudness normalisation or a noise floor | **recommend against** | measured: each worse or mixed, local attention also non-deterministic | reversible | +| Do not switch to Canary-1b-v2 | **recommend against** | measured: hallucination loops (up to 417 invented words) and twice the memory | reversible | +| Do not use the 1.1B or CTC Parakeets | **recommend against** | read at source: no punctuation or casing, which Scriberr's segmentation needs | reversible | + +## 8. The integration patch + +`stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch` +(on top of 0001, upstream `a353078`; sha256 `e3098bf2…7e62`). It is in +`proposed/`, so `scripts/scriberr-rebuild` does **not** apply it until it is moved +up a directory. What it does: + +1. **Gap retry** (§ 6) in both scripts, `--retry-gaps SECS` (default 3, 0 off). + Go never passes the flag, so the default is the behaviour; the JSON gains + `retried_gaps`. +2. **`PARAKEET_MODEL_PATH`**: an absolute path (or one relative to the env) to + the `.nemo` both scripts load; default unchanged. The JSON `model` field now + reports the file actually loaded, and the Go adapter records it as + `ModelUsed` instead of the hardcoded "parakeet-tdt-0.6b-v3". **No model is + ever swapped in under v3's filename.** +3. Tests: 18 new unit tests (39 in total), and upstream's own standard and + buffered tests pass in the built image. Reviewed at high effort; all ten + findings fixed (§ 6). + +**Validated as a build** (`scripts/scriberr-rebuild --suffix dropout2 --budget +5600 --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`): all +stages PASS, including the embed of both scripts, 39 unit tests, the JSON seam +for the long- and the short-audio script, and memory (5,506 MiB). The image +`scriberr:local-blackwell-a353078-dropout2` exists on fv-ml1 and is **not +deployed**. + +**To ship it (Prime's call):** either deploy the already-built +`scriberr:local-blackwell-a353078-dropout2` per `patches/README.md` § Deploy, or +move the patch up into `stacks/scriberr/patches/` first (so it becomes part of +the carried set) and rebuild under a new suffix with `--budget 5600` (see the +note on the budget). + +**To also switch weights (only if Prime chooses v2 or, later, another weight):** +in `stacks/scriberr/compose.yaml`, mount the shared model cache read-only and +name the pinned file. The revision is visible in the path and recorded in every +transcript's metadata: + +```yaml + volumes: + - /tank/aimodels/huggingface:/models:ro + environment: + - PARAKEET_MODEL_PATH=/models/hub/models--nvidia--parakeet-tdt-0.6b-v2/snapshots/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/parakeet-tdt-0.6b-v2.nemo +``` + +**How it survives upgrades:** `scriberr-rebuild` re-applies 0001 and 0002 to any +pinned upstream sha and stops on a conflict; the model choice lives in our +compose file, not in the image or the env directory, so an upgrade cannot +silently change it. If upstream ever grows its own model selection, 0002 is +dropped in favour of it. + +**Budget note:** the rebuild script's default memory budget is still GPU 1's old +5,496 MiB. With Scriberr on GPU 3 that number no longer protects anything, but +the patched build peaks at 5,506, so either pass `--budget` or retire the old +default (a one-line change; left for Prime or the coordinator). + +## 9. Harness, data and reproduction + +All code is in fv-ml1 `/tank/spikes/scriberr-slicer/code/dropout/`: `lab.py` +(the production pipeline with every knob), `gtscore.py`, `boot.py`, +`adjudicate.py`, `adj_score.py`, `characterize.py`, `probe.py`, `fit.sh`, +`fitlab.sh`, plus the run scripts and configs. Metrics are in `…/metrics/`. The +public audio, ground truth and transcripts are in `…/public/` and `…/gt/`; +Prime's are in `…/private/` (mode 700). Throwaway envs: `…/envs/nemo300` +(NeMo 3.0.0, for the unified model). + +Every run was a transient `--rm` container on GPU 3, with Scriberr's env and the +model cache mounted read-only. GPU 1 and the live Scriberr container were not +touched. diff --git a/persistent-memory.md b/persistent-memory.md index a6d7087..50f7b05 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -181,7 +181,7 @@ _As of 2026-09-30 ~0120 PT._ - **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness. - **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha `. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. - **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`). - - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout1`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. + - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. - Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`). - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. diff --git a/scripts/scriberr-rebuild b/scripts/scriberr-rebuild index a776d02..380ed0f 100755 --- a/scripts/scriberr-rebuild +++ b/scripts/scriberr-rebuild @@ -31,6 +31,12 @@ # Usage: # scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N] # [--budget MIB] [--memory-audio PATH] [--reuse-image] +# [--patches DIR] +# +# --patches DIR[:DIR...] applies every *.patch in those directories, sorted by file +# name, instead of the carried set in stacks/scriberr/patches. To test-build a +# proposed patch on top of the carried ones, under its own --suffix: +# --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed # # Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496, # --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB @@ -44,10 +50,11 @@ HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54} ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1 SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +STD_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe.py TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav -SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 +SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 PATCH_DIR_ARG="" MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav while [ $# -gt 0 ]; do case $1 in @@ -57,6 +64,7 @@ while [ $# -gt 0 ]; do --budget) BUDGET=$2; shift 2 ;; --memory-audio) MEM_AUDIO=$2; shift 2 ;; --reuse-image) REUSE_IMAGE=1; shift ;; + --patches) PATCH_DIR_ARG=$2; shift 2 ;; -h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;; *) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;; esac @@ -66,8 +74,10 @@ done [[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; } REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches -mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort) +PATCH_DIR=${PATCH_DIR_ARG:-$REPO_ROOT/stacks/scriberr/patches} +IFS=: read -r -a PATCH_DIRS <<<"$PATCH_DIR" +mapfile -t PATCHES < <(for d in "${PATCH_DIRS[@]}"; do find "$d" -maxdepth 1 -name '*.patch'; done \ + | awk -F/ '{print $NF "\t" $0}' | sort | cut -f2-) [ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; } # The only paths a reused build dir may differ from upstream in. mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u) @@ -192,14 +202,14 @@ else fi # ── embed ────────────────────────────────────────────────────────────────── -if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF' +if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" "$STD_REL" 2>&1 <<'EOF' docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c " import sys -script = open('/src/$3', 'rb').read() -sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)" +binary = open('/app/scriberr', 'rb').read() +sys.exit(0 if all(open('/src/' + f, 'rb').read() in binary for f in ('$3', '$4')) else 1)" EOF ); then - pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr" + pass embed "both Parakeet scripts are byte-identical inside /app/scriberr" else fail embed "the binary does not embed the patched script ${out:+($out)}" fi @@ -240,6 +250,12 @@ gpu_idle() { # "room", not "idle": Scriberr itself may be running a job on this gpu_idle seam if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")" else fail seam "$out"; fi +# The short-audio script (files under PARAKEET_CHUNK_THRESHOLD_SECS) has its own entry point. +if out=$("${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \ + -v $BUILD_DIR/tests/data:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$STD_REL /audio/$(basename "$SEAM_AUDIO_REL") \ + --output /tmp/out.json --context-left 256 --context-right 256 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \ + python3 /tools/seam-check.py /tmp/out.json --standard'" 2>&1); then pass seam-short "$(tail -1 <<<"$out")" +else fail seam-short "$out"; fi # ── memory ───────────────────────────────────────────────────────────────── "${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1" diff --git a/scripts/scriberr-seam-check.py b/scripts/scriberr-seam-check.py index 18ef4ef..3d0057c 100755 --- a/scripts/scriberr-seam-check.py +++ b/scripts/scriberr-seam-check.py @@ -7,7 +7,8 @@ job. This checks the shape Go reads plus the stitching invariants the slicer patch promises. Stdlib only, so it runs under any python3. Prints counts, never transcript text. -usage: scriberr-seam-check.py RESULT.json [--min-chunks N] +usage: scriberr-seam-check.py RESULT.json [--min-chunks N] [--standard] + --standard the short-audio script's result (no buffered/num_chunks keys) """ import argparse import json @@ -42,6 +43,8 @@ def main(): parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py") parser.add_argument("--min-chunks", type=int, default=1, help="fail unless the run used at least this many chunks") + parser.add_argument("--standard", action="store_true", + help="the short-audio script's result: no buffered/num_chunks keys") args = parser.parse_args() min_chunks = args.min_chunks try: @@ -56,12 +59,13 @@ def main(): for key, kind in required.items(): if not isinstance(data.get(key), kind): fail(f"'{key}' missing or not {kind.__name__}") - if data.get("buffered") is not True: - fail("'buffered' is not true") - if not isinstance(data.get("chunk_duration_secs"), NUMBER): - fail("'chunk_duration_secs' is not a number") - if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks: - fail(f"'num_chunks' is not an integer >= {min_chunks}") + if not args.standard: + if data.get("buffered") is not True: + fail("'buffered' is not true") + if not isinstance(data.get("chunk_duration_secs"), NUMBER): + fail("'chunk_duration_secs' is not a number") + if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks: + fail(f"'num_chunks' is not an integer >= {min_chunks}") words, segments = data["word_timestamps"], data["segment_timestamps"] if not words or not data["transcription"].strip(): @@ -79,7 +83,8 @@ def main(): fail("segments do not cover the stitched words exactly once, in order") print(f"SEAM OK: {len(words)} words, {len(segments)} segments, " - f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points") + f"{data.get('num_chunks', 1)} chunks, cuts at {len(data.get('cut_times', []))} points" + f"{', model ' + data['model'] if data.get('model') else ''}") if __name__ == "__main__": diff --git a/stacks/scriberr/README.md b/stacks/scriberr/README.md index e10e666..242b708 100644 --- a/stacks/scriberr/README.md +++ b/stacks/scriberr/README.md @@ -181,4 +181,7 @@ Peak GPU memory on a 35-minute file: peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the - patch; that is still open. + patch. **Investigated 2026-09-30** (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`): + the losses are real (against ground truth) and belong to the v2/v3 weights over + long windows; a proposed patch, `patches/proposed/0002`, re-transcribes speech that + got no words and cuts them 80–90 %. It is not deployed; that is Prime's call. diff --git a/stacks/scriberr/patches/README.md b/stacks/scriberr/patches/README.md index ae832aa..0a3b40d 100644 --- a/stacks/scriberr/patches/README.md +++ b/stacks/scriberr/patches/README.md @@ -8,6 +8,8 @@ distinctly tagged image, and proves it before anyone deploys it. | patch | against | status | |---|---|---| | `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) | +| `proposed/0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **proposed, not applied** (the rebuild script reads only this directory, not `proposed/`); built and tested as `scriberr:local-blackwell-a353078-dropout2`, not deployed. Why and how: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` | + Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`, @@ -104,6 +106,17 @@ each other; the default is the simplest of them. of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12 transcripts, for upstream's slicer too). See the bench doc. +### 0002 (proposed) — gap retry and an explicit model path + +Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long +chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the +audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which +cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the +`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes +the Go adapter record it as `ModelUsed`. To adopt: move it up into this directory and +rebuild. To test-build it on top of the carried set under its own suffix: +`scripts/scriberr-rebuild --suffix --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`. + ### Upgrading upstream ```bash diff --git a/stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch b/stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch new file mode 100644 index 0000000..5ab4835 --- /dev/null +++ b/stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch @@ -0,0 +1,719 @@ +From d253aa2b0ea2562ed4e1702380fac7e22bd11828 Mon Sep 17 00:00:00 2001 +From: Vuong Hoang +Date: Wed, 30 Sep 2026 15:00:46 -0700 +Subject: [PATCH] feat(parakeet): selectable model and a retry for speech that + got no words + +Parakeet v2/v3 sometimes stop emitting for tens of seconds deep inside a long +full-attention chunk while someone is talking; the same audio transcribed on +its own is usually fine. After stitching, any stretch of at least +--retry-gaps seconds (default 3) with no word, where most 25 ms frames sit +within 12 dB of the recording's typical speech level, is re-transcribed on +its own (in pieces of at most 60 s) and the words that land inside it are +spliced in; segments are then rebuilt from the words with the model's own +rule (end after . ? !). Output that triggers no retry is unchanged. The +same retry runs in the short-audio script, which shares the helpers. + +The retry is best effort: a failing piece is skipped and a failing retry +keeps the first pass. A piece that looks like a hallucination loop (mostly +one repeated token, or more than 7 words a second) is not spliced in, and a +retried word that repeats its neighbour at the gap edge or a piece boundary +is dropped (the first pass's copy wins). Chunk and retry audio go through a +per-run temporary directory instead of fixed /tmp paths. + +PARAKEET_MODEL_PATH (absolute, or relative to the env) selects the .nemo +both scripts load; the default stays parakeet-tdt-0.6b-v3.nemo. The JSON +"model" field reports the file actually loaded, and the Go adapter records +it as ModelUsed instead of a hardcoded name. + +The CLI and JSON seam is otherwise unchanged: the new flag is optional, and +the JSON gains retried_gaps. +--- + .../adapters/parakeet_adapter.go | 7 +- + .../adapters/py/nvidia/parakeet_transcribe.py | 53 +++- + .../py/nvidia/parakeet_transcribe_buffered.py | 231 +++++++++++++----- + .../py/nvidia/tests/test_parakeet_slicing.py | 159 +++++++++++- + 4 files changed, 380 insertions(+), 70 deletions(-) + +diff --git a/internal/transcription/adapters/parakeet_adapter.go b/internal/transcription/adapters/parakeet_adapter.go +index 4fa252d..5f3bb9e 100644 +--- a/internal/transcription/adapters/parakeet_adapter.go ++++ b/internal/transcription/adapters/parakeet_adapter.go +@@ -342,7 +342,9 @@ func (p *ParakeetAdapter) Transcribe(ctx context.Context, input interfaces.Audio + } + + result.ProcessingTime = time.Since(startTime) +- result.ModelUsed = "parakeet-tdt-0.6b-v3" ++ if result.ModelUsed == "" { ++ result.ModelUsed = "parakeet-tdt-0.6b-v3" ++ } + result.Metadata = p.CreateDefaultMetadata(params) + + logger.Info("Parakeet transcription completed", +@@ -525,6 +527,8 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu + End float64 `json:"end"` + } `json:"segment_timestamps"` + Confidence interface{} `json:"confidence,omitempty"` ++ // The scripts report the model they actually loaded (PARAKEET_MODEL_PATH can change it). ++ Model string `json:"model"` + } + + if err := json.Unmarshal(data, ¶keetResult); err != nil { +@@ -538,6 +542,7 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu + Segments: make([]interfaces.TranscriptSegment, len(parakeetResult.SegmentTimestamps)), + WordSegments: make([]interfaces.TranscriptWord, len(parakeetResult.WordTimestamps)), + Confidence: 0.0, // Default confidence ++ ModelUsed: parakeetResult.Model, + } + + // Convert segments +diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py +index 235e52a..6b2539e 100644 +--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py ++++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py +@@ -7,9 +7,17 @@ import argparse + import json + import sys + import os ++import shutil ++import tempfile + from pathlib import Path ++import librosa + import nemo.collections.asr as nemo_asr + ++# Shared with the long-audio script, which lives in the same directory. ++from parakeet_transcribe_buffered import ( ++ DEFAULT_RETRY_GAP_SECS, resolve_model_path, retry_speech_gaps, transcribe_chunk, ++) ++ + + def transcribe_audio( + audio_path: str, +@@ -18,14 +26,11 @@ def transcribe_audio( + context_left: int = 256, + context_right: int = 256, + include_confidence: bool = True, ++ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS, + ): + """ + Transcribe audio using NVIDIA Parakeet model. + """ +- # Determine model path +- model_filename = "parakeet-tdt-0.6b-v3.nemo" +- model_path = None +- + # Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv + virtual_env = os.environ.get("VIRTUAL_ENV") + if not virtual_env: +@@ -33,10 +38,11 @@ def transcribe_audio( + sys.exit(1) + + project_root = os.path.dirname(virtual_env) +- model_path = os.path.join(project_root, model_filename) ++ model_path = resolve_model_path(project_root) ++ model_name = os.path.splitext(os.path.basename(model_path))[0] + + if not os.path.exists(model_path): +- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}") ++ print(f"Error during transcription: Can't find model file: {model_path}") + sys.exit(1) + + print(f"Loading NVIDIA Parakeet model from: {model_path}") +@@ -80,6 +86,27 @@ def transcribe_audio( + text = result_data.text + word_timestamps = result_data.timestamp.get("word", []) + segment_timestamps = result_data.timestamp.get("segment", []) ++ retried = 0 ++ words_added = False ++ if retry_gap_secs > 0: ++ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp ++ try: ++ audio, sr = librosa.load(audio_path, sr=None, mono=True) ++ ++ def transcribe_span(start_sample, end_sample): ++ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr, ++ start_sample / sr, os.path.join(workdir, "retry.wav")) ++ before = len(word_timestamps) ++ word_timestamps, segment_timestamps, retried = retry_speech_gaps( ++ word_timestamps, segment_timestamps, audio, sr, retry_gap_secs, transcribe_span ++ ) ++ words_added = len(word_timestamps) != before ++ if words_added: ++ text = " ".join(w["word"] for w in word_timestamps) ++ except Exception as err: # best effort: never lose a finished transcription ++ print(f"Warning: gap retry failed ({err}); keeping the first pass") ++ finally: ++ shutil.rmtree(workdir, ignore_errors=True) + + print(f"Transcription: {text}") + +@@ -90,14 +117,15 @@ def transcribe_audio( + "word_timestamps": word_timestamps, + "segment_timestamps": segment_timestamps, + "audio_file": audio_path, +- "model": "parakeet-tdt-0.6b-v3", ++ "model": model_name, ++ "retried_gaps": retried, + "context": { + "left": context_left, + "right": context_right + } + } + +- if include_confidence: ++ if include_confidence and not words_added: # first-pass scores no longer line up + # Add confidence scores if available + if hasattr(result_data, 'confidence') and result_data.confidence: + output_data["confidence"] = result_data.confidence +@@ -119,7 +147,7 @@ def transcribe_audio( + "transcription": text, + "language": "en", + "audio_file": audio_path, +- "model": "parakeet-tdt-0.6b-v3" ++ "model": model_name + } + + if output_file: +@@ -154,6 +182,12 @@ def main(): + "--context-right", type=int, default=256, + help="Right attention context size (default: 256)" + ) ++ parser.add_argument( ++ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS, ++ help=f"Re-transcribe on its own any stretch of at least this many seconds where " ++ f"the audio holds speech but the model produced no words " ++ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)" ++ ) + parser.add_argument( + "--include-confidence", action="store_true", default=True, + help="Include confidence scores" +@@ -178,6 +212,7 @@ def main(): + context_left=args.context_left, + context_right=args.context_right, + include_confidence=args.include_confidence, ++ retry_gap_secs=args.retry_gaps, + ) + except Exception as e: + print(f"Error during transcription: {e}") +diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +index ba755c1..49bb72e 100644 +--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py ++++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +@@ -12,15 +12,21 @@ import argparse + import json + import sys + import os ++import shutil ++import tempfile + import librosa + import soundfile as sf + import numpy as np + from pathlib import Path + ++DEFAULT_MODEL = "parakeet-tdt-0.6b-v3.nemo" + DEFAULT_OVERLAP_SECS = 4.0 + DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap + QUIET_WINDOW_SECS = 0.3 + SAME_WORD_SECS = 0.5 ++DEFAULT_RETRY_GAP_SECS = 3.0 ++RETRY_PIECE_SECS = 60.0 ++RETRY_MAX_WORDS_PER_SEC = 7.0 + + + def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0): +@@ -103,14 +109,7 @@ def stitch_slices(slice_results, cut_times, chunk_spans=None): + elif len(kept) == stop - start: + segments.append(seg) + elif kept: +- segments.append({ +- **seg, +- "segment": " ".join(w["word"] for w in kept), +- "start_offset": kept[0]["start_offset"], +- "end_offset": kept[-1]["end_offset"], +- "start": kept[0]["start"], +- "end": kept[-1]["end"], +- }) ++ segments.append({**seg, **_segment_of(kept)}) + return words, segments + + +@@ -159,6 +158,134 @@ def _segment_ranges(words, segments): + return ranges + + ++def resolve_model_path(project_root): ++ """The .nemo to load: $PARAKEET_MODEL_PATH (absolute, or relative to the env), ++ else the bundled v3 weights.""" ++ chosen = os.environ.get("PARAKEET_MODEL_PATH", "").strip() or DEFAULT_MODEL ++ return chosen if os.path.isabs(chosen) else os.path.join(project_root, chosen) ++ ++ ++def find_speech_gaps(words, audio, sr, min_gap): ++ """Stretches of at least `min_gap` s with no word, where at least half the ++ 25 ms frames are within 12 dB of the recording's typical speech level: the ++ model went quiet while someone was talking. Parakeet sometimes does this for ++ tens of seconds deep inside a long chunk; the same audio transcribed on its ++ own is usually fine.""" ++ hop, win = int(0.010 * sr), int(0.025 * sr) ++ frames = max(0, (len(audio) - win) // hop) ++ if frames == 0: ++ return [] ++ # A strided view, reduced in blocks: bounded memory even for hours of audio. ++ view = np.lib.stride_tricks.sliding_window_view(audio, win)[::hop][:frames] ++ power = np.empty(frames) ++ for k in range(0, frames, 8192): ++ power[k:k + 8192] = np.mean(view[k:k + 8192].astype(np.float64) ** 2, axis=1) ++ level = 10 * np.log10(power + 1e-12) ++ speech = float(np.median(level[level >= np.median(level)])) ++ duration = len(audio) / sr ++ gaps = [] ++ for a, b in zip([0.0] + [w["end"] for w in words], [w["start"] for w in words] + [duration]): ++ if b - a >= min_gap: ++ span = level[int(a * sr) // hop:int(b * sr) // hop] ++ if len(span) and np.mean(span >= speech - 12) >= 0.5: ++ gaps.append((a, b)) ++ return gaps ++ ++ ++def retry_speech_gaps(words, segments, audio, sr, min_gap, transcribe): ++ """Re-transcribe each speech gap on its own, in pieces of at most ++ RETRY_PIECE_SECS, and splice in the words that land inside it. ++ ++ `transcribe(start_sample, end_sample)` returns (text, words, segments) in ++ absolute time. Returns (words, segments, number of gaps retried); when words ++ were added, segments are rebuilt from the words with the model's own rule. ++ """ ++ gaps = find_speech_gaps(words, audio, sr, min_gap) ++ if not gaps: ++ return words, segments, 0 ++ found = [] ++ for a, b in gaps: ++ t = a ++ while t < b: ++ e = min(b, t + RETRY_PIECE_SECS) ++ try: ++ _, piece, _ = transcribe(max(0, int((t - 0.5) * sr)), min(len(audio), int((e + 0.5) * sr))) ++ except Exception as err: # best effort: the first pass already succeeded ++ print(f"Warning: retry of {t:.1f}-{e:.1f}s failed ({err}); keeping the first pass there") ++ piece = [] ++ piece = [w for w in piece if t <= w["start"] < e] ++ if _plausible(piece, e - t): ++ found.extend(piece) ++ t = e ++ if not found: ++ return words, segments, len(gaps) ++ merged = _without_retried_duplicates(sorted(words + found, key=lambda w: w["start"]), ++ {id(w) for w in found}) ++ return merged, segments_from_words(merged), len(gaps) ++ ++ ++def _plausible(piece, seconds): ++ """False for what looks like a hallucination loop rather than speech: many words ++ that are mostly one repeated token, or more words per second than anyone says.""" ++ if len(piece) >= 10 and len({_normalize(w["word"]) for w in piece}) < 0.3 * len(piece): ++ return False ++ return len(piece) <= RETRY_MAX_WORDS_PER_SEC * max(seconds, 1.0) ++ ++ ++def _without_retried_duplicates(merged, retried): ++ """The retry's padding re-hears the words either side of a gap, and a word ++ near a piece boundary is heard by both pieces. Drop a retried word that ++ repeats its neighbour within SAME_WORD_SECS; the first pass's copy wins.""" ++ out = [] ++ for w in merged: ++ if out and _normalize(out[-1]["word"]) == _normalize(w["word"]) \ ++ and w["start"] - out[-1]["start"] <= SAME_WORD_SECS: ++ if id(w) in retried: ++ continue ++ if id(out[-1]) in retried: ++ out[-1] = w ++ continue ++ out.append(w) ++ return out ++ ++ ++def segments_from_words(words): ++ """Segments as Parakeet's decoding config makes them: a segment ends after a ++ word ending in '.', '?' or '!' (segment_seperators, no gap threshold).""" ++ segments, current = [], [] ++ for w in words: ++ current.append(w) ++ if w["word"].endswith((".", "?", "!")): ++ segments.append(_segment_of(current)) ++ current = [] ++ if current: ++ segments.append(_segment_of(current)) ++ return segments ++ ++ ++def _segment_of(words): ++ return {"segment": " ".join(w["word"] for w in words), ++ "start_offset": words[0]["start_offset"], "end_offset": words[-1]["end_offset"], ++ "start": words[0]["start"], "end": words[-1]["end"]} ++ ++ ++def transcribe_chunk(asr_model, audio, sr, start_time, path): ++ """Transcribe one chunk via a WAV file; word and segment times become absolute.""" ++ sf.write(path, audio, sr) ++ try: ++ result = asr_model.transcribe([path], batch_size=1, timestamps=True)[0] ++ finally: ++ if os.path.exists(path): ++ os.remove(path) ++ words, segments = [], [] ++ if getattr(result, "timestamp", None): ++ words = [dict(w, start=w["start"] + start_time, end=w["end"] + start_time) ++ for w in result.timestamp.get("word", [])] ++ segments = [dict(g, start=g["start"] + start_time, end=g["end"] + start_time) ++ for g in result.timestamp.get("segment", [])] ++ return result.text, words, segments ++ ++ + def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0): + """Split audio file into chunks of at most chunk_duration_secs.""" + audio, sr = librosa.load(audio_path, sr=None, mono=True) +@@ -173,7 +300,7 @@ def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, sear + 'duration': len(chunk_audio) / sr + }) + +- return chunks, sr, [cut / sr for cut in cuts] ++ return chunks, sr, [cut / sr for cut in cuts], audio + + + def transcribe_buffered( +@@ -182,16 +309,13 @@ def transcribe_buffered( + chunk_duration_secs: float = 300, # 5 minutes default + overlap_secs: float = DEFAULT_OVERLAP_SECS, + pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS, ++ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS, + ): + """ + Transcribe long audio by splitting into chunks and merging results. + """ + import nemo.collections.asr as nemo_asr + +- # Determine model path +- model_filename = "parakeet-tdt-0.6b-v3.nemo" +- model_path = None +- + # Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv + virtual_env = os.environ.get("VIRTUAL_ENV") + if not virtual_env: +@@ -199,10 +323,11 @@ def transcribe_buffered( + sys.exit(1) + + project_root = os.path.dirname(virtual_env) +- model_path = os.path.join(project_root, model_filename) ++ model_path = resolve_model_path(project_root) ++ model_name = os.path.splitext(os.path.basename(model_path))[0] + + if not os.path.exists(model_path): +- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}") ++ print(f"Error during transcription: Can't find model file: {model_path}") + sys.exit(1) + + print(f"Loading NVIDIA Parakeet model from: {model_path}") +@@ -226,58 +351,25 @@ def transcribe_buffered( + + print(f"Splitting audio into chunks of at most {chunk_duration_secs}s " + f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...") +- chunks, sr, cut_times = split_audio_file( ++ chunks, sr, cut_times, audio = split_audio_file( + audio_path, chunk_duration_secs, overlap_secs, pause_search_secs + ) + print(f"Created {len(chunks)} chunks") + + slice_results = [] + chunk_texts = [] ++ retried = 0 ++ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp + + for i, chunk_info in enumerate(chunks): + print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...") +- +- # Save chunk to temporary file +- chunk_path = f"/tmp/chunk_{i}.wav" +- sf.write(chunk_path, chunk_info['audio'], sr) +- +- try: +- # Transcribe chunk +- output = asr_model.transcribe( +- [chunk_path], +- batch_size=1, +- timestamps=True, +- ) +- +- result_data = output[0] +- chunk_text = result_data.text +- chunk_texts.append(chunk_text) +- chunk_words = [] +- chunk_segments = [] +- +- # Extract and adjust timestamps +- if hasattr(result_data, 'timestamp') and result_data.timestamp: +- # Adjust timestamps by chunk start time +- for word in result_data.timestamp.get("word", []): +- word_copy = dict(word) +- word_copy['start'] += chunk_info['start_time'] +- word_copy['end'] += chunk_info['start_time'] +- chunk_words.append(word_copy) +- +- for segment in result_data.timestamp.get("segment", []): +- seg_copy = dict(segment) +- seg_copy['start'] += chunk_info['start_time'] +- seg_copy['end'] += chunk_info['start_time'] +- chunk_segments.append(seg_copy) +- +- slice_results.append((chunk_words, chunk_segments)) +- +- print(f"Chunk {i+1} complete: {len(chunk_text)} characters") +- +- finally: +- # Clean up temp file +- if os.path.exists(chunk_path): +- os.remove(chunk_path) ++ chunk_text, chunk_words, chunk_segments = transcribe_chunk( ++ asr_model, chunk_info['audio'], sr, chunk_info['start_time'], ++ os.path.join(workdir, f"chunk_{i}.wav"), ++ ) ++ chunk_texts.append(chunk_text) ++ slice_results.append((chunk_words, chunk_segments)) ++ print(f"Chunk {i+1} complete: {len(chunk_text)} characters") + + chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks] + all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans) +@@ -287,7 +379,20 @@ def transcribe_buffered( + print("Warning: a chunk has text but no word timestamps; joining chunk texts") + final_text = " ".join(chunk_texts) + else: ++ if retry_gap_secs > 0: ++ def transcribe_span(start_sample, end_sample): ++ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr, ++ start_sample / sr, os.path.join(workdir, "retry.wav")) ++ try: ++ all_words, all_segments, retried = retry_speech_gaps( ++ all_words, all_segments, audio, sr, retry_gap_secs, transcribe_span ++ ) ++ except Exception as err: # best effort: never lose a finished transcription ++ print(f"Warning: gap retry failed ({err}); keeping the first pass") ++ if retried: ++ print(f"Re-transcribed {retried} stretch(es) of speech that got no words") + final_text = " ".join(w["word"] for w in all_words) ++ shutil.rmtree(workdir, ignore_errors=True) + print(f"Transcription complete: {len(final_text)} characters total") + + output_data = { +@@ -296,13 +401,14 @@ def transcribe_buffered( + "word_timestamps": all_words, + "segment_timestamps": all_segments, + "audio_file": audio_path, +- "model": "parakeet-tdt-0.6b-v3", ++ "model": model_name, + "buffered": True, + "chunk_duration_secs": chunk_duration_secs, + "num_chunks": len(chunks), + "overlap_secs": overlap_secs, + "pause_search_secs": pause_search_secs, + "cut_times": cut_times, ++ "retried_gaps": retried, + } + + if output_file: +@@ -328,6 +434,12 @@ def main(): + help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len " + f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)" + ) ++ parser.add_argument( ++ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS, ++ help=f"Re-transcribe on its own any stretch of at least this many seconds where " ++ f"the audio holds speech but the model produced no words " ++ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)" ++ ) + parser.add_argument( + "--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS, + help=f"Seconds before each chunk limit searched for the quietest point to cut at, " +@@ -346,6 +458,7 @@ def main(): + chunk_duration_secs=args.chunk_len, + overlap_secs=args.overlap, + pause_search_secs=args.pause_search, ++ retry_gap_secs=args.retry_gaps, + ) + + +diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py +index 6a35947..c3d7537 100644 +--- a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py ++++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py +@@ -10,7 +10,10 @@ import numpy as np + import pytest + + sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +-from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402 ++from parakeet_transcribe_buffered import ( # noqa: E402 ++ find_speech_gaps, plan_slices, resolve_model_path, retry_speech_gaps, ++ segments_from_words, stitch_slices, ++) + + SR = 16000 + +@@ -327,3 +330,157 @@ def test_punctuation_alone_is_never_an_anchor(): + right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)] + words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) + assert words[0] is left[0] and texts(words) == ["-", "no"] ++ ++ ++# -- model and attention selection ---------------------------------------------- ++ ++ ++def test_the_bundled_v3_model_is_the_default(monkeypatch): ++ monkeypatch.delenv("PARAKEET_MODEL_PATH", raising=False) ++ assert resolve_model_path("/env") == "/env/parakeet-tdt-0.6b-v3.nemo" ++ ++ ++def test_parakeet_model_path_overrides_the_model(monkeypatch): ++ monkeypatch.setenv("PARAKEET_MODEL_PATH", "/models/parakeet-tdt-0.6b-v2.nemo") ++ assert resolve_model_path("/env") == "/models/parakeet-tdt-0.6b-v2.nemo" ++ monkeypatch.setenv("PARAKEET_MODEL_PATH", "other.nemo") ++ assert resolve_model_path("/env") == "/env/other.nemo" ++ ++ ++# -- retrying stretches where the model went quiet -------------------------------- ++ ++ ++def talk(seconds, level=0.1, seed=3): ++ return (level * np.random.default_rng(seed).standard_normal(int(seconds * SR))).astype(np.float32) ++ ++ ++def w_at(text, start, end): ++ return {"word": text, "start_offset": 0, "end_offset": 0, "start": start, "end": end} ++ ++ ++def test_a_long_wordless_stretch_over_speech_is_a_gap(): ++ audio = talk(40) ++ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 38.0, 38.5)] ++ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (10.0, 20.0), (20.5, 38.0)] ++ ++ ++def test_a_wordless_stretch_over_silence_or_a_short_pause_is_not(): ++ audio = talk(40) ++ audio[int(10 * SR):int(20 * SR)] = 0.0 ++ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 22.0, 22.5), ++ w_at("e", 24.0, 24.5), w_at("f", 39.0, 39.5)] ++ # 10-20 s is silent; 20.5-22 and 22.5-24 are too short; 1.5-9.5 and 24.5-39 are talk ++ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (24.5, 39.0)] ++ ++ ++def test_speech_before_the_first_word_counts(): ++ audio = talk(20) ++ assert find_speech_gaps([w_at("late", 15.0, 15.5)], audio, SR, min_gap=3.0) == [(0.0, 15.0), (15.5, 20.0)] ++ ++ ++def test_no_words_at_all_over_speech_is_one_gap(): ++ assert find_speech_gaps([], talk(12), SR, min_gap=3.0) == [(0.0, 12.0)] ++ ++ ++def test_retried_words_are_spliced_into_the_gap_and_segments_rebuilt(): ++ audio = talk(31) ++ words = [w_at("Hello", 1.0, 1.5), w_at("there.", 2.0, 2.5), w_at("Bye.", 30.0, 30.5)] ++ segments = [segment(words[:2]), segment(words[2:])] ++ calls = [] ++ ++ def transcribe(s0, s1): # stands in for Parakeet: finds speech the main pass missed ++ calls.append((s0 / SR, s1 / SR)) ++ found = [w_at("We", 5.0, 5.2), w_at("missed", 6.0, 6.3), w_at("this.", 7.0, 7.4), ++ w_at("Again", 12.0, 12.4), w_at("outside", 40.5, 41.0)] ++ return "", [w for w in found if s0 / SR <= w["start"] < s1 / SR], [] ++ new_words, new_segments, n = retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe) ++ assert [w["word"] for w in new_words] == ["Hello", "there.", "We", "missed", "this.", "Again", "Bye."] ++ assert n == 1 and calls == [(2.0, 30.5)] # one gap, 2.5-30.0 s, padded by 0.5 s ++ assert [g["segment"] for g in new_segments] == ["Hello there.", "We missed this.", "Again Bye."] ++ assert " ".join(g["segment"] for g in new_segments) == " ".join(w["word"] for w in new_words) ++ ++ ++def test_without_gaps_nothing_is_retried_and_output_is_untouched(): ++ audio = talk(10) ++ words = [w_at("One", 0.5, 1.0), w_at("two", 2.5, 3.0), w_at("three.", 5.0, 5.5), w_at("four", 8.0, 8.5)] ++ segments = [segment(words[:3]), segment(words[3:])] ++ ++ def transcribe(s0, s1): ++ raise AssertionError("must not be called") ++ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe) == (words, segments, 0) ++ ++ ++def test_long_gaps_are_retried_in_pieces(): ++ audio = talk(191) ++ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)] ++ calls = [] ++ ++ def transcribe(s0, s1): ++ calls.append(round((s1 - s0) / SR, 1)) ++ return "", [], [] ++ retry_speech_gaps(words, [], audio, SR, 3.0, transcribe) ++ assert max(calls) <= 61.0 and len(calls) == 4 # 1.5-190 s in <= 60 s pieces (+0.5 s padding) ++ ++ ++def test_segments_from_words_split_after_sentence_punctuation(): ++ ws = [w_at("Yes.", 0, 1), w_at("Is", 1, 2), w_at("it?", 2, 3), w_at("Go", 3, 4), w_at("now", 4, 5)] ++ assert [g["segment"] for g in segments_from_words(ws)] == ["Yes.", "Is it?", "Go now"] ++ assert segments_from_words([]) == [] ++ ++ ++def test_a_retried_copy_of_the_word_after_the_gap_is_not_kept_twice(): ++ # The retry piece is padded, so it re-hears the next kept word and may ++ # timestamp it just inside the gap. ++ audio = talk(31) ++ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)] ++ ++ def transcribe(s0, s1): ++ return "", [w_at("We", 5.0, 5.2), w_at("left.", 6.0, 6.3), w_at("bye.", 29.7, 30.3)], [] ++ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe) ++ assert [w["word"] for w in new_words] == ["Hello.", "We", "left.", "Bye."] ++ assert new_words[-1] is words[-1] ++ ++ ++def test_a_word_heard_by_two_retry_pieces_is_kept_once(): ++ audio = talk(191) ++ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)] ++ ++ def transcribe(s0, s1): # both pieces around 61.5 s hear "boundary" ++ t0, t1 = s0 / SR, s1 / SR ++ out = [w_at("boundary", 61.3, 61.8)] if t0 <= 61.3 < t1 else [] ++ out += [w_at("boundary", 61.55, 61.9)] if t0 <= 61.55 < t1 and t0 > 60 else [] ++ return "", out, [] ++ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe) ++ assert [w["word"] for w in new_words] == ["a", "boundary", "z"] ++ ++ ++def test_a_failing_retry_keeps_the_first_pass(): ++ audio = talk(31) ++ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)] ++ segments = [segment(words[:1]), segment(words[1:])] ++ ++ def transcribe(s0, s1): ++ raise RuntimeError("CUDA out of memory") ++ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe)[:2] == (words, segments) ++ ++ ++def test_a_looping_retry_is_not_spliced_in(): ++ audio = talk(31) ++ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)] ++ ++ def transcribe(s0, s1): # a hallucination loop over music-like audio ++ return "", [w_at("la", 3.0 + k, 3.2 + k) for k in range(20)], [] ++ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe) ++ assert [w["word"] for w in new_words] == ["Hello.", "Bye."] ++ ++ ++def test_gap_levels_line_up_with_time_at_sample_rates_other_than_16k(): ++ # At 11025 Hz a 10 ms hop is 110 samples (9.977 ms); indexing frames by ++ # time/10ms would drift ~2 s by 900 s and read the silence before the gap. ++ sr = 11025 ++ audio = (0.1 * np.random.default_rng(5).standard_normal(1000 * sr)).astype(np.float32) ++ audio[int(896.9 * sr):int(900.0 * sr)] = 0.0 ++ words = [w_at("w", t, t + 0.5) for t in np.arange(0.0, 1000.0, 1.0) if not 900.0 <= t < 903.0] ++ words = [w for w in words if not (899.9 < w["start"] < 900.5)] + [w_at("w", 899.5, 900.0)] ++ words.sort(key=lambda w: w["start"]) ++ assert (900.0, 903.0) in find_speech_gaps(words, audio, sr, 3.0) +-- +2.39.5 +