docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)

Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
This commit is contained in:
vh
2026-09-30 15:48:02 -07:00
parent 1189adbf18
commit 38015a1977
7 changed files with 1162 additions and 17 deletions
@@ -0,0 +1,389 @@
# Parakeet dropout investigation (2026-09-30)
Prime's ask, via the coordinator, 2026-09-30: investigate the "Parakeet skips
stretches of speech" finding from the slicer bench
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md`), and include a different
Parakeet weight. **Investigation only:** nothing here changed the live Scriberr
container or its `.env`; deploying anything is Prime's call.
**Privacy:** two of the four recordings are Prime's. Their audio, transcripts and
diarization stay in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700).
This document carries metrics only.
## Verdict
- **The drops are real.** Against ground truth, today's Parakeet (v3, shipped
slicer) loses whole stretches of clean speech: 140 words per transcript on a
clean audiobook, 66 on a court argument, and about 50 on each of Prime's two
recordings (Whisper-referenced, Canary-confirmed). The old "reference" was
mostly right; on SCOTUS it drops speech too.
- **Root cause:** the Granary-era 0.6B weights (v3 and v2) sometimes stop
producing words for tens of seconds deep inside a long full-attention window.
The encoder output there is degraded (a fresh decoder recovers only 22 to 51 %),
the same audio transcribed alone is fine, and older Parakeets with the **same
TDT decoder** do not do it. It is chaotic in cut placement and not explained by
loudness, SNR, speaking rate or language. Decoding settings, beam search,
shorter slices, loudness normalisation and resampling do not fix it.
- **Fix that works:** re-transcribe any ≥ 3 s stretch where the audio holds
speech but the model produced no words. That cuts the loss 80 to 90 % on all
four recordings (to 19, 13, 7 and 5 words per transcript) with no invented
text, at +10 MiB and no measurable time. It is patch 0002 (proposed, built and
tested, **not deployed**), which also adds an explicit `PARAKEET_MODEL_PATH`.
- **Weights:** v2 loses nothing on Prime's recordings or SCOTUS but collapses on
read speech; `parakeet-unified-en-0.6b` (newest) is the best-balanced
punctuating Parakeet but needs NeMo 3.0.0 and a licence change. The 1.1B and
CTC Parakeets never collapse but produce no punctuation.
- **Separate, smaller issue:** overlapping speakers (SCOTUS interjections) are
lost by every single-stream model. That is not this bug.
## 1. Is the instrument right? Ground truth, not a model reference
The first finding measured losses against the whole-file local-attention
transcript. That reference could itself be inserting text, so the public
recordings were re-scored against real ground truth.
| file | ground truth | provenance |
|---|---|---|
| scotus (first 30 min of No. 22-451) | the official argument transcript, `supremecourt.gov/oral_arguments/argument_transcripts/2023/22-451_114p.pdf` (sha256 `4feb7786…78a2`) | cover pages, page and line numbers, running headers, time stamps, argument headings and speaker labels removed; cut where the audio ends by alignment (5,627 of 14,764 transcript words) |
| wilde (LibriVox section 1) | Project Gutenberg #38916, *The Trial of Oscar Wilde, from the Shorthand Reports* (the LibriVox page's own "online text" link), sha256 `2271f271…b3a9` | section 1 is the book's Preface, read by the narrator alone ("It is wrong for us…" to "…came too late."), plus LibriVox's standard spoken intro and outro |
**Normalisation:** Whisper's English normaliser (transformers 4.53.3, the version
in Scriberr's env, no spelling map), applied word by word to both sides so every
hypothesis token keeps its timestamp. Fillers are dropped, numbers become digits
and contractions are expanded.
**Alignment and the dropout detector:** difflib opcodes, then exact Levenshtein
inside each mismatch. A **dropout** is a stretch between solid matches (≥ 3
tokens) holding ≥ 10 reference words where the hypothesis emitted fewer than half
as many. An **insertion run** is the mirror image. Matching islands shorter than
3 tokens count as part of the stretch, so one spurious "the" cannot split a
skipped paragraph in two.
**Timed ground truth:** every ground-truth token gets the median start time of
the transcripts that matched it (14 transcripts); the 143 (scotus) and 61
(wilde) tokens that no transcript matched are interpolated. There are no time
reversals over 1 s.
### Result: the drops are real, measured against ground truth
The same outputs the first finding used (the shipped slicer and upstream's fixed
cutter, 3 placements each), now scored against ground truth:
| file | system | WER | dropouts | words dropped |
|---|---|---|---|---|
| wilde | whole-file local attention (the old stand-in reference) | 2.12 % | 0 | 0 |
| wilde | upstream fixed 120 s cutter | 2.57 % | 0 | 0 |
| wilde | shipped slicer, at 120 / 110 / 100 s | 2.15 / **13.02** / 2.40 % | 0 / 3 / 0 | 0 / **394** / 0 |
| scotus | whole-file local attention | 6.01 % | 6 | 146 |
| scotus | upstream fixed 120 s cutter | 4.54 % | 4 | 46 |
| scotus | shipped slicer, at 120 / 110 / 100 s | 5.42 / 5.89 / 5.30 % | 4 / 5 / 4 | 94 / 137 / 111 |
On Wilde, a clean single-narrator audiobook, the stand-in reference was right
and the chunked runs genuinely lost whole paragraphs; which paragraphs depends
only on where the cuts fall. On SCOTUS the stand-in reference **also** drops
speech (146 words), so the first finding's "reference insertions" there were
really reference deletions.
### Adjudicating Prime's recordings without ground truth
Two independent models give a second opinion on every disputed stretch (≥ 10
words one system has and another lacks): **Whisper large-v3** (openai, pinned
`06f233fe`, transformers sequential long-form decoding; its word times are
spread evenly within segments, so it gets a ±3 s window) and **Canary-1b-v2**
(the copy inside Scriberr's env, sha-matched to HF `d4557063`; cut in pauses,
no overlap, ±1 s window). "Speech is real" needs both to have ≥ 50 % of the
disputed words; "no speech" needs both under 20 %; anything else is ambiguous.
**The adjudicator is calibrated on the public files first**, where ground truth
gives the right answer. It decided 129 of 140 disputed stretches and **got all
129 right**; it abstained on 11, mostly crosstalk that the official transcript
renders differently. Whisper large-v3 on its own scores 1.96 % (wilde) and 3.93 %
(scotus) WER against ground truth with **zero** dropouts, so for scoring Prime's
files it serves as the reference, and a dropout counts only if Canary
independently has the words. On the public files that metric reproduces the
ground-truth clean-speech dropout counts within a few percent (1,060 vs 1,116;
287 vs 310; 240 vs 246; 5,450 vs 5,513; 515 vs 526 words).
### Controls and floor
- **A-vs-A.** Production v3 (full attention) is byte-identical across three
separate processes at every placement tested. **Local attention is not:** at
the same SCOTUS placement three runs dropped 68, 169 and 100 words (WER 4.73 to
6.32 %). Its numbers below carry that run-to-run noise.
- **Positive control.** Copies of both public files with 30, 15, 5 and 4 s of
audio replaced by digital silence (the ground truth still holds those words):
the 30, 15 and 5 s stretches (15 to 73 words) were reported in all 36 runs; the
4 s stretch (11 to 14 words) in 11 of 12.
- **Null control.** In those same runs, chunks that touch no silenced stretch get
byte-identical input; with full attention their words were identical and their
dropouts equal in all 6 runs, so the instrument manufactures nothing. (With
local attention they differ; that is the non-determinism above.)
- **A "should-not-matter" perturbation** (−0.5 dB gain) left SCOTUS identical and
Wilde within its placement spread (1,027 vs 1,116 words). Parakeet normalises
each mel feature per chunk, so a constant gain is nearly a no-op by design.
- **Sensitivity floor:** a dropout of ≥ 15 words is caught every time (36/36);
10 to 14 words, 11/12; shorter losses are not counted as dropouts at all (they
still count in WER). Because the failure is chaotic in cut placement, every
configuration is run at 8 placements (4 for the private files and the 1.1B
diagnostics); totals over 8 placements that differ by less than about a
quarter are not a difference.
## 2. What the real drops look like
Production v3 (the shipped slicer: 120 s chunks, 4 s overlap), 8 cut placements
per public file, each dropout compared with 20 random windows of the same length
from the same file (clean speech; SCOTUS crosstalk split out below):
| | Wilde (one narrator) | SCOTUS (argument) |
|---|---|---|
| dropouts / words | 17 / 1,116 | 29 / 730 |
| start, seconds into its chunk | median 47, **never before 16** | median 53, never before 13 |
| runs to the end of its chunk | 35 % | 10 % |
| level vs file speech level (drops / controls) | −1.0 / −1.35 dB | −0.1 / −2.0 dB |
| local SNR (drops / controls) | 42 / 41 dB | 29 / 27 dB |
| speaking rate (drops / controls) | 2.4 / 2.5 words/s | 4.0 / 3.2 words/s |
| pause just before (drops / controls) | 0.62 / 0.37 s | 0.02 / 0.02 s |
| overlapped speech, Sortformer (drops / controls) | 0 / 0 | **16 % / 0 %** of the window |
| non-English words within ±10 s | 0 | 0 |
Two different things are being counted:
- **A. Long-context collapse.** Clean speech lost in stretches of tens of seconds
(up to 60 s), beginning at a sentence boundary well into a long chunk, often
running to the chunk's end. It is **not explained by the audio**: level, SNR
and rate match the controls, there is one speaker, and nothing is non-English.
Which stretches go is chaotic: it moves with the cut placement, and on Wilde
no word is lost by more than 75 % of placements.
- **B. Crosstalk.** On SCOTUS, a justice's interjection over counsel ("Well,
wait a minute…") is lost at the same few spots in almost every run. A
single-stream model transcribes one voice; the official transcript records
both. That is a limit of single-stream ASR and of the transcript convention,
not the failure this investigation is about, so it is **separated out**
(diarized overlap ≥ 10 % of the stretch) in every comparison below.
Every comparison below counts **clean-speech dropped words per transcript**, mean
with a 95 % bootstrap interval over cut placements.
**Language ID is not the cause.** v3's only non-ASCII output is legitimate French
names in the text (Mallarmé, Comédie), never near a drop, and English-only v2
collapses too. Prime's recordings are English (function-word share 0.35 to 0.39,
the same as the public English files at 0.39 to 0.44; no non-ASCII letters).
## 3. Mechanism
- **The encoder output is degraded, not just the decoder.** On the three chunks
that went fully quiet (0 words after the drop began), a fresh decoder state
started on the same full-attention encoder output recovered only 22 to 51 % of
the ground-truth words there. The same audio encoded on its own recovered 72 to
102 %, and local attention over the full chunk 97 to 102 % (a fourth, partial
drop: 98 %, 57 %, 98 %; a control chunk: all ≈ production). That is n = 4 hand-picked
chunks, so it points the way; the unbiased tests below carry the weight.
- **It belongs to the weights, not to TDT decoding.** On the same slices,
`parakeet-tdt-1.1b` (TDT), `parakeet-rnnt-1.1b`, `parakeet-ctc-1.1b`,
`parakeet-ctc-0.6b` and both heads (TDT and CTC) of `parakeet-tdt_ctc-1.1b` all
lose **0** words on Wilde. v3 and v2, the Granary-era 0.6B TDT models, collapse.
The new `parakeet-unified-en-0.6b` (RNN-T) collapses once in 8.
- **It is a knife edge.** Deterministic for a given input (v3 at full attention is
byte-identical across processes), but the float-level non-determinism of the
local-attention kernel is enough to flip whether a stretch is transcribed
(68 / 169 / 100 words at one placement).
- **Recording style matters, differently per weight.** Wilde and p2 are gated
recordings (pauses near −64 and −74 dBFS); SCOTUS and p1 never go quiet (room
tone at −40 and −36 dBFS). v2 collapses badly on Wilde but never on the other
three; a −50 dBFS noise floor more than halves v3's and v2's Wilde losses but
makes SCOTUS worse. Not a clean lever.
## 4. Hypotheses tested (unbiased: full files at 8 placements; clean-speech words per transcript)
| hypothesis | test | Wilde | SCOTUS | Prime p1 / p2 | verdict |
|---|---|---|---|---|---|
| — | **production v3** | 140 [48–240] | 66 [41–88] | 50 [40–63] / 51 [25–82] | baseline |
| (a) decoding | CUDA graphs on | identical output | identical | | no effect |
| (a) | greedy (per-frame) instead of greedy_batch | identical output | identical | | no effect |
| (a) | max_symbols 20 | identical at its 3 placements | identical | | no effect |
| (a) | TDT beam 4 (text only; NeMo 2.5.3 cannot give timestamps with TDT beam), all dropouts, vs greedy under the same slicing | 194 vs 136 | 488 vs 57 | | **worse**, 3–13× slower, unusable in Scriberr |
| (b) context | 60 s slices | 88 [41–138] | 61 [41–85] | 31 / 31 | within the floor |
| (b) | 30 s slices | 165 [95–236] | 77 [49–105] | 32 / 13 | worse on Wilde |
| (b) | local attention ±64 frames in 120 s slices | 34 [21–47] | 97 [72–121] | | trades one file for the other |
| (b) | local attention ±128 | 39 [24–51] | 128 [92–168] | 0 / 3 | trades; **non-deterministic** |
| (b) | local attention ±256 | 31 [15–47] | 80 [43–120] | 60 / 64 | no help on Prime's files |
| (c) preprocessing | −0.5 dB gain (null) | 128 | identical | 54 / 30 | no effect (the null) |
| (c) | loudness normalisation to −20 dBFS, peak-safe | **498** [382–613] | 66 | | **worse** on quiet audio |
| (c) | soxr resampling instead of ffmpeg (all dropouts; production on the same measure: 140 / 91) | 206 | 92 | | within the floor |
| (c) | −50 dBFS noise floor | 59 [30–87] | 87 [74–100] | | helps Wilde, hurts SCOTUS |
| (d) weights | see § 5 | | | | the lever |
| new | **re-transcribe speech gaps ≥ 3 s** (patch 0002) | **19 [0–38]** | **13 [3–25]** | **7 [0–22] / 5 [0–15]** | **fixes most of it** |
WER against ground truth moves the same way: production v3 4.83 % / 5.33 % median
(Wilde / SCOTUS) against 2.40 % / 4.47 % with the gap retry.
## 5. Candidates
Clean-speech **dropped words per transcript** (mean, 95 % bootstrap interval over
placements; public: against ground truth, 8 placements; Prime's: against Whisper,
confirmed by Canary, 4 placements) and **median WER** (public: against ground
truth; Prime's: against Whisper). **Memory**: per-process peak on the 35-min file
(p1), nvidia-smi every 0.2 s, only the run's own container PIDs, n = 3, zero
spread in every case. **Speed**: median wall time per job for the 35.3-min file
including model load, n = 3 (spread ±5 s). "CLI" rows ran the actual patched
Scriberr scripts under Scriberr's invocation; "lab" rows ran the lab harness
because Scriberr's env cannot run them as-is.
| candidate | Wilde words / WER | SCOTUS words / WER | p1 words / WER | p2 words / WER | peak MiB | job time | old GPU 1 budget (5,496) | GPU 3 (+Blender 270 MiB) | punctuation |
|---|---|---|---|---|---|---|---|---|---|
| **v3, production** (0001) | 140 [48–240] / 4.83 % | 66 [41–88] / 5.33 % | 50 / 2.96 % | 51 / 2.93 % | 5,496 (CLI) | 53 s | fits (0 spare) | fits | yes |
| **v3 + gap retry** (0001+0002) | **19** [0–38] / **2.40 %** | **13** [3–25] / **4.47 %** | **7** / 2.52 % | **5** / 2.38 % | 5,506 (CLI) | 48 s | **10 MiB over** | fits | yes |
| v2 | 689 [490–971] / 17.6 % | **0** / 4.55 % | **0** / 2.22 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
| v2 + gap retry | 151 [116–192] / 5.84 % | **0** / 4.58 % | **0** / 2.25 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
| parakeet-unified-en-0.6b (NeMo 3.0.0) | 31 [0–92] / **2.21 %** | 22 [14–32] / 4.92 % | **0** / **2.19 %** | **0** / **1.83 %** | 5,438 (lab) | 34 s | fits | fits | yes |
| parakeet-tdt-1.1b | **0** / 2.31 % | 3 / 6.78 % | **0** / 3.30 % | **0** / 2.65 % | 8,884 (lab) | 46 s | no | fits | **no** |
| parakeet-ctc-0.6b | **0** / 2.46 % | 3 / 6.41 % | – | – | 5,360 (lab) | 31 s | fits | fits | **no** |
| canary-1b-v2 (30 s pause cuts) | 11 [4–21] / 9.64 % ⚠ | 5 / 5.25 % | **0** / 3.88 % | **0** / 2.84 % | 10,504 (lab) | 112 s | no | fits | yes |
| *Whisper large-v3 (reference, 1 run)* | *0 / 1.96 %* | *0 / 3.93 %* | – | – | – | – | – | – | *yes* |
⚠ Canary's Wilde WER is its **hallucination loops**: 6 insertion runs across 8
placements, one of 417 words. It barely drops speech but invents it.
Other 1.1B diagnostics (no punctuation, so not candidates): `parakeet-rnnt-1.1b`
0 / 7, `parakeet-ctc-1.1b` 0 / 3, `parakeet-tdt_ctc-1.1b` TDT head 0 / 6 and CTC
head 0 / 0 (Wilde / SCOTUS words per transcript, 4 placements).
Provenance of every weight (all pulled revision-pinned, sha256 equal to the HF
LFS oid, in `/tank/aimodels/huggingface/hub/`):
| repo | revision | licence (read at the raw card) | file sha256 |
|---|---|---|---|
| nvidia/parakeet-tdt-0.6b-v3 (in Scriberr's env) | `541d1f99` | CC-BY-4.0 | `3cbdc858…` |
| nvidia/parakeet-tdt-0.6b-v2 | `ae9ad070` | CC-BY-4.0 | `d99e3995…` |
| nvidia/parakeet-unified-en-0.6b (2026-04-07, newest Parakeet) | `fe53cd88` | **NVIDIA Open Model License** | `ec23ed91…` |
| nvidia/parakeet-tdt-1.1b | `53276c64` | CC-BY-4.0 | `9c563d52…` |
| nvidia/parakeet-rnnt-1.1b | `2acc4c61` | CC-BY-4.0 | `535896f0…` |
| nvidia/parakeet-ctc-1.1b | `20e63a0f` | CC-BY-4.0 | `8e91253d…` |
| nvidia/parakeet-tdt_ctc-1.1b | `675e7868` | CC-BY-4.0 | `4e7ccfdd…` |
| nvidia/parakeet-ctc-0.6b | `ad09ba1c` | CC-BY-4.0 | `bc01f3f8…` |
| nvidia/canary-1b-v2 (in Scriberr's env) | `d4557063` | CC-BY-4.0 | `ae5ef1bf…` |
| openai/whisper-large-v3 (adjudicator only) | `06f233fe` | Apache-2.0 | `a8e94b85…` |
All repo ids were verified with an authenticated HF API call before any pull (no
phantoms). `parakeet-unified-en-0.6b` needs **NeMo 3.0.0**: its card says 2.7.3,
but released 2.7.3 lacks its encoder argument (`att_chunk_context_size`), and its
`.nemo` ships without a `validation_ds` config that `transcribe()` reads (a
two-line shim). It ran in a throwaway env, not Scriberr's (NeMo 2.5.3).
## 6. The fix that works on any weight: re-transcribe speech gaps
The mechanism says the lost audio is fine on its own; only its long-window
context breaks. So after stitching, the buffered script looks for stretches of
**≥ 3 s with no word where at least half the 25 ms frames sit within 12 dB of the
recording's typical speech level**, re-transcribes each on its own (pieces of at
most 60 s, 0.5 s of padding) and splices in the words that land inside it.
Output that triggers no retry is byte-identical to today's.
- **Effect:** v3's clean-speech losses fall 80 to 90 % on all four recordings
(140 → 19, 66 → 13, 50 → 7, 51 → 5 words per transcript) and WER falls with
them (Wilde 4.83 → 2.40 %, SCOTUS 5.33 → 4.47 %). It adds **no insertion runs**,
so it is not making text up.
- **Trigger:** 3 s beat 6 s on Prime's files (p1 7 vs 27 words per transcript)
and matched it on the public ones; 6 retries per 35-minute transcript were
typical.
- **Cost:** peak 5,506 MiB vs 5,496 (the retry's shorter chunk shifts the
allocator by 10 MiB); job time unchanged within the ±5 s run-to-run spread.
- **Residual:** stretches where the model still emits a few words (no clean gap),
and retries that come back short. It does not rescue v2 on the audiobook
(689 → 151).
- **Validity:** the production implementation (after code review) reproduces
the measured lab version at every shared placement: identical dropped words
on all 16 file-placement pairs tried and WER equal or up to 0.02 points lower
(it now removes the odd retried word that repeated its neighbour). It is
deterministic across processes, and the JSON seam checks pass for both
scripts with the rebuilt segments.
- **Hardening from review:** a failing retry piece is skipped and a failing retry
keeps the first pass (never loses a finished transcript); a piece that looks
like a hallucination loop (mostly one repeated token, or > 7 words/s) is not
spliced in; retried copies of the words at a gap's edge are dropped; audio goes
through a per-run temp directory; frame levels use bounded memory (the first
version needed ~1.7 GB of RAM for 35 min); frame indexing is correct at any
sample rate. **Residual risk:** the gap detector is an energy test, so a loud
non-speech stretch (a music bed) is retried; the loop guard is the only check
on what comes back. None of the four recordings has music.
## 7. Recommendation
| claim | strength | basis | reversibility |
|---|---|---|---|
| The drops are real speech lost by Parakeet, not a reference artefact | **insist** | measured against ground truth on 2 files, 129/129 adjudications correct on the calibration | n/a |
| Ship patch 0002 (gap retry on by default, v3 weights) | **strongly recommend** | measured: 80–90 % fewer lost words on all 4 recordings, lower WER, no new insertions, n = 4–8 placements each | reversible (image rollback) |
| Keep v3 rather than switch to v2 | **lean** | measured: v2 is perfect on Prime's two files and SCOTUS but loses 151 words per transcript on read speech even with the retry, and is English-only; the right answer depends on what Prime transcribes | reversible (one env var) |
| Evaluate `parakeet-unified-en-0.6b` as the next weight, as its own project | **lean** | measured: the best-balanced punctuating Parakeet (0 / 0 on Prime's files, lowest WER there); but it needs NeMo 3.0.0 in Scriberr's env (Canary and Sortformer too), a packaging shim, and a licence change from CC-BY-4.0 to the NVIDIA Open Model License | costly to reverse (env rebuild) |
| Do not use local attention, beam search, shorter slices, loudness normalisation or a noise floor | **recommend against** | measured: each worse or mixed, local attention also non-deterministic | reversible |
| Do not switch to Canary-1b-v2 | **recommend against** | measured: hallucination loops (up to 417 invented words) and twice the memory | reversible |
| Do not use the 1.1B or CTC Parakeets | **recommend against** | read at source: no punctuation or casing, which Scriberr's segmentation needs | reversible |
## 8. The integration patch
`stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch`
(on top of 0001, upstream `a353078`; sha256 `e3098bf2…7e62`). It is in
`proposed/`, so `scripts/scriberr-rebuild` does **not** apply it until it is moved
up a directory. What it does:
1. **Gap retry** (§ 6) in both scripts, `--retry-gaps SECS` (default 3, 0 off).
Go never passes the flag, so the default is the behaviour; the JSON gains
`retried_gaps`.
2. **`PARAKEET_MODEL_PATH`**: an absolute path (or one relative to the env) to
the `.nemo` both scripts load; default unchanged. The JSON `model` field now
reports the file actually loaded, and the Go adapter records it as
`ModelUsed` instead of the hardcoded "parakeet-tdt-0.6b-v3". **No model is
ever swapped in under v3's filename.**
3. Tests: 18 new unit tests (39 in total), and upstream's own standard and
buffered tests pass in the built image. Reviewed at high effort; all ten
findings fixed (§ 6).
**Validated as a build** (`scripts/scriberr-rebuild --suffix dropout2 --budget
5600 --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`): all
stages PASS, including the embed of both scripts, 39 unit tests, the JSON seam
for the long- and the short-audio script, and memory (5,506 MiB). The image
`scriberr:local-blackwell-a353078-dropout2` exists on fv-ml1 and is **not
deployed**.
**To ship it (Prime's call):** either deploy the already-built
`scriberr:local-blackwell-a353078-dropout2` per `patches/README.md` § Deploy, or
move the patch up into `stacks/scriberr/patches/` first (so it becomes part of
the carried set) and rebuild under a new suffix with `--budget 5600` (see the
note on the budget).
**To also switch weights (only if Prime chooses v2 or, later, another weight):**
in `stacks/scriberr/compose.yaml`, mount the shared model cache read-only and
name the pinned file. The revision is visible in the path and recorded in every
transcript's metadata:
```yaml
volumes:
- /tank/aimodels/huggingface:/models:ro
environment:
- PARAKEET_MODEL_PATH=/models/hub/models--nvidia--parakeet-tdt-0.6b-v2/snapshots/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/parakeet-tdt-0.6b-v2.nemo
```
**How it survives upgrades:** `scriberr-rebuild` re-applies 0001 and 0002 to any
pinned upstream sha and stops on a conflict; the model choice lives in our
compose file, not in the image or the env directory, so an upgrade cannot
silently change it. If upstream ever grows its own model selection, 0002 is
dropped in favour of it.
**Budget note:** the rebuild script's default memory budget is still GPU 1's old
5,496 MiB. With Scriberr on GPU 3 that number no longer protects anything, but
the patched build peaks at 5,506, so either pass `--budget` or retire the old
default (a one-line change; left for Prime or the coordinator).
## 9. Harness, data and reproduction
All code is in fv-ml1 `/tank/spikes/scriberr-slicer/code/dropout/`: `lab.py`
(the production pipeline with every knob), `gtscore.py`, `boot.py`,
`adjudicate.py`, `adj_score.py`, `characterize.py`, `probe.py`, `fit.sh`,
`fitlab.sh`, plus the run scripts and configs. Metrics are in `…/metrics/`. The
public audio, ground truth and transcripts are in `…/public/` and `…/gt/`;
Prime's are in `…/private/` (mode 700). Throwaway envs: `…/envs/nemo300`
(NeMo 3.0.0, for the unified model).
Every run was a transient `--rm` container on GPU 3, with Scriberr's env and the
model cache mounted read-only. GPU 1 and the live Scriberr container were not
touched.
+1 -1
View File
@@ -181,7 +181,7 @@ _As of 2026-09-30 ~0120 PT._
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness. - **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
- **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. - **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`.
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`). - **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout1`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`). - Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
+23 -7
View File
@@ -31,6 +31,12 @@
# Usage: # Usage:
# scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N] # scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N]
# [--budget MIB] [--memory-audio PATH] [--reuse-image] # [--budget MIB] [--memory-audio PATH] [--reuse-image]
# [--patches DIR]
#
# --patches DIR[:DIR...] applies every *.patch in those directories, sorted by file
# name, instead of the carried set in stacks/scriberr/patches. To test-build a
# proposed patch on top of the carried ones, under its own --suffix:
# --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed
# #
# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496, # Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496,
# --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB # --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB
@@ -44,10 +50,11 @@ HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54}
ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY
TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1 TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1
SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
STD_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 PATCH_DIR_ARG=""
MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav
while [ $# -gt 0 ]; do while [ $# -gt 0 ]; do
case $1 in case $1 in
@@ -57,6 +64,7 @@ while [ $# -gt 0 ]; do
--budget) BUDGET=$2; shift 2 ;; --budget) BUDGET=$2; shift 2 ;;
--memory-audio) MEM_AUDIO=$2; shift 2 ;; --memory-audio) MEM_AUDIO=$2; shift 2 ;;
--reuse-image) REUSE_IMAGE=1; shift ;; --reuse-image) REUSE_IMAGE=1; shift ;;
--patches) PATCH_DIR_ARG=$2; shift 2 ;;
-h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;; -h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;;
*) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;; *) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;;
esac esac
@@ -66,8 +74,10 @@ done
[[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; } [[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; }
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches PATCH_DIR=${PATCH_DIR_ARG:-$REPO_ROOT/stacks/scriberr/patches}
mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort) IFS=: read -r -a PATCH_DIRS <<<"$PATCH_DIR"
mapfile -t PATCHES < <(for d in "${PATCH_DIRS[@]}"; do find "$d" -maxdepth 1 -name '*.patch'; done \
| awk -F/ '{print $NF "\t" $0}' | sort | cut -f2-)
[ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; } [ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; }
# The only paths a reused build dir may differ from upstream in. # The only paths a reused build dir may differ from upstream in.
mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u) mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u)
@@ -192,14 +202,14 @@ else
fi fi
# ── embed ────────────────────────────────────────────────────────────────── # ── embed ──────────────────────────────────────────────────────────────────
if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF' if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" "$STD_REL" 2>&1 <<'EOF'
docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c " docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c "
import sys import sys
script = open('/src/$3', 'rb').read() binary = open('/app/scriberr', 'rb').read()
sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)" sys.exit(0 if all(open('/src/' + f, 'rb').read() in binary for f in ('$3', '$4')) else 1)"
EOF EOF
); then ); then
pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr" pass embed "both Parakeet scripts are byte-identical inside /app/scriberr"
else else
fail embed "the binary does not embed the patched script ${out:+($out)}" fail embed "the binary does not embed the patched script ${out:+($out)}"
fi fi
@@ -240,6 +250,12 @@ gpu_idle() { # "room", not "idle": Scriberr itself may be running a job on this
gpu_idle seam gpu_idle seam
if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")" if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")"
else fail seam "$out"; fi else fail seam "$out"; fi
# The short-audio script (files under PARAKEET_CHUNK_THRESHOLD_SECS) has its own entry point.
if out=$("${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \
-v $BUILD_DIR/tests/data:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$STD_REL /audio/$(basename "$SEAM_AUDIO_REL") \
--output /tmp/out.json --context-left 256 --context-right 256 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \
python3 /tools/seam-check.py /tmp/out.json --standard'" 2>&1); then pass seam-short "$(tail -1 <<<"$out")"
else fail seam-short "$out"; fi
# ── memory ───────────────────────────────────────────────────────────────── # ── memory ─────────────────────────────────────────────────────────────────
"${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1" "${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1"
+7 -2
View File
@@ -7,7 +7,8 @@ job. This checks the shape Go reads plus the stitching invariants the slicer
patch promises. Stdlib only, so it runs under any python3. Prints counts, never patch promises. Stdlib only, so it runs under any python3. Prints counts, never
transcript text. transcript text.
usage: scriberr-seam-check.py RESULT.json [--min-chunks N] usage: scriberr-seam-check.py RESULT.json [--min-chunks N] [--standard]
--standard the short-audio script's result (no buffered/num_chunks keys)
""" """
import argparse import argparse
import json import json
@@ -42,6 +43,8 @@ def main():
parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py") parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py")
parser.add_argument("--min-chunks", type=int, default=1, parser.add_argument("--min-chunks", type=int, default=1,
help="fail unless the run used at least this many chunks") help="fail unless the run used at least this many chunks")
parser.add_argument("--standard", action="store_true",
help="the short-audio script's result: no buffered/num_chunks keys")
args = parser.parse_args() args = parser.parse_args()
min_chunks = args.min_chunks min_chunks = args.min_chunks
try: try:
@@ -56,6 +59,7 @@ def main():
for key, kind in required.items(): for key, kind in required.items():
if not isinstance(data.get(key), kind): if not isinstance(data.get(key), kind):
fail(f"'{key}' missing or not {kind.__name__}") fail(f"'{key}' missing or not {kind.__name__}")
if not args.standard:
if data.get("buffered") is not True: if data.get("buffered") is not True:
fail("'buffered' is not true") fail("'buffered' is not true")
if not isinstance(data.get("chunk_duration_secs"), NUMBER): if not isinstance(data.get("chunk_duration_secs"), NUMBER):
@@ -79,7 +83,8 @@ def main():
fail("segments do not cover the stitched words exactly once, in order") fail("segments do not cover the stitched words exactly once, in order")
print(f"SEAM OK: {len(words)} words, {len(segments)} segments, " print(f"SEAM OK: {len(words)} words, {len(segments)} segments, "
f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points") f"{data.get('num_chunks', 1)} chunks, cuts at {len(data.get('cut_times', []))} points"
f"{', model ' + data['model'] if data.get('model') else ''}")
if __name__ == "__main__": if __name__ == "__main__":
+4 -1
View File
@@ -181,4 +181,7 @@ Peak GPU memory on a 35-minute file:
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
patch; that is still open. patch. **Investigated 2026-09-30** (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`):
the losses are real (against ground truth) and belong to the v2/v3 weights over
long windows; a proposed patch, `patches/proposed/0002`, re-transcribes speech that
got no words and cuts them 80–90 %. It is not deployed; that is Prime's call.
+13
View File
@@ -8,6 +8,8 @@ distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status | | patch | against | status |
|---|---|---| |---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) | | `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `proposed/0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **proposed, not applied** (the rebuild script reads only this directory, not `proposed/`); built and tested as `scriberr:local-blackwell-a353078-dropout2`, not deployed. Why and how: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`, unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
@@ -104,6 +106,17 @@ each other; the default is the simplest of them.
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12 of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc. transcripts, for upstream's slicer too). See the bench doc.
### 0002 (proposed) — gap retry and an explicit model path
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
the Go adapter record it as `ModelUsed`. To adopt: move it up into this directory and
rebuild. To test-build it on top of the carried set under its own suffix:
`scripts/scriberr-rebuild --suffix <name> --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
### Upgrading upstream ### Upgrading upstream
```bash ```bash
@@ -0,0 +1,719 @@
From d253aa2b0ea2562ed4e1702380fac7e22bd11828 Mon Sep 17 00:00:00 2001
From: Vuong Hoang <vh@phasefinal.com>
Date: Wed, 30 Sep 2026 15:00:46 -0700
Subject: [PATCH] feat(parakeet): selectable model and a retry for speech that
got no words
Parakeet v2/v3 sometimes stop emitting for tens of seconds deep inside a long
full-attention chunk while someone is talking; the same audio transcribed on
its own is usually fine. After stitching, any stretch of at least
--retry-gaps seconds (default 3) with no word, where most 25 ms frames sit
within 12 dB of the recording's typical speech level, is re-transcribed on
its own (in pieces of at most 60 s) and the words that land inside it are
spliced in; segments are then rebuilt from the words with the model's own
rule (end after . ? !). Output that triggers no retry is unchanged. The
same retry runs in the short-audio script, which shares the helpers.
The retry is best effort: a failing piece is skipped and a failing retry
keeps the first pass. A piece that looks like a hallucination loop (mostly
one repeated token, or more than 7 words a second) is not spliced in, and a
retried word that repeats its neighbour at the gap edge or a piece boundary
is dropped (the first pass's copy wins). Chunk and retry audio go through a
per-run temporary directory instead of fixed /tmp paths.
PARAKEET_MODEL_PATH (absolute, or relative to the env) selects the .nemo
both scripts load; the default stays parakeet-tdt-0.6b-v3.nemo. The JSON
"model" field reports the file actually loaded, and the Go adapter records
it as ModelUsed instead of a hardcoded name.
The CLI and JSON seam is otherwise unchanged: the new flag is optional, and
the JSON gains retried_gaps.
---
.../adapters/parakeet_adapter.go | 7 +-
.../adapters/py/nvidia/parakeet_transcribe.py | 53 +++-
.../py/nvidia/parakeet_transcribe_buffered.py | 231 +++++++++++++-----
.../py/nvidia/tests/test_parakeet_slicing.py | 159 +++++++++++-
4 files changed, 380 insertions(+), 70 deletions(-)
diff --git a/internal/transcription/adapters/parakeet_adapter.go b/internal/transcription/adapters/parakeet_adapter.go
index 4fa252d..5f3bb9e 100644
--- a/internal/transcription/adapters/parakeet_adapter.go
+++ b/internal/transcription/adapters/parakeet_adapter.go
@@ -342,7 +342,9 @@ func (p *ParakeetAdapter) Transcribe(ctx context.Context, input interfaces.Audio
}
result.ProcessingTime = time.Since(startTime)
- result.ModelUsed = "parakeet-tdt-0.6b-v3"
+ if result.ModelUsed == "" {
+ result.ModelUsed = "parakeet-tdt-0.6b-v3"
+ }
result.Metadata = p.CreateDefaultMetadata(params)
logger.Info("Parakeet transcription completed",
@@ -525,6 +527,8 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu
End float64 `json:"end"`
} `json:"segment_timestamps"`
Confidence interface{} `json:"confidence,omitempty"`
+ // The scripts report the model they actually loaded (PARAKEET_MODEL_PATH can change it).
+ Model string `json:"model"`
}
if err := json.Unmarshal(data, &parakeetResult); err != nil {
@@ -538,6 +542,7 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu
Segments: make([]interfaces.TranscriptSegment, len(parakeetResult.SegmentTimestamps)),
WordSegments: make([]interfaces.TranscriptWord, len(parakeetResult.WordTimestamps)),
Confidence: 0.0, // Default confidence
+ ModelUsed: parakeetResult.Model,
}
// Convert segments
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
index 235e52a..6b2539e 100644
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
@@ -7,9 +7,17 @@ import argparse
import json
import sys
import os
+import shutil
+import tempfile
from pathlib import Path
+import librosa
import nemo.collections.asr as nemo_asr
+# Shared with the long-audio script, which lives in the same directory.
+from parakeet_transcribe_buffered import (
+ DEFAULT_RETRY_GAP_SECS, resolve_model_path, retry_speech_gaps, transcribe_chunk,
+)
+
def transcribe_audio(
audio_path: str,
@@ -18,14 +26,11 @@ def transcribe_audio(
context_left: int = 256,
context_right: int = 256,
include_confidence: bool = True,
+ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS,
):
"""
Transcribe audio using NVIDIA Parakeet model.
"""
- # Determine model path
- model_filename = "parakeet-tdt-0.6b-v3.nemo"
- model_path = None
-
# Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv
virtual_env = os.environ.get("VIRTUAL_ENV")
if not virtual_env:
@@ -33,10 +38,11 @@ def transcribe_audio(
sys.exit(1)
project_root = os.path.dirname(virtual_env)
- model_path = os.path.join(project_root, model_filename)
+ model_path = resolve_model_path(project_root)
+ model_name = os.path.splitext(os.path.basename(model_path))[0]
if not os.path.exists(model_path):
- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}")
+ print(f"Error during transcription: Can't find model file: {model_path}")
sys.exit(1)
print(f"Loading NVIDIA Parakeet model from: {model_path}")
@@ -80,6 +86,27 @@ def transcribe_audio(
text = result_data.text
word_timestamps = result_data.timestamp.get("word", [])
segment_timestamps = result_data.timestamp.get("segment", [])
+ retried = 0
+ words_added = False
+ if retry_gap_secs > 0:
+ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp
+ try:
+ audio, sr = librosa.load(audio_path, sr=None, mono=True)
+
+ def transcribe_span(start_sample, end_sample):
+ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr,
+ start_sample / sr, os.path.join(workdir, "retry.wav"))
+ before = len(word_timestamps)
+ word_timestamps, segment_timestamps, retried = retry_speech_gaps(
+ word_timestamps, segment_timestamps, audio, sr, retry_gap_secs, transcribe_span
+ )
+ words_added = len(word_timestamps) != before
+ if words_added:
+ text = " ".join(w["word"] for w in word_timestamps)
+ except Exception as err: # best effort: never lose a finished transcription
+ print(f"Warning: gap retry failed ({err}); keeping the first pass")
+ finally:
+ shutil.rmtree(workdir, ignore_errors=True)
print(f"Transcription: {text}")
@@ -90,14 +117,15 @@ def transcribe_audio(
"word_timestamps": word_timestamps,
"segment_timestamps": segment_timestamps,
"audio_file": audio_path,
- "model": "parakeet-tdt-0.6b-v3",
+ "model": model_name,
+ "retried_gaps": retried,
"context": {
"left": context_left,
"right": context_right
}
}
- if include_confidence:
+ if include_confidence and not words_added: # first-pass scores no longer line up
# Add confidence scores if available
if hasattr(result_data, 'confidence') and result_data.confidence:
output_data["confidence"] = result_data.confidence
@@ -119,7 +147,7 @@ def transcribe_audio(
"transcription": text,
"language": "en",
"audio_file": audio_path,
- "model": "parakeet-tdt-0.6b-v3"
+ "model": model_name
}
if output_file:
@@ -154,6 +182,12 @@ def main():
"--context-right", type=int, default=256,
help="Right attention context size (default: 256)"
)
+ parser.add_argument(
+ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS,
+ help=f"Re-transcribe on its own any stretch of at least this many seconds where "
+ f"the audio holds speech but the model produced no words "
+ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)"
+ )
parser.add_argument(
"--include-confidence", action="store_true", default=True,
help="Include confidence scores"
@@ -178,6 +212,7 @@ def main():
context_left=args.context_left,
context_right=args.context_right,
include_confidence=args.include_confidence,
+ retry_gap_secs=args.retry_gaps,
)
except Exception as e:
print(f"Error during transcription: {e}")
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
index ba755c1..49bb72e 100644
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
@@ -12,15 +12,21 @@ import argparse
import json
import sys
import os
+import shutil
+import tempfile
import librosa
import soundfile as sf
import numpy as np
from pathlib import Path
+DEFAULT_MODEL = "parakeet-tdt-0.6b-v3.nemo"
DEFAULT_OVERLAP_SECS = 4.0
DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap
QUIET_WINDOW_SECS = 0.3
SAME_WORD_SECS = 0.5
+DEFAULT_RETRY_GAP_SECS = 3.0
+RETRY_PIECE_SECS = 60.0
+RETRY_MAX_WORDS_PER_SEC = 7.0
def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0):
@@ -103,14 +109,7 @@ def stitch_slices(slice_results, cut_times, chunk_spans=None):
elif len(kept) == stop - start:
segments.append(seg)
elif kept:
- segments.append({
- **seg,
- "segment": " ".join(w["word"] for w in kept),
- "start_offset": kept[0]["start_offset"],
- "end_offset": kept[-1]["end_offset"],
- "start": kept[0]["start"],
- "end": kept[-1]["end"],
- })
+ segments.append({**seg, **_segment_of(kept)})
return words, segments
@@ -159,6 +158,134 @@ def _segment_ranges(words, segments):
return ranges
+def resolve_model_path(project_root):
+ """The .nemo to load: $PARAKEET_MODEL_PATH (absolute, or relative to the env),
+ else the bundled v3 weights."""
+ chosen = os.environ.get("PARAKEET_MODEL_PATH", "").strip() or DEFAULT_MODEL
+ return chosen if os.path.isabs(chosen) else os.path.join(project_root, chosen)
+
+
+def find_speech_gaps(words, audio, sr, min_gap):
+ """Stretches of at least `min_gap` s with no word, where at least half the
+ 25 ms frames are within 12 dB of the recording's typical speech level: the
+ model went quiet while someone was talking. Parakeet sometimes does this for
+ tens of seconds deep inside a long chunk; the same audio transcribed on its
+ own is usually fine."""
+ hop, win = int(0.010 * sr), int(0.025 * sr)
+ frames = max(0, (len(audio) - win) // hop)
+ if frames == 0:
+ return []
+ # A strided view, reduced in blocks: bounded memory even for hours of audio.
+ view = np.lib.stride_tricks.sliding_window_view(audio, win)[::hop][:frames]
+ power = np.empty(frames)
+ for k in range(0, frames, 8192):
+ power[k:k + 8192] = np.mean(view[k:k + 8192].astype(np.float64) ** 2, axis=1)
+ level = 10 * np.log10(power + 1e-12)
+ speech = float(np.median(level[level >= np.median(level)]))
+ duration = len(audio) / sr
+ gaps = []
+ for a, b in zip([0.0] + [w["end"] for w in words], [w["start"] for w in words] + [duration]):
+ if b - a >= min_gap:
+ span = level[int(a * sr) // hop:int(b * sr) // hop]
+ if len(span) and np.mean(span >= speech - 12) >= 0.5:
+ gaps.append((a, b))
+ return gaps
+
+
+def retry_speech_gaps(words, segments, audio, sr, min_gap, transcribe):
+ """Re-transcribe each speech gap on its own, in pieces of at most
+ RETRY_PIECE_SECS, and splice in the words that land inside it.
+
+ `transcribe(start_sample, end_sample)` returns (text, words, segments) in
+ absolute time. Returns (words, segments, number of gaps retried); when words
+ were added, segments are rebuilt from the words with the model's own rule.
+ """
+ gaps = find_speech_gaps(words, audio, sr, min_gap)
+ if not gaps:
+ return words, segments, 0
+ found = []
+ for a, b in gaps:
+ t = a
+ while t < b:
+ e = min(b, t + RETRY_PIECE_SECS)
+ try:
+ _, piece, _ = transcribe(max(0, int((t - 0.5) * sr)), min(len(audio), int((e + 0.5) * sr)))
+ except Exception as err: # best effort: the first pass already succeeded
+ print(f"Warning: retry of {t:.1f}-{e:.1f}s failed ({err}); keeping the first pass there")
+ piece = []
+ piece = [w for w in piece if t <= w["start"] < e]
+ if _plausible(piece, e - t):
+ found.extend(piece)
+ t = e
+ if not found:
+ return words, segments, len(gaps)
+ merged = _without_retried_duplicates(sorted(words + found, key=lambda w: w["start"]),
+ {id(w) for w in found})
+ return merged, segments_from_words(merged), len(gaps)
+
+
+def _plausible(piece, seconds):
+ """False for what looks like a hallucination loop rather than speech: many words
+ that are mostly one repeated token, or more words per second than anyone says."""
+ if len(piece) >= 10 and len({_normalize(w["word"]) for w in piece}) < 0.3 * len(piece):
+ return False
+ return len(piece) <= RETRY_MAX_WORDS_PER_SEC * max(seconds, 1.0)
+
+
+def _without_retried_duplicates(merged, retried):
+ """The retry's padding re-hears the words either side of a gap, and a word
+ near a piece boundary is heard by both pieces. Drop a retried word that
+ repeats its neighbour within SAME_WORD_SECS; the first pass's copy wins."""
+ out = []
+ for w in merged:
+ if out and _normalize(out[-1]["word"]) == _normalize(w["word"]) \
+ and w["start"] - out[-1]["start"] <= SAME_WORD_SECS:
+ if id(w) in retried:
+ continue
+ if id(out[-1]) in retried:
+ out[-1] = w
+ continue
+ out.append(w)
+ return out
+
+
+def segments_from_words(words):
+ """Segments as Parakeet's decoding config makes them: a segment ends after a
+ word ending in '.', '?' or '!' (segment_seperators, no gap threshold)."""
+ segments, current = [], []
+ for w in words:
+ current.append(w)
+ if w["word"].endswith((".", "?", "!")):
+ segments.append(_segment_of(current))
+ current = []
+ if current:
+ segments.append(_segment_of(current))
+ return segments
+
+
+def _segment_of(words):
+ return {"segment": " ".join(w["word"] for w in words),
+ "start_offset": words[0]["start_offset"], "end_offset": words[-1]["end_offset"],
+ "start": words[0]["start"], "end": words[-1]["end"]}
+
+
+def transcribe_chunk(asr_model, audio, sr, start_time, path):
+ """Transcribe one chunk via a WAV file; word and segment times become absolute."""
+ sf.write(path, audio, sr)
+ try:
+ result = asr_model.transcribe([path], batch_size=1, timestamps=True)[0]
+ finally:
+ if os.path.exists(path):
+ os.remove(path)
+ words, segments = [], []
+ if getattr(result, "timestamp", None):
+ words = [dict(w, start=w["start"] + start_time, end=w["end"] + start_time)
+ for w in result.timestamp.get("word", [])]
+ segments = [dict(g, start=g["start"] + start_time, end=g["end"] + start_time)
+ for g in result.timestamp.get("segment", [])]
+ return result.text, words, segments
+
+
def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0):
"""Split audio file into chunks of at most chunk_duration_secs."""
audio, sr = librosa.load(audio_path, sr=None, mono=True)
@@ -173,7 +300,7 @@ def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, sear
'duration': len(chunk_audio) / sr
})
- return chunks, sr, [cut / sr for cut in cuts]
+ return chunks, sr, [cut / sr for cut in cuts], audio
def transcribe_buffered(
@@ -182,16 +309,13 @@ def transcribe_buffered(
chunk_duration_secs: float = 300, # 5 minutes default
overlap_secs: float = DEFAULT_OVERLAP_SECS,
pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS,
+ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS,
):
"""
Transcribe long audio by splitting into chunks and merging results.
"""
import nemo.collections.asr as nemo_asr
- # Determine model path
- model_filename = "parakeet-tdt-0.6b-v3.nemo"
- model_path = None
-
# Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv
virtual_env = os.environ.get("VIRTUAL_ENV")
if not virtual_env:
@@ -199,10 +323,11 @@ def transcribe_buffered(
sys.exit(1)
project_root = os.path.dirname(virtual_env)
- model_path = os.path.join(project_root, model_filename)
+ model_path = resolve_model_path(project_root)
+ model_name = os.path.splitext(os.path.basename(model_path))[0]
if not os.path.exists(model_path):
- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}")
+ print(f"Error during transcription: Can't find model file: {model_path}")
sys.exit(1)
print(f"Loading NVIDIA Parakeet model from: {model_path}")
@@ -226,58 +351,25 @@ def transcribe_buffered(
print(f"Splitting audio into chunks of at most {chunk_duration_secs}s "
f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...")
- chunks, sr, cut_times = split_audio_file(
+ chunks, sr, cut_times, audio = split_audio_file(
audio_path, chunk_duration_secs, overlap_secs, pause_search_secs
)
print(f"Created {len(chunks)} chunks")
slice_results = []
chunk_texts = []
+ retried = 0
+ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp
for i, chunk_info in enumerate(chunks):
print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...")
-
- # Save chunk to temporary file
- chunk_path = f"/tmp/chunk_{i}.wav"
- sf.write(chunk_path, chunk_info['audio'], sr)
-
- try:
- # Transcribe chunk
- output = asr_model.transcribe(
- [chunk_path],
- batch_size=1,
- timestamps=True,
- )
-
- result_data = output[0]
- chunk_text = result_data.text
- chunk_texts.append(chunk_text)
- chunk_words = []
- chunk_segments = []
-
- # Extract and adjust timestamps
- if hasattr(result_data, 'timestamp') and result_data.timestamp:
- # Adjust timestamps by chunk start time
- for word in result_data.timestamp.get("word", []):
- word_copy = dict(word)
- word_copy['start'] += chunk_info['start_time']
- word_copy['end'] += chunk_info['start_time']
- chunk_words.append(word_copy)
-
- for segment in result_data.timestamp.get("segment", []):
- seg_copy = dict(segment)
- seg_copy['start'] += chunk_info['start_time']
- seg_copy['end'] += chunk_info['start_time']
- chunk_segments.append(seg_copy)
-
- slice_results.append((chunk_words, chunk_segments))
-
- print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
-
- finally:
- # Clean up temp file
- if os.path.exists(chunk_path):
- os.remove(chunk_path)
+ chunk_text, chunk_words, chunk_segments = transcribe_chunk(
+ asr_model, chunk_info['audio'], sr, chunk_info['start_time'],
+ os.path.join(workdir, f"chunk_{i}.wav"),
+ )
+ chunk_texts.append(chunk_text)
+ slice_results.append((chunk_words, chunk_segments))
+ print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks]
all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans)
@@ -287,7 +379,20 @@ def transcribe_buffered(
print("Warning: a chunk has text but no word timestamps; joining chunk texts")
final_text = " ".join(chunk_texts)
else:
+ if retry_gap_secs > 0:
+ def transcribe_span(start_sample, end_sample):
+ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr,
+ start_sample / sr, os.path.join(workdir, "retry.wav"))
+ try:
+ all_words, all_segments, retried = retry_speech_gaps(
+ all_words, all_segments, audio, sr, retry_gap_secs, transcribe_span
+ )
+ except Exception as err: # best effort: never lose a finished transcription
+ print(f"Warning: gap retry failed ({err}); keeping the first pass")
+ if retried:
+ print(f"Re-transcribed {retried} stretch(es) of speech that got no words")
final_text = " ".join(w["word"] for w in all_words)
+ shutil.rmtree(workdir, ignore_errors=True)
print(f"Transcription complete: {len(final_text)} characters total")
output_data = {
@@ -296,13 +401,14 @@ def transcribe_buffered(
"word_timestamps": all_words,
"segment_timestamps": all_segments,
"audio_file": audio_path,
- "model": "parakeet-tdt-0.6b-v3",
+ "model": model_name,
"buffered": True,
"chunk_duration_secs": chunk_duration_secs,
"num_chunks": len(chunks),
"overlap_secs": overlap_secs,
"pause_search_secs": pause_search_secs,
"cut_times": cut_times,
+ "retried_gaps": retried,
}
if output_file:
@@ -328,6 +434,12 @@ def main():
help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len "
f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)"
)
+ parser.add_argument(
+ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS,
+ help=f"Re-transcribe on its own any stretch of at least this many seconds where "
+ f"the audio holds speech but the model produced no words "
+ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)"
+ )
parser.add_argument(
"--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS,
help=f"Seconds before each chunk limit searched for the quietest point to cut at, "
@@ -346,6 +458,7 @@ def main():
chunk_duration_secs=args.chunk_len,
overlap_secs=args.overlap,
pause_search_secs=args.pause_search,
+ retry_gap_secs=args.retry_gaps,
)
diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
index 6a35947..c3d7537 100644
--- a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
+++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
@@ -10,7 +10,10 @@ import numpy as np
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
-from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402
+from parakeet_transcribe_buffered import ( # noqa: E402
+ find_speech_gaps, plan_slices, resolve_model_path, retry_speech_gaps,
+ segments_from_words, stitch_slices,
+)
SR = 16000
@@ -327,3 +330,157 @@ def test_punctuation_alone_is_never_an_anchor():
right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)]
words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
assert words[0] is left[0] and texts(words) == ["-", "no"]
+
+
+# -- model and attention selection ----------------------------------------------
+
+
+def test_the_bundled_v3_model_is_the_default(monkeypatch):
+ monkeypatch.delenv("PARAKEET_MODEL_PATH", raising=False)
+ assert resolve_model_path("/env") == "/env/parakeet-tdt-0.6b-v3.nemo"
+
+
+def test_parakeet_model_path_overrides_the_model(monkeypatch):
+ monkeypatch.setenv("PARAKEET_MODEL_PATH", "/models/parakeet-tdt-0.6b-v2.nemo")
+ assert resolve_model_path("/env") == "/models/parakeet-tdt-0.6b-v2.nemo"
+ monkeypatch.setenv("PARAKEET_MODEL_PATH", "other.nemo")
+ assert resolve_model_path("/env") == "/env/other.nemo"
+
+
+# -- retrying stretches where the model went quiet --------------------------------
+
+
+def talk(seconds, level=0.1, seed=3):
+ return (level * np.random.default_rng(seed).standard_normal(int(seconds * SR))).astype(np.float32)
+
+
+def w_at(text, start, end):
+ return {"word": text, "start_offset": 0, "end_offset": 0, "start": start, "end": end}
+
+
+def test_a_long_wordless_stretch_over_speech_is_a_gap():
+ audio = talk(40)
+ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 38.0, 38.5)]
+ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (10.0, 20.0), (20.5, 38.0)]
+
+
+def test_a_wordless_stretch_over_silence_or_a_short_pause_is_not():
+ audio = talk(40)
+ audio[int(10 * SR):int(20 * SR)] = 0.0
+ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 22.0, 22.5),
+ w_at("e", 24.0, 24.5), w_at("f", 39.0, 39.5)]
+ # 10-20 s is silent; 20.5-22 and 22.5-24 are too short; 1.5-9.5 and 24.5-39 are talk
+ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (24.5, 39.0)]
+
+
+def test_speech_before_the_first_word_counts():
+ audio = talk(20)
+ assert find_speech_gaps([w_at("late", 15.0, 15.5)], audio, SR, min_gap=3.0) == [(0.0, 15.0), (15.5, 20.0)]
+
+
+def test_no_words_at_all_over_speech_is_one_gap():
+ assert find_speech_gaps([], talk(12), SR, min_gap=3.0) == [(0.0, 12.0)]
+
+
+def test_retried_words_are_spliced_into_the_gap_and_segments_rebuilt():
+ audio = talk(31)
+ words = [w_at("Hello", 1.0, 1.5), w_at("there.", 2.0, 2.5), w_at("Bye.", 30.0, 30.5)]
+ segments = [segment(words[:2]), segment(words[2:])]
+ calls = []
+
+ def transcribe(s0, s1): # stands in for Parakeet: finds speech the main pass missed
+ calls.append((s0 / SR, s1 / SR))
+ found = [w_at("We", 5.0, 5.2), w_at("missed", 6.0, 6.3), w_at("this.", 7.0, 7.4),
+ w_at("Again", 12.0, 12.4), w_at("outside", 40.5, 41.0)]
+ return "", [w for w in found if s0 / SR <= w["start"] < s1 / SR], []
+ new_words, new_segments, n = retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe)
+ assert [w["word"] for w in new_words] == ["Hello", "there.", "We", "missed", "this.", "Again", "Bye."]
+ assert n == 1 and calls == [(2.0, 30.5)] # one gap, 2.5-30.0 s, padded by 0.5 s
+ assert [g["segment"] for g in new_segments] == ["Hello there.", "We missed this.", "Again Bye."]
+ assert " ".join(g["segment"] for g in new_segments) == " ".join(w["word"] for w in new_words)
+
+
+def test_without_gaps_nothing_is_retried_and_output_is_untouched():
+ audio = talk(10)
+ words = [w_at("One", 0.5, 1.0), w_at("two", 2.5, 3.0), w_at("three.", 5.0, 5.5), w_at("four", 8.0, 8.5)]
+ segments = [segment(words[:3]), segment(words[3:])]
+
+ def transcribe(s0, s1):
+ raise AssertionError("must not be called")
+ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe) == (words, segments, 0)
+
+
+def test_long_gaps_are_retried_in_pieces():
+ audio = talk(191)
+ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)]
+ calls = []
+
+ def transcribe(s0, s1):
+ calls.append(round((s1 - s0) / SR, 1))
+ return "", [], []
+ retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
+ assert max(calls) <= 61.0 and len(calls) == 4 # 1.5-190 s in <= 60 s pieces (+0.5 s padding)
+
+
+def test_segments_from_words_split_after_sentence_punctuation():
+ ws = [w_at("Yes.", 0, 1), w_at("Is", 1, 2), w_at("it?", 2, 3), w_at("Go", 3, 4), w_at("now", 4, 5)]
+ assert [g["segment"] for g in segments_from_words(ws)] == ["Yes.", "Is it?", "Go now"]
+ assert segments_from_words([]) == []
+
+
+def test_a_retried_copy_of_the_word_after_the_gap_is_not_kept_twice():
+ # The retry piece is padded, so it re-hears the next kept word and may
+ # timestamp it just inside the gap.
+ audio = talk(31)
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
+
+ def transcribe(s0, s1):
+ return "", [w_at("We", 5.0, 5.2), w_at("left.", 6.0, 6.3), w_at("bye.", 29.7, 30.3)], []
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
+ assert [w["word"] for w in new_words] == ["Hello.", "We", "left.", "Bye."]
+ assert new_words[-1] is words[-1]
+
+
+def test_a_word_heard_by_two_retry_pieces_is_kept_once():
+ audio = talk(191)
+ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)]
+
+ def transcribe(s0, s1): # both pieces around 61.5 s hear "boundary"
+ t0, t1 = s0 / SR, s1 / SR
+ out = [w_at("boundary", 61.3, 61.8)] if t0 <= 61.3 < t1 else []
+ out += [w_at("boundary", 61.55, 61.9)] if t0 <= 61.55 < t1 and t0 > 60 else []
+ return "", out, []
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
+ assert [w["word"] for w in new_words] == ["a", "boundary", "z"]
+
+
+def test_a_failing_retry_keeps_the_first_pass():
+ audio = talk(31)
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
+ segments = [segment(words[:1]), segment(words[1:])]
+
+ def transcribe(s0, s1):
+ raise RuntimeError("CUDA out of memory")
+ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe)[:2] == (words, segments)
+
+
+def test_a_looping_retry_is_not_spliced_in():
+ audio = talk(31)
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
+
+ def transcribe(s0, s1): # a hallucination loop over music-like audio
+ return "", [w_at("la", 3.0 + k, 3.2 + k) for k in range(20)], []
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
+ assert [w["word"] for w in new_words] == ["Hello.", "Bye."]
+
+
+def test_gap_levels_line_up_with_time_at_sample_rates_other_than_16k():
+ # At 11025 Hz a 10 ms hop is 110 samples (9.977 ms); indexing frames by
+ # time/10ms would drift ~2 s by 900 s and read the silence before the gap.
+ sr = 11025
+ audio = (0.1 * np.random.default_rng(5).standard_normal(1000 * sr)).astype(np.float32)
+ audio[int(896.9 * sr):int(900.0 * sr)] = 0.0
+ words = [w_at("w", t, t + 0.5) for t in np.arange(0.0, 1000.0, 1.0) if not 900.0 <= t < 903.0]
+ words = [w for w in words if not (899.9 < w["start"] < 900.5)] + [w_at("w", 899.5, 900.0)]
+ words.sort(key=lambda w: w["start"])
+ assert (900.0, 903.0) in find_speech_gaps(words, audio, sr, 3.0)
--
2.39.5