docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips stretches of speech" finding, including other Parakeet weights. Investigation only; nothing deployed. Against ground truth (official SCOTUS transcript, Gutenberg #38916) the drops are real: production v3 loses 140 / 66 clean words per transcript on the two public files and ~50 on each private one (Whisper-referenced, Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause: the v2/v3 0.6B weights collapse deep inside long full-attention windows; the encoder output is degraded, the audio alone transcribes fine, and 1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy variants, max_symbols, beam), slice length, local attention, loudness, resampling and a noise floor do not fix it. Controls: A-vs-A, silence positive control (>=15 words 36/36), null control, bootstrap floor. Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds speech but no word came out (-80 to -90 % lost words on all four recordings, lower WER, no invented text, +10 MiB) and adds an explicit PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed. Reviewed at high effort, all findings fixed; built and tested as scriberr:local-blackwell-a353078-dropout2, not deployed. scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks both Parakeet scripts (seam-check --standard for the short-audio one).
This commit is contained in:
@@ -0,0 +1,389 @@
|
||||
# Parakeet dropout investigation (2026-09-30)
|
||||
|
||||
Prime's ask, via the coordinator, 2026-09-30: investigate the "Parakeet skips
|
||||
stretches of speech" finding from the slicer bench
|
||||
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md`), and include a different
|
||||
Parakeet weight. **Investigation only:** nothing here changed the live Scriberr
|
||||
container or its `.env`; deploying anything is Prime's call.
|
||||
|
||||
**Privacy:** two of the four recordings are Prime's. Their audio, transcripts and
|
||||
diarization stay in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700).
|
||||
This document carries metrics only.
|
||||
|
||||
## Verdict
|
||||
|
||||
- **The drops are real.** Against ground truth, today's Parakeet (v3, shipped
|
||||
slicer) loses whole stretches of clean speech: 140 words per transcript on a
|
||||
clean audiobook, 66 on a court argument, and about 50 on each of Prime's two
|
||||
recordings (Whisper-referenced, Canary-confirmed). The old "reference" was
|
||||
mostly right; on SCOTUS it drops speech too.
|
||||
- **Root cause:** the Granary-era 0.6B weights (v3 and v2) sometimes stop
|
||||
producing words for tens of seconds deep inside a long full-attention window.
|
||||
The encoder output there is degraded (a fresh decoder recovers only 22 to 51 %),
|
||||
the same audio transcribed alone is fine, and older Parakeets with the **same
|
||||
TDT decoder** do not do it. It is chaotic in cut placement and not explained by
|
||||
loudness, SNR, speaking rate or language. Decoding settings, beam search,
|
||||
shorter slices, loudness normalisation and resampling do not fix it.
|
||||
- **Fix that works:** re-transcribe any ≥ 3 s stretch where the audio holds
|
||||
speech but the model produced no words. That cuts the loss 80 to 90 % on all
|
||||
four recordings (to 19, 13, 7 and 5 words per transcript) with no invented
|
||||
text, at +10 MiB and no measurable time. It is patch 0002 (proposed, built and
|
||||
tested, **not deployed**), which also adds an explicit `PARAKEET_MODEL_PATH`.
|
||||
- **Weights:** v2 loses nothing on Prime's recordings or SCOTUS but collapses on
|
||||
read speech; `parakeet-unified-en-0.6b` (newest) is the best-balanced
|
||||
punctuating Parakeet but needs NeMo 3.0.0 and a licence change. The 1.1B and
|
||||
CTC Parakeets never collapse but produce no punctuation.
|
||||
- **Separate, smaller issue:** overlapping speakers (SCOTUS interjections) are
|
||||
lost by every single-stream model. That is not this bug.
|
||||
|
||||
## 1. Is the instrument right? Ground truth, not a model reference
|
||||
|
||||
The first finding measured losses against the whole-file local-attention
|
||||
transcript. That reference could itself be inserting text, so the public
|
||||
recordings were re-scored against real ground truth.
|
||||
|
||||
| file | ground truth | provenance |
|
||||
|---|---|---|
|
||||
| scotus (first 30 min of No. 22-451) | the official argument transcript, `supremecourt.gov/oral_arguments/argument_transcripts/2023/22-451_114p.pdf` (sha256 `4feb7786…78a2`) | cover pages, page and line numbers, running headers, time stamps, argument headings and speaker labels removed; cut where the audio ends by alignment (5,627 of 14,764 transcript words) |
|
||||
| wilde (LibriVox section 1) | Project Gutenberg #38916, *The Trial of Oscar Wilde, from the Shorthand Reports* (the LibriVox page's own "online text" link), sha256 `2271f271…b3a9` | section 1 is the book's Preface, read by the narrator alone ("It is wrong for us…" to "…came too late."), plus LibriVox's standard spoken intro and outro |
|
||||
|
||||
**Normalisation:** Whisper's English normaliser (transformers 4.53.3, the version
|
||||
in Scriberr's env, no spelling map), applied word by word to both sides so every
|
||||
hypothesis token keeps its timestamp. Fillers are dropped, numbers become digits
|
||||
and contractions are expanded.
|
||||
|
||||
**Alignment and the dropout detector:** difflib opcodes, then exact Levenshtein
|
||||
inside each mismatch. A **dropout** is a stretch between solid matches (≥ 3
|
||||
tokens) holding ≥ 10 reference words where the hypothesis emitted fewer than half
|
||||
as many. An **insertion run** is the mirror image. Matching islands shorter than
|
||||
3 tokens count as part of the stretch, so one spurious "the" cannot split a
|
||||
skipped paragraph in two.
|
||||
|
||||
**Timed ground truth:** every ground-truth token gets the median start time of
|
||||
the transcripts that matched it (14 transcripts); the 143 (scotus) and 61
|
||||
(wilde) tokens that no transcript matched are interpolated. There are no time
|
||||
reversals over 1 s.
|
||||
|
||||
### Result: the drops are real, measured against ground truth
|
||||
|
||||
The same outputs the first finding used (the shipped slicer and upstream's fixed
|
||||
cutter, 3 placements each), now scored against ground truth:
|
||||
|
||||
| file | system | WER | dropouts | words dropped |
|
||||
|---|---|---|---|---|
|
||||
| wilde | whole-file local attention (the old stand-in reference) | 2.12 % | 0 | 0 |
|
||||
| wilde | upstream fixed 120 s cutter | 2.57 % | 0 | 0 |
|
||||
| wilde | shipped slicer, at 120 / 110 / 100 s | 2.15 / **13.02** / 2.40 % | 0 / 3 / 0 | 0 / **394** / 0 |
|
||||
| scotus | whole-file local attention | 6.01 % | 6 | 146 |
|
||||
| scotus | upstream fixed 120 s cutter | 4.54 % | 4 | 46 |
|
||||
| scotus | shipped slicer, at 120 / 110 / 100 s | 5.42 / 5.89 / 5.30 % | 4 / 5 / 4 | 94 / 137 / 111 |
|
||||
|
||||
On Wilde, a clean single-narrator audiobook, the stand-in reference was right
|
||||
and the chunked runs genuinely lost whole paragraphs; which paragraphs depends
|
||||
only on where the cuts fall. On SCOTUS the stand-in reference **also** drops
|
||||
speech (146 words), so the first finding's "reference insertions" there were
|
||||
really reference deletions.
|
||||
|
||||
### Adjudicating Prime's recordings without ground truth
|
||||
|
||||
Two independent models give a second opinion on every disputed stretch (≥ 10
|
||||
words one system has and another lacks): **Whisper large-v3** (openai, pinned
|
||||
`06f233fe`, transformers sequential long-form decoding; its word times are
|
||||
spread evenly within segments, so it gets a ±3 s window) and **Canary-1b-v2**
|
||||
(the copy inside Scriberr's env, sha-matched to HF `d4557063`; cut in pauses,
|
||||
no overlap, ±1 s window). "Speech is real" needs both to have ≥ 50 % of the
|
||||
disputed words; "no speech" needs both under 20 %; anything else is ambiguous.
|
||||
|
||||
**The adjudicator is calibrated on the public files first**, where ground truth
|
||||
gives the right answer. It decided 129 of 140 disputed stretches and **got all
|
||||
129 right**; it abstained on 11, mostly crosstalk that the official transcript
|
||||
renders differently. Whisper large-v3 on its own scores 1.96 % (wilde) and 3.93 %
|
||||
(scotus) WER against ground truth with **zero** dropouts, so for scoring Prime's
|
||||
files it serves as the reference, and a dropout counts only if Canary
|
||||
independently has the words. On the public files that metric reproduces the
|
||||
ground-truth clean-speech dropout counts within a few percent (1,060 vs 1,116;
|
||||
287 vs 310; 240 vs 246; 5,450 vs 5,513; 515 vs 526 words).
|
||||
|
||||
### Controls and floor
|
||||
|
||||
- **A-vs-A.** Production v3 (full attention) is byte-identical across three
|
||||
separate processes at every placement tested. **Local attention is not:** at
|
||||
the same SCOTUS placement three runs dropped 68, 169 and 100 words (WER 4.73 to
|
||||
6.32 %). Its numbers below carry that run-to-run noise.
|
||||
- **Positive control.** Copies of both public files with 30, 15, 5 and 4 s of
|
||||
audio replaced by digital silence (the ground truth still holds those words):
|
||||
the 30, 15 and 5 s stretches (15 to 73 words) were reported in all 36 runs; the
|
||||
4 s stretch (11 to 14 words) in 11 of 12.
|
||||
- **Null control.** In those same runs, chunks that touch no silenced stretch get
|
||||
byte-identical input; with full attention their words were identical and their
|
||||
dropouts equal in all 6 runs, so the instrument manufactures nothing. (With
|
||||
local attention they differ; that is the non-determinism above.)
|
||||
- **A "should-not-matter" perturbation** (−0.5 dB gain) left SCOTUS identical and
|
||||
Wilde within its placement spread (1,027 vs 1,116 words). Parakeet normalises
|
||||
each mel feature per chunk, so a constant gain is nearly a no-op by design.
|
||||
- **Sensitivity floor:** a dropout of ≥ 15 words is caught every time (36/36);
|
||||
10 to 14 words, 11/12; shorter losses are not counted as dropouts at all (they
|
||||
still count in WER). Because the failure is chaotic in cut placement, every
|
||||
configuration is run at 8 placements (4 for the private files and the 1.1B
|
||||
diagnostics); totals over 8 placements that differ by less than about a
|
||||
quarter are not a difference.
|
||||
|
||||
## 2. What the real drops look like
|
||||
|
||||
Production v3 (the shipped slicer: 120 s chunks, 4 s overlap), 8 cut placements
|
||||
per public file, each dropout compared with 20 random windows of the same length
|
||||
from the same file (clean speech; SCOTUS crosstalk split out below):
|
||||
|
||||
| | Wilde (one narrator) | SCOTUS (argument) |
|
||||
|---|---|---|
|
||||
| dropouts / words | 17 / 1,116 | 29 / 730 |
|
||||
| start, seconds into its chunk | median 47, **never before 16** | median 53, never before 13 |
|
||||
| runs to the end of its chunk | 35 % | 10 % |
|
||||
| level vs file speech level (drops / controls) | −1.0 / −1.35 dB | −0.1 / −2.0 dB |
|
||||
| local SNR (drops / controls) | 42 / 41 dB | 29 / 27 dB |
|
||||
| speaking rate (drops / controls) | 2.4 / 2.5 words/s | 4.0 / 3.2 words/s |
|
||||
| pause just before (drops / controls) | 0.62 / 0.37 s | 0.02 / 0.02 s |
|
||||
| overlapped speech, Sortformer (drops / controls) | 0 / 0 | **16 % / 0 %** of the window |
|
||||
| non-English words within ±10 s | 0 | 0 |
|
||||
|
||||
Two different things are being counted:
|
||||
|
||||
- **A. Long-context collapse.** Clean speech lost in stretches of tens of seconds
|
||||
(up to 60 s), beginning at a sentence boundary well into a long chunk, often
|
||||
running to the chunk's end. It is **not explained by the audio**: level, SNR
|
||||
and rate match the controls, there is one speaker, and nothing is non-English.
|
||||
Which stretches go is chaotic: it moves with the cut placement, and on Wilde
|
||||
no word is lost by more than 75 % of placements.
|
||||
- **B. Crosstalk.** On SCOTUS, a justice's interjection over counsel ("Well,
|
||||
wait a minute…") is lost at the same few spots in almost every run. A
|
||||
single-stream model transcribes one voice; the official transcript records
|
||||
both. That is a limit of single-stream ASR and of the transcript convention,
|
||||
not the failure this investigation is about, so it is **separated out**
|
||||
(diarized overlap ≥ 10 % of the stretch) in every comparison below.
|
||||
|
||||
Every comparison below counts **clean-speech dropped words per transcript**, mean
|
||||
with a 95 % bootstrap interval over cut placements.
|
||||
|
||||
**Language ID is not the cause.** v3's only non-ASCII output is legitimate French
|
||||
names in the text (Mallarmé, Comédie), never near a drop, and English-only v2
|
||||
collapses too. Prime's recordings are English (function-word share 0.35 to 0.39,
|
||||
the same as the public English files at 0.39 to 0.44; no non-ASCII letters).
|
||||
|
||||
## 3. Mechanism
|
||||
|
||||
- **The encoder output is degraded, not just the decoder.** On the three chunks
|
||||
that went fully quiet (0 words after the drop began), a fresh decoder state
|
||||
started on the same full-attention encoder output recovered only 22 to 51 % of
|
||||
the ground-truth words there. The same audio encoded on its own recovered 72 to
|
||||
102 %, and local attention over the full chunk 97 to 102 % (a fourth, partial
|
||||
drop: 98 %, 57 %, 98 %; a control chunk: all ≈ production). That is n = 4 hand-picked
|
||||
chunks, so it points the way; the unbiased tests below carry the weight.
|
||||
- **It belongs to the weights, not to TDT decoding.** On the same slices,
|
||||
`parakeet-tdt-1.1b` (TDT), `parakeet-rnnt-1.1b`, `parakeet-ctc-1.1b`,
|
||||
`parakeet-ctc-0.6b` and both heads (TDT and CTC) of `parakeet-tdt_ctc-1.1b` all
|
||||
lose **0** words on Wilde. v3 and v2, the Granary-era 0.6B TDT models, collapse.
|
||||
The new `parakeet-unified-en-0.6b` (RNN-T) collapses once in 8.
|
||||
- **It is a knife edge.** Deterministic for a given input (v3 at full attention is
|
||||
byte-identical across processes), but the float-level non-determinism of the
|
||||
local-attention kernel is enough to flip whether a stretch is transcribed
|
||||
(68 / 169 / 100 words at one placement).
|
||||
- **Recording style matters, differently per weight.** Wilde and p2 are gated
|
||||
recordings (pauses near −64 and −74 dBFS); SCOTUS and p1 never go quiet (room
|
||||
tone at −40 and −36 dBFS). v2 collapses badly on Wilde but never on the other
|
||||
three; a −50 dBFS noise floor more than halves v3's and v2's Wilde losses but
|
||||
makes SCOTUS worse. Not a clean lever.
|
||||
|
||||
## 4. Hypotheses tested (unbiased: full files at 8 placements; clean-speech words per transcript)
|
||||
|
||||
| hypothesis | test | Wilde | SCOTUS | Prime p1 / p2 | verdict |
|
||||
|---|---|---|---|---|---|
|
||||
| — | **production v3** | 140 [48–240] | 66 [41–88] | 50 [40–63] / 51 [25–82] | baseline |
|
||||
| (a) decoding | CUDA graphs on | identical output | identical | | no effect |
|
||||
| (a) | greedy (per-frame) instead of greedy_batch | identical output | identical | | no effect |
|
||||
| (a) | max_symbols 20 | identical at its 3 placements | identical | | no effect |
|
||||
| (a) | TDT beam 4 (text only; NeMo 2.5.3 cannot give timestamps with TDT beam), all dropouts, vs greedy under the same slicing | 194 vs 136 | 488 vs 57 | | **worse**, 3–13× slower, unusable in Scriberr |
|
||||
| (b) context | 60 s slices | 88 [41–138] | 61 [41–85] | 31 / 31 | within the floor |
|
||||
| (b) | 30 s slices | 165 [95–236] | 77 [49–105] | 32 / 13 | worse on Wilde |
|
||||
| (b) | local attention ±64 frames in 120 s slices | 34 [21–47] | 97 [72–121] | | trades one file for the other |
|
||||
| (b) | local attention ±128 | 39 [24–51] | 128 [92–168] | 0 / 3 | trades; **non-deterministic** |
|
||||
| (b) | local attention ±256 | 31 [15–47] | 80 [43–120] | 60 / 64 | no help on Prime's files |
|
||||
| (c) preprocessing | −0.5 dB gain (null) | 128 | identical | 54 / 30 | no effect (the null) |
|
||||
| (c) | loudness normalisation to −20 dBFS, peak-safe | **498** [382–613] | 66 | | **worse** on quiet audio |
|
||||
| (c) | soxr resampling instead of ffmpeg (all dropouts; production on the same measure: 140 / 91) | 206 | 92 | | within the floor |
|
||||
| (c) | −50 dBFS noise floor | 59 [30–87] | 87 [74–100] | | helps Wilde, hurts SCOTUS |
|
||||
| (d) weights | see § 5 | | | | the lever |
|
||||
| new | **re-transcribe speech gaps ≥ 3 s** (patch 0002) | **19 [0–38]** | **13 [3–25]** | **7 [0–22] / 5 [0–15]** | **fixes most of it** |
|
||||
|
||||
WER against ground truth moves the same way: production v3 4.83 % / 5.33 % median
|
||||
(Wilde / SCOTUS) against 2.40 % / 4.47 % with the gap retry.
|
||||
|
||||
## 5. Candidates
|
||||
|
||||
Clean-speech **dropped words per transcript** (mean, 95 % bootstrap interval over
|
||||
placements; public: against ground truth, 8 placements; Prime's: against Whisper,
|
||||
confirmed by Canary, 4 placements) and **median WER** (public: against ground
|
||||
truth; Prime's: against Whisper). **Memory**: per-process peak on the 35-min file
|
||||
(p1), nvidia-smi every 0.2 s, only the run's own container PIDs, n = 3, zero
|
||||
spread in every case. **Speed**: median wall time per job for the 35.3-min file
|
||||
including model load, n = 3 (spread ±5 s). "CLI" rows ran the actual patched
|
||||
Scriberr scripts under Scriberr's invocation; "lab" rows ran the lab harness
|
||||
because Scriberr's env cannot run them as-is.
|
||||
|
||||
| candidate | Wilde words / WER | SCOTUS words / WER | p1 words / WER | p2 words / WER | peak MiB | job time | old GPU 1 budget (5,496) | GPU 3 (+Blender 270 MiB) | punctuation |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| **v3, production** (0001) | 140 [48–240] / 4.83 % | 66 [41–88] / 5.33 % | 50 / 2.96 % | 51 / 2.93 % | 5,496 (CLI) | 53 s | fits (0 spare) | fits | yes |
|
||||
| **v3 + gap retry** (0001+0002) | **19** [0–38] / **2.40 %** | **13** [3–25] / **4.47 %** | **7** / 2.52 % | **5** / 2.38 % | 5,506 (CLI) | 48 s | **10 MiB over** | fits | yes |
|
||||
| v2 | 689 [490–971] / 17.6 % | **0** / 4.55 % | **0** / 2.22 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
|
||||
| v2 + gap retry | 151 [116–192] / 5.84 % | **0** / 4.58 % | **0** / 2.25 % | **0** / 2.09 % | 5,438 (CLI) | 49 s | fits | fits | yes |
|
||||
| parakeet-unified-en-0.6b (NeMo 3.0.0) | 31 [0–92] / **2.21 %** | 22 [14–32] / 4.92 % | **0** / **2.19 %** | **0** / **1.83 %** | 5,438 (lab) | 34 s | fits | fits | yes |
|
||||
| parakeet-tdt-1.1b | **0** / 2.31 % | 3 / 6.78 % | **0** / 3.30 % | **0** / 2.65 % | 8,884 (lab) | 46 s | no | fits | **no** |
|
||||
| parakeet-ctc-0.6b | **0** / 2.46 % | 3 / 6.41 % | – | – | 5,360 (lab) | 31 s | fits | fits | **no** |
|
||||
| canary-1b-v2 (30 s pause cuts) | 11 [4–21] / 9.64 % ⚠ | 5 / 5.25 % | **0** / 3.88 % | **0** / 2.84 % | 10,504 (lab) | 112 s | no | fits | yes |
|
||||
| *Whisper large-v3 (reference, 1 run)* | *0 / 1.96 %* | *0 / 3.93 %* | – | – | – | – | – | – | *yes* |
|
||||
|
||||
⚠ Canary's Wilde WER is its **hallucination loops**: 6 insertion runs across 8
|
||||
placements, one of 417 words. It barely drops speech but invents it.
|
||||
|
||||
Other 1.1B diagnostics (no punctuation, so not candidates): `parakeet-rnnt-1.1b`
|
||||
0 / 7, `parakeet-ctc-1.1b` 0 / 3, `parakeet-tdt_ctc-1.1b` TDT head 0 / 6 and CTC
|
||||
head 0 / 0 (Wilde / SCOTUS words per transcript, 4 placements).
|
||||
|
||||
Provenance of every weight (all pulled revision-pinned, sha256 equal to the HF
|
||||
LFS oid, in `/tank/aimodels/huggingface/hub/`):
|
||||
|
||||
| repo | revision | licence (read at the raw card) | file sha256 |
|
||||
|---|---|---|---|
|
||||
| nvidia/parakeet-tdt-0.6b-v3 (in Scriberr's env) | `541d1f99` | CC-BY-4.0 | `3cbdc858…` |
|
||||
| nvidia/parakeet-tdt-0.6b-v2 | `ae9ad070` | CC-BY-4.0 | `d99e3995…` |
|
||||
| nvidia/parakeet-unified-en-0.6b (2026-04-07, newest Parakeet) | `fe53cd88` | **NVIDIA Open Model License** | `ec23ed91…` |
|
||||
| nvidia/parakeet-tdt-1.1b | `53276c64` | CC-BY-4.0 | `9c563d52…` |
|
||||
| nvidia/parakeet-rnnt-1.1b | `2acc4c61` | CC-BY-4.0 | `535896f0…` |
|
||||
| nvidia/parakeet-ctc-1.1b | `20e63a0f` | CC-BY-4.0 | `8e91253d…` |
|
||||
| nvidia/parakeet-tdt_ctc-1.1b | `675e7868` | CC-BY-4.0 | `4e7ccfdd…` |
|
||||
| nvidia/parakeet-ctc-0.6b | `ad09ba1c` | CC-BY-4.0 | `bc01f3f8…` |
|
||||
| nvidia/canary-1b-v2 (in Scriberr's env) | `d4557063` | CC-BY-4.0 | `ae5ef1bf…` |
|
||||
| openai/whisper-large-v3 (adjudicator only) | `06f233fe` | Apache-2.0 | `a8e94b85…` |
|
||||
|
||||
All repo ids were verified with an authenticated HF API call before any pull (no
|
||||
phantoms). `parakeet-unified-en-0.6b` needs **NeMo 3.0.0**: its card says 2.7.3,
|
||||
but released 2.7.3 lacks its encoder argument (`att_chunk_context_size`), and its
|
||||
`.nemo` ships without a `validation_ds` config that `transcribe()` reads (a
|
||||
two-line shim). It ran in a throwaway env, not Scriberr's (NeMo 2.5.3).
|
||||
|
||||
## 6. The fix that works on any weight: re-transcribe speech gaps
|
||||
|
||||
The mechanism says the lost audio is fine on its own; only its long-window
|
||||
context breaks. So after stitching, the buffered script looks for stretches of
|
||||
**≥ 3 s with no word where at least half the 25 ms frames sit within 12 dB of the
|
||||
recording's typical speech level**, re-transcribes each on its own (pieces of at
|
||||
most 60 s, 0.5 s of padding) and splices in the words that land inside it.
|
||||
Output that triggers no retry is byte-identical to today's.
|
||||
|
||||
- **Effect:** v3's clean-speech losses fall 80 to 90 % on all four recordings
|
||||
(140 → 19, 66 → 13, 50 → 7, 51 → 5 words per transcript) and WER falls with
|
||||
them (Wilde 4.83 → 2.40 %, SCOTUS 5.33 → 4.47 %). It adds **no insertion runs**,
|
||||
so it is not making text up.
|
||||
- **Trigger:** 3 s beat 6 s on Prime's files (p1 7 vs 27 words per transcript)
|
||||
and matched it on the public ones; 6 retries per 35-minute transcript were
|
||||
typical.
|
||||
- **Cost:** peak 5,506 MiB vs 5,496 (the retry's shorter chunk shifts the
|
||||
allocator by 10 MiB); job time unchanged within the ±5 s run-to-run spread.
|
||||
- **Residual:** stretches where the model still emits a few words (no clean gap),
|
||||
and retries that come back short. It does not rescue v2 on the audiobook
|
||||
(689 → 151).
|
||||
- **Validity:** the production implementation (after code review) reproduces
|
||||
the measured lab version at every shared placement: identical dropped words
|
||||
on all 16 file-placement pairs tried and WER equal or up to 0.02 points lower
|
||||
(it now removes the odd retried word that repeated its neighbour). It is
|
||||
deterministic across processes, and the JSON seam checks pass for both
|
||||
scripts with the rebuilt segments.
|
||||
- **Hardening from review:** a failing retry piece is skipped and a failing retry
|
||||
keeps the first pass (never loses a finished transcript); a piece that looks
|
||||
like a hallucination loop (mostly one repeated token, or > 7 words/s) is not
|
||||
spliced in; retried copies of the words at a gap's edge are dropped; audio goes
|
||||
through a per-run temp directory; frame levels use bounded memory (the first
|
||||
version needed ~1.7 GB of RAM for 35 min); frame indexing is correct at any
|
||||
sample rate. **Residual risk:** the gap detector is an energy test, so a loud
|
||||
non-speech stretch (a music bed) is retried; the loop guard is the only check
|
||||
on what comes back. None of the four recordings has music.
|
||||
|
||||
## 7. Recommendation
|
||||
|
||||
| claim | strength | basis | reversibility |
|
||||
|---|---|---|---|
|
||||
| The drops are real speech lost by Parakeet, not a reference artefact | **insist** | measured against ground truth on 2 files, 129/129 adjudications correct on the calibration | n/a |
|
||||
| Ship patch 0002 (gap retry on by default, v3 weights) | **strongly recommend** | measured: 80–90 % fewer lost words on all 4 recordings, lower WER, no new insertions, n = 4–8 placements each | reversible (image rollback) |
|
||||
| Keep v3 rather than switch to v2 | **lean** | measured: v2 is perfect on Prime's two files and SCOTUS but loses 151 words per transcript on read speech even with the retry, and is English-only; the right answer depends on what Prime transcribes | reversible (one env var) |
|
||||
| Evaluate `parakeet-unified-en-0.6b` as the next weight, as its own project | **lean** | measured: the best-balanced punctuating Parakeet (0 / 0 on Prime's files, lowest WER there); but it needs NeMo 3.0.0 in Scriberr's env (Canary and Sortformer too), a packaging shim, and a licence change from CC-BY-4.0 to the NVIDIA Open Model License | costly to reverse (env rebuild) |
|
||||
| Do not use local attention, beam search, shorter slices, loudness normalisation or a noise floor | **recommend against** | measured: each worse or mixed, local attention also non-deterministic | reversible |
|
||||
| Do not switch to Canary-1b-v2 | **recommend against** | measured: hallucination loops (up to 417 invented words) and twice the memory | reversible |
|
||||
| Do not use the 1.1B or CTC Parakeets | **recommend against** | read at source: no punctuation or casing, which Scriberr's segmentation needs | reversible |
|
||||
|
||||
## 8. The integration patch
|
||||
|
||||
`stacks/scriberr/patches/proposed/0002-parakeet-model-path-and-gap-retry.patch`
|
||||
(on top of 0001, upstream `a353078`; sha256 `e3098bf2…7e62`). It is in
|
||||
`proposed/`, so `scripts/scriberr-rebuild` does **not** apply it until it is moved
|
||||
up a directory. What it does:
|
||||
|
||||
1. **Gap retry** (§ 6) in both scripts, `--retry-gaps SECS` (default 3, 0 off).
|
||||
Go never passes the flag, so the default is the behaviour; the JSON gains
|
||||
`retried_gaps`.
|
||||
2. **`PARAKEET_MODEL_PATH`**: an absolute path (or one relative to the env) to
|
||||
the `.nemo` both scripts load; default unchanged. The JSON `model` field now
|
||||
reports the file actually loaded, and the Go adapter records it as
|
||||
`ModelUsed` instead of the hardcoded "parakeet-tdt-0.6b-v3". **No model is
|
||||
ever swapped in under v3's filename.**
|
||||
3. Tests: 18 new unit tests (39 in total), and upstream's own standard and
|
||||
buffered tests pass in the built image. Reviewed at high effort; all ten
|
||||
findings fixed (§ 6).
|
||||
|
||||
**Validated as a build** (`scripts/scriberr-rebuild --suffix dropout2 --budget
|
||||
5600 --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`): all
|
||||
stages PASS, including the embed of both scripts, 39 unit tests, the JSON seam
|
||||
for the long- and the short-audio script, and memory (5,506 MiB). The image
|
||||
`scriberr:local-blackwell-a353078-dropout2` exists on fv-ml1 and is **not
|
||||
deployed**.
|
||||
|
||||
**To ship it (Prime's call):** either deploy the already-built
|
||||
`scriberr:local-blackwell-a353078-dropout2` per `patches/README.md` § Deploy, or
|
||||
move the patch up into `stacks/scriberr/patches/` first (so it becomes part of
|
||||
the carried set) and rebuild under a new suffix with `--budget 5600` (see the
|
||||
note on the budget).
|
||||
|
||||
**To also switch weights (only if Prime chooses v2 or, later, another weight):**
|
||||
in `stacks/scriberr/compose.yaml`, mount the shared model cache read-only and
|
||||
name the pinned file. The revision is visible in the path and recorded in every
|
||||
transcript's metadata:
|
||||
|
||||
```yaml
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/models:ro
|
||||
environment:
|
||||
- PARAKEET_MODEL_PATH=/models/hub/models--nvidia--parakeet-tdt-0.6b-v2/snapshots/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/parakeet-tdt-0.6b-v2.nemo
|
||||
```
|
||||
|
||||
**How it survives upgrades:** `scriberr-rebuild` re-applies 0001 and 0002 to any
|
||||
pinned upstream sha and stops on a conflict; the model choice lives in our
|
||||
compose file, not in the image or the env directory, so an upgrade cannot
|
||||
silently change it. If upstream ever grows its own model selection, 0002 is
|
||||
dropped in favour of it.
|
||||
|
||||
**Budget note:** the rebuild script's default memory budget is still GPU 1's old
|
||||
5,496 MiB. With Scriberr on GPU 3 that number no longer protects anything, but
|
||||
the patched build peaks at 5,506, so either pass `--budget` or retire the old
|
||||
default (a one-line change; left for Prime or the coordinator).
|
||||
|
||||
## 9. Harness, data and reproduction
|
||||
|
||||
All code is in fv-ml1 `/tank/spikes/scriberr-slicer/code/dropout/`: `lab.py`
|
||||
(the production pipeline with every knob), `gtscore.py`, `boot.py`,
|
||||
`adjudicate.py`, `adj_score.py`, `characterize.py`, `probe.py`, `fit.sh`,
|
||||
`fitlab.sh`, plus the run scripts and configs. Metrics are in `…/metrics/`. The
|
||||
public audio, ground truth and transcripts are in `…/public/` and `…/gt/`;
|
||||
Prime's are in `…/private/` (mode 700). Throwaway envs: `…/envs/nemo300`
|
||||
(NeMo 3.0.0, for the unified model).
|
||||
|
||||
Every run was a transient `--rm` container on GPU 3, with Scriberr's env and the
|
||||
model cache mounted read-only. GPU 1 and the live Scriberr container were not
|
||||
touched.
|
||||
@@ -181,7 +181,7 @@ _As of 2026-09-30 ~0120 PT._
|
||||
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
||||
- **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`.
|
||||
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout1`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
|
||||
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
|
||||
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
|
||||
|
||||
@@ -31,6 +31,12 @@
|
||||
# Usage:
|
||||
# scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N]
|
||||
# [--budget MIB] [--memory-audio PATH] [--reuse-image]
|
||||
# [--patches DIR]
|
||||
#
|
||||
# --patches DIR[:DIR...] applies every *.patch in those directories, sorted by file
|
||||
# name, instead of the carried set in stacks/scriberr/patches. To test-build a
|
||||
# proposed patch on top of the carried ones, under its own --suffix:
|
||||
# --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed
|
||||
#
|
||||
# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496,
|
||||
# --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB
|
||||
@@ -44,10 +50,11 @@ HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54}
|
||||
ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY
|
||||
TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1
|
||||
SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||
STD_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
|
||||
TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||
SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav
|
||||
|
||||
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0
|
||||
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 PATCH_DIR_ARG=""
|
||||
MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav
|
||||
while [ $# -gt 0 ]; do
|
||||
case $1 in
|
||||
@@ -57,6 +64,7 @@ while [ $# -gt 0 ]; do
|
||||
--budget) BUDGET=$2; shift 2 ;;
|
||||
--memory-audio) MEM_AUDIO=$2; shift 2 ;;
|
||||
--reuse-image) REUSE_IMAGE=1; shift ;;
|
||||
--patches) PATCH_DIR_ARG=$2; shift 2 ;;
|
||||
-h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;;
|
||||
*) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;;
|
||||
esac
|
||||
@@ -66,8 +74,10 @@ done
|
||||
[[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; }
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches
|
||||
mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort)
|
||||
PATCH_DIR=${PATCH_DIR_ARG:-$REPO_ROOT/stacks/scriberr/patches}
|
||||
IFS=: read -r -a PATCH_DIRS <<<"$PATCH_DIR"
|
||||
mapfile -t PATCHES < <(for d in "${PATCH_DIRS[@]}"; do find "$d" -maxdepth 1 -name '*.patch'; done \
|
||||
| awk -F/ '{print $NF "\t" $0}' | sort | cut -f2-)
|
||||
[ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; }
|
||||
# The only paths a reused build dir may differ from upstream in.
|
||||
mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u)
|
||||
@@ -192,14 +202,14 @@ else
|
||||
fi
|
||||
|
||||
# ── embed ──────────────────────────────────────────────────────────────────
|
||||
if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF'
|
||||
if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" "$STD_REL" 2>&1 <<'EOF'
|
||||
docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c "
|
||||
import sys
|
||||
script = open('/src/$3', 'rb').read()
|
||||
sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)"
|
||||
binary = open('/app/scriberr', 'rb').read()
|
||||
sys.exit(0 if all(open('/src/' + f, 'rb').read() in binary for f in ('$3', '$4')) else 1)"
|
||||
EOF
|
||||
); then
|
||||
pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr"
|
||||
pass embed "both Parakeet scripts are byte-identical inside /app/scriberr"
|
||||
else
|
||||
fail embed "the binary does not embed the patched script ${out:+($out)}"
|
||||
fi
|
||||
@@ -240,6 +250,12 @@ gpu_idle() { # "room", not "idle": Scriberr itself may be running a job on this
|
||||
gpu_idle seam
|
||||
if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")"
|
||||
else fail seam "$out"; fi
|
||||
# The short-audio script (files under PARAKEET_CHUNK_THRESHOLD_SECS) has its own entry point.
|
||||
if out=$("${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \
|
||||
-v $BUILD_DIR/tests/data:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$STD_REL /audio/$(basename "$SEAM_AUDIO_REL") \
|
||||
--output /tmp/out.json --context-left 256 --context-right 256 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \
|
||||
python3 /tools/seam-check.py /tmp/out.json --standard'" 2>&1); then pass seam-short "$(tail -1 <<<"$out")"
|
||||
else fail seam-short "$out"; fi
|
||||
|
||||
# ── memory ─────────────────────────────────────────────────────────────────
|
||||
"${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1"
|
||||
|
||||
@@ -7,7 +7,8 @@ job. This checks the shape Go reads plus the stitching invariants the slicer
|
||||
patch promises. Stdlib only, so it runs under any python3. Prints counts, never
|
||||
transcript text.
|
||||
|
||||
usage: scriberr-seam-check.py RESULT.json [--min-chunks N]
|
||||
usage: scriberr-seam-check.py RESULT.json [--min-chunks N] [--standard]
|
||||
--standard the short-audio script's result (no buffered/num_chunks keys)
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
@@ -42,6 +43,8 @@ def main():
|
||||
parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py")
|
||||
parser.add_argument("--min-chunks", type=int, default=1,
|
||||
help="fail unless the run used at least this many chunks")
|
||||
parser.add_argument("--standard", action="store_true",
|
||||
help="the short-audio script's result: no buffered/num_chunks keys")
|
||||
args = parser.parse_args()
|
||||
min_chunks = args.min_chunks
|
||||
try:
|
||||
@@ -56,12 +59,13 @@ def main():
|
||||
for key, kind in required.items():
|
||||
if not isinstance(data.get(key), kind):
|
||||
fail(f"'{key}' missing or not {kind.__name__}")
|
||||
if data.get("buffered") is not True:
|
||||
fail("'buffered' is not true")
|
||||
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
|
||||
fail("'chunk_duration_secs' is not a number")
|
||||
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
|
||||
fail(f"'num_chunks' is not an integer >= {min_chunks}")
|
||||
if not args.standard:
|
||||
if data.get("buffered") is not True:
|
||||
fail("'buffered' is not true")
|
||||
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
|
||||
fail("'chunk_duration_secs' is not a number")
|
||||
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
|
||||
fail(f"'num_chunks' is not an integer >= {min_chunks}")
|
||||
|
||||
words, segments = data["word_timestamps"], data["segment_timestamps"]
|
||||
if not words or not data["transcription"].strip():
|
||||
@@ -79,7 +83,8 @@ def main():
|
||||
fail("segments do not cover the stitched words exactly once, in order")
|
||||
|
||||
print(f"SEAM OK: {len(words)} words, {len(segments)} segments, "
|
||||
f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points")
|
||||
f"{data.get('num_chunks', 1)} chunks, cuts at {len(data.get('cut_times', []))} points"
|
||||
f"{', model ' + data['model'] if data.get('model') else ''}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -181,4 +181,7 @@ Peak GPU memory on a 35-minute file:
|
||||
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
|
||||
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
|
||||
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
|
||||
patch; that is still open.
|
||||
patch. **Investigated 2026-09-30** (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`):
|
||||
the losses are real (against ground truth) and belong to the v2/v3 weights over
|
||||
long windows; a proposed patch, `patches/proposed/0002`, re-transcribes speech that
|
||||
got no words and cuts them 80–90 %. It is not deployed; that is Prime's call.
|
||||
|
||||
@@ -8,6 +8,8 @@ distinctly tagged image, and proves it before anyone deploys it.
|
||||
| patch | against | status |
|
||||
|---|---|---|
|
||||
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
|
||||
| `proposed/0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **proposed, not applied** (the rebuild script reads only this directory, not `proposed/`); built and tested as `scriberr:local-blackwell-a353078-dropout2`, not deployed. Why and how: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
|
||||
|
||||
|
||||
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
|
||||
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
|
||||
@@ -104,6 +106,17 @@ each other; the default is the simplest of them.
|
||||
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
|
||||
transcripts, for upstream's slicer too). See the bench doc.
|
||||
|
||||
### 0002 (proposed) — gap retry and an explicit model path
|
||||
|
||||
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
|
||||
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
|
||||
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
|
||||
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
|
||||
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
|
||||
the Go adapter record it as `ModelUsed`. To adopt: move it up into this directory and
|
||||
rebuild. To test-build it on top of the carried set under its own suffix:
|
||||
`scripts/scriberr-rebuild --suffix <name> --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
|
||||
|
||||
### Upgrading upstream
|
||||
|
||||
```bash
|
||||
|
||||
@@ -0,0 +1,719 @@
|
||||
From d253aa2b0ea2562ed4e1702380fac7e22bd11828 Mon Sep 17 00:00:00 2001
|
||||
From: Vuong Hoang <vh@phasefinal.com>
|
||||
Date: Wed, 30 Sep 2026 15:00:46 -0700
|
||||
Subject: [PATCH] feat(parakeet): selectable model and a retry for speech that
|
||||
got no words
|
||||
|
||||
Parakeet v2/v3 sometimes stop emitting for tens of seconds deep inside a long
|
||||
full-attention chunk while someone is talking; the same audio transcribed on
|
||||
its own is usually fine. After stitching, any stretch of at least
|
||||
--retry-gaps seconds (default 3) with no word, where most 25 ms frames sit
|
||||
within 12 dB of the recording's typical speech level, is re-transcribed on
|
||||
its own (in pieces of at most 60 s) and the words that land inside it are
|
||||
spliced in; segments are then rebuilt from the words with the model's own
|
||||
rule (end after . ? !). Output that triggers no retry is unchanged. The
|
||||
same retry runs in the short-audio script, which shares the helpers.
|
||||
|
||||
The retry is best effort: a failing piece is skipped and a failing retry
|
||||
keeps the first pass. A piece that looks like a hallucination loop (mostly
|
||||
one repeated token, or more than 7 words a second) is not spliced in, and a
|
||||
retried word that repeats its neighbour at the gap edge or a piece boundary
|
||||
is dropped (the first pass's copy wins). Chunk and retry audio go through a
|
||||
per-run temporary directory instead of fixed /tmp paths.
|
||||
|
||||
PARAKEET_MODEL_PATH (absolute, or relative to the env) selects the .nemo
|
||||
both scripts load; the default stays parakeet-tdt-0.6b-v3.nemo. The JSON
|
||||
"model" field reports the file actually loaded, and the Go adapter records
|
||||
it as ModelUsed instead of a hardcoded name.
|
||||
|
||||
The CLI and JSON seam is otherwise unchanged: the new flag is optional, and
|
||||
the JSON gains retried_gaps.
|
||||
---
|
||||
.../adapters/parakeet_adapter.go | 7 +-
|
||||
.../adapters/py/nvidia/parakeet_transcribe.py | 53 +++-
|
||||
.../py/nvidia/parakeet_transcribe_buffered.py | 231 +++++++++++++-----
|
||||
.../py/nvidia/tests/test_parakeet_slicing.py | 159 +++++++++++-
|
||||
4 files changed, 380 insertions(+), 70 deletions(-)
|
||||
|
||||
diff --git a/internal/transcription/adapters/parakeet_adapter.go b/internal/transcription/adapters/parakeet_adapter.go
|
||||
index 4fa252d..5f3bb9e 100644
|
||||
--- a/internal/transcription/adapters/parakeet_adapter.go
|
||||
+++ b/internal/transcription/adapters/parakeet_adapter.go
|
||||
@@ -342,7 +342,9 @@ func (p *ParakeetAdapter) Transcribe(ctx context.Context, input interfaces.Audio
|
||||
}
|
||||
|
||||
result.ProcessingTime = time.Since(startTime)
|
||||
- result.ModelUsed = "parakeet-tdt-0.6b-v3"
|
||||
+ if result.ModelUsed == "" {
|
||||
+ result.ModelUsed = "parakeet-tdt-0.6b-v3"
|
||||
+ }
|
||||
result.Metadata = p.CreateDefaultMetadata(params)
|
||||
|
||||
logger.Info("Parakeet transcription completed",
|
||||
@@ -525,6 +527,8 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu
|
||||
End float64 `json:"end"`
|
||||
} `json:"segment_timestamps"`
|
||||
Confidence interface{} `json:"confidence,omitempty"`
|
||||
+ // The scripts report the model they actually loaded (PARAKEET_MODEL_PATH can change it).
|
||||
+ Model string `json:"model"`
|
||||
}
|
||||
|
||||
if err := json.Unmarshal(data, ¶keetResult); err != nil {
|
||||
@@ -538,6 +542,7 @@ func (p *ParakeetAdapter) parseResult(tempDir string, input interfaces.AudioInpu
|
||||
Segments: make([]interfaces.TranscriptSegment, len(parakeetResult.SegmentTimestamps)),
|
||||
WordSegments: make([]interfaces.TranscriptWord, len(parakeetResult.WordTimestamps)),
|
||||
Confidence: 0.0, // Default confidence
|
||||
+ ModelUsed: parakeetResult.Model,
|
||||
}
|
||||
|
||||
// Convert segments
|
||||
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
|
||||
index 235e52a..6b2539e 100644
|
||||
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
|
||||
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
|
||||
@@ -7,9 +7,17 @@ import argparse
|
||||
import json
|
||||
import sys
|
||||
import os
|
||||
+import shutil
|
||||
+import tempfile
|
||||
from pathlib import Path
|
||||
+import librosa
|
||||
import nemo.collections.asr as nemo_asr
|
||||
|
||||
+# Shared with the long-audio script, which lives in the same directory.
|
||||
+from parakeet_transcribe_buffered import (
|
||||
+ DEFAULT_RETRY_GAP_SECS, resolve_model_path, retry_speech_gaps, transcribe_chunk,
|
||||
+)
|
||||
+
|
||||
|
||||
def transcribe_audio(
|
||||
audio_path: str,
|
||||
@@ -18,14 +26,11 @@ def transcribe_audio(
|
||||
context_left: int = 256,
|
||||
context_right: int = 256,
|
||||
include_confidence: bool = True,
|
||||
+ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS,
|
||||
):
|
||||
"""
|
||||
Transcribe audio using NVIDIA Parakeet model.
|
||||
"""
|
||||
- # Determine model path
|
||||
- model_filename = "parakeet-tdt-0.6b-v3.nemo"
|
||||
- model_path = None
|
||||
-
|
||||
# Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv
|
||||
virtual_env = os.environ.get("VIRTUAL_ENV")
|
||||
if not virtual_env:
|
||||
@@ -33,10 +38,11 @@ def transcribe_audio(
|
||||
sys.exit(1)
|
||||
|
||||
project_root = os.path.dirname(virtual_env)
|
||||
- model_path = os.path.join(project_root, model_filename)
|
||||
+ model_path = resolve_model_path(project_root)
|
||||
+ model_name = os.path.splitext(os.path.basename(model_path))[0]
|
||||
|
||||
if not os.path.exists(model_path):
|
||||
- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}")
|
||||
+ print(f"Error during transcription: Can't find model file: {model_path}")
|
||||
sys.exit(1)
|
||||
|
||||
print(f"Loading NVIDIA Parakeet model from: {model_path}")
|
||||
@@ -80,6 +86,27 @@ def transcribe_audio(
|
||||
text = result_data.text
|
||||
word_timestamps = result_data.timestamp.get("word", [])
|
||||
segment_timestamps = result_data.timestamp.get("segment", [])
|
||||
+ retried = 0
|
||||
+ words_added = False
|
||||
+ if retry_gap_secs > 0:
|
||||
+ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp
|
||||
+ try:
|
||||
+ audio, sr = librosa.load(audio_path, sr=None, mono=True)
|
||||
+
|
||||
+ def transcribe_span(start_sample, end_sample):
|
||||
+ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr,
|
||||
+ start_sample / sr, os.path.join(workdir, "retry.wav"))
|
||||
+ before = len(word_timestamps)
|
||||
+ word_timestamps, segment_timestamps, retried = retry_speech_gaps(
|
||||
+ word_timestamps, segment_timestamps, audio, sr, retry_gap_secs, transcribe_span
|
||||
+ )
|
||||
+ words_added = len(word_timestamps) != before
|
||||
+ if words_added:
|
||||
+ text = " ".join(w["word"] for w in word_timestamps)
|
||||
+ except Exception as err: # best effort: never lose a finished transcription
|
||||
+ print(f"Warning: gap retry failed ({err}); keeping the first pass")
|
||||
+ finally:
|
||||
+ shutil.rmtree(workdir, ignore_errors=True)
|
||||
|
||||
print(f"Transcription: {text}")
|
||||
|
||||
@@ -90,14 +117,15 @@ def transcribe_audio(
|
||||
"word_timestamps": word_timestamps,
|
||||
"segment_timestamps": segment_timestamps,
|
||||
"audio_file": audio_path,
|
||||
- "model": "parakeet-tdt-0.6b-v3",
|
||||
+ "model": model_name,
|
||||
+ "retried_gaps": retried,
|
||||
"context": {
|
||||
"left": context_left,
|
||||
"right": context_right
|
||||
}
|
||||
}
|
||||
|
||||
- if include_confidence:
|
||||
+ if include_confidence and not words_added: # first-pass scores no longer line up
|
||||
# Add confidence scores if available
|
||||
if hasattr(result_data, 'confidence') and result_data.confidence:
|
||||
output_data["confidence"] = result_data.confidence
|
||||
@@ -119,7 +147,7 @@ def transcribe_audio(
|
||||
"transcription": text,
|
||||
"language": "en",
|
||||
"audio_file": audio_path,
|
||||
- "model": "parakeet-tdt-0.6b-v3"
|
||||
+ "model": model_name
|
||||
}
|
||||
|
||||
if output_file:
|
||||
@@ -154,6 +182,12 @@ def main():
|
||||
"--context-right", type=int, default=256,
|
||||
help="Right attention context size (default: 256)"
|
||||
)
|
||||
+ parser.add_argument(
|
||||
+ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS,
|
||||
+ help=f"Re-transcribe on its own any stretch of at least this many seconds where "
|
||||
+ f"the audio holds speech but the model produced no words "
|
||||
+ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)"
|
||||
+ )
|
||||
parser.add_argument(
|
||||
"--include-confidence", action="store_true", default=True,
|
||||
help="Include confidence scores"
|
||||
@@ -178,6 +212,7 @@ def main():
|
||||
context_left=args.context_left,
|
||||
context_right=args.context_right,
|
||||
include_confidence=args.include_confidence,
|
||||
+ retry_gap_secs=args.retry_gaps,
|
||||
)
|
||||
except Exception as e:
|
||||
print(f"Error during transcription: {e}")
|
||||
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||
index ba755c1..49bb72e 100644
|
||||
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||
@@ -12,15 +12,21 @@ import argparse
|
||||
import json
|
||||
import sys
|
||||
import os
|
||||
+import shutil
|
||||
+import tempfile
|
||||
import librosa
|
||||
import soundfile as sf
|
||||
import numpy as np
|
||||
from pathlib import Path
|
||||
|
||||
+DEFAULT_MODEL = "parakeet-tdt-0.6b-v3.nemo"
|
||||
DEFAULT_OVERLAP_SECS = 4.0
|
||||
DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap
|
||||
QUIET_WINDOW_SECS = 0.3
|
||||
SAME_WORD_SECS = 0.5
|
||||
+DEFAULT_RETRY_GAP_SECS = 3.0
|
||||
+RETRY_PIECE_SECS = 60.0
|
||||
+RETRY_MAX_WORDS_PER_SEC = 7.0
|
||||
|
||||
|
||||
def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0):
|
||||
@@ -103,14 +109,7 @@ def stitch_slices(slice_results, cut_times, chunk_spans=None):
|
||||
elif len(kept) == stop - start:
|
||||
segments.append(seg)
|
||||
elif kept:
|
||||
- segments.append({
|
||||
- **seg,
|
||||
- "segment": " ".join(w["word"] for w in kept),
|
||||
- "start_offset": kept[0]["start_offset"],
|
||||
- "end_offset": kept[-1]["end_offset"],
|
||||
- "start": kept[0]["start"],
|
||||
- "end": kept[-1]["end"],
|
||||
- })
|
||||
+ segments.append({**seg, **_segment_of(kept)})
|
||||
return words, segments
|
||||
|
||||
|
||||
@@ -159,6 +158,134 @@ def _segment_ranges(words, segments):
|
||||
return ranges
|
||||
|
||||
|
||||
+def resolve_model_path(project_root):
|
||||
+ """The .nemo to load: $PARAKEET_MODEL_PATH (absolute, or relative to the env),
|
||||
+ else the bundled v3 weights."""
|
||||
+ chosen = os.environ.get("PARAKEET_MODEL_PATH", "").strip() or DEFAULT_MODEL
|
||||
+ return chosen if os.path.isabs(chosen) else os.path.join(project_root, chosen)
|
||||
+
|
||||
+
|
||||
+def find_speech_gaps(words, audio, sr, min_gap):
|
||||
+ """Stretches of at least `min_gap` s with no word, where at least half the
|
||||
+ 25 ms frames are within 12 dB of the recording's typical speech level: the
|
||||
+ model went quiet while someone was talking. Parakeet sometimes does this for
|
||||
+ tens of seconds deep inside a long chunk; the same audio transcribed on its
|
||||
+ own is usually fine."""
|
||||
+ hop, win = int(0.010 * sr), int(0.025 * sr)
|
||||
+ frames = max(0, (len(audio) - win) // hop)
|
||||
+ if frames == 0:
|
||||
+ return []
|
||||
+ # A strided view, reduced in blocks: bounded memory even for hours of audio.
|
||||
+ view = np.lib.stride_tricks.sliding_window_view(audio, win)[::hop][:frames]
|
||||
+ power = np.empty(frames)
|
||||
+ for k in range(0, frames, 8192):
|
||||
+ power[k:k + 8192] = np.mean(view[k:k + 8192].astype(np.float64) ** 2, axis=1)
|
||||
+ level = 10 * np.log10(power + 1e-12)
|
||||
+ speech = float(np.median(level[level >= np.median(level)]))
|
||||
+ duration = len(audio) / sr
|
||||
+ gaps = []
|
||||
+ for a, b in zip([0.0] + [w["end"] for w in words], [w["start"] for w in words] + [duration]):
|
||||
+ if b - a >= min_gap:
|
||||
+ span = level[int(a * sr) // hop:int(b * sr) // hop]
|
||||
+ if len(span) and np.mean(span >= speech - 12) >= 0.5:
|
||||
+ gaps.append((a, b))
|
||||
+ return gaps
|
||||
+
|
||||
+
|
||||
+def retry_speech_gaps(words, segments, audio, sr, min_gap, transcribe):
|
||||
+ """Re-transcribe each speech gap on its own, in pieces of at most
|
||||
+ RETRY_PIECE_SECS, and splice in the words that land inside it.
|
||||
+
|
||||
+ `transcribe(start_sample, end_sample)` returns (text, words, segments) in
|
||||
+ absolute time. Returns (words, segments, number of gaps retried); when words
|
||||
+ were added, segments are rebuilt from the words with the model's own rule.
|
||||
+ """
|
||||
+ gaps = find_speech_gaps(words, audio, sr, min_gap)
|
||||
+ if not gaps:
|
||||
+ return words, segments, 0
|
||||
+ found = []
|
||||
+ for a, b in gaps:
|
||||
+ t = a
|
||||
+ while t < b:
|
||||
+ e = min(b, t + RETRY_PIECE_SECS)
|
||||
+ try:
|
||||
+ _, piece, _ = transcribe(max(0, int((t - 0.5) * sr)), min(len(audio), int((e + 0.5) * sr)))
|
||||
+ except Exception as err: # best effort: the first pass already succeeded
|
||||
+ print(f"Warning: retry of {t:.1f}-{e:.1f}s failed ({err}); keeping the first pass there")
|
||||
+ piece = []
|
||||
+ piece = [w for w in piece if t <= w["start"] < e]
|
||||
+ if _plausible(piece, e - t):
|
||||
+ found.extend(piece)
|
||||
+ t = e
|
||||
+ if not found:
|
||||
+ return words, segments, len(gaps)
|
||||
+ merged = _without_retried_duplicates(sorted(words + found, key=lambda w: w["start"]),
|
||||
+ {id(w) for w in found})
|
||||
+ return merged, segments_from_words(merged), len(gaps)
|
||||
+
|
||||
+
|
||||
+def _plausible(piece, seconds):
|
||||
+ """False for what looks like a hallucination loop rather than speech: many words
|
||||
+ that are mostly one repeated token, or more words per second than anyone says."""
|
||||
+ if len(piece) >= 10 and len({_normalize(w["word"]) for w in piece}) < 0.3 * len(piece):
|
||||
+ return False
|
||||
+ return len(piece) <= RETRY_MAX_WORDS_PER_SEC * max(seconds, 1.0)
|
||||
+
|
||||
+
|
||||
+def _without_retried_duplicates(merged, retried):
|
||||
+ """The retry's padding re-hears the words either side of a gap, and a word
|
||||
+ near a piece boundary is heard by both pieces. Drop a retried word that
|
||||
+ repeats its neighbour within SAME_WORD_SECS; the first pass's copy wins."""
|
||||
+ out = []
|
||||
+ for w in merged:
|
||||
+ if out and _normalize(out[-1]["word"]) == _normalize(w["word"]) \
|
||||
+ and w["start"] - out[-1]["start"] <= SAME_WORD_SECS:
|
||||
+ if id(w) in retried:
|
||||
+ continue
|
||||
+ if id(out[-1]) in retried:
|
||||
+ out[-1] = w
|
||||
+ continue
|
||||
+ out.append(w)
|
||||
+ return out
|
||||
+
|
||||
+
|
||||
+def segments_from_words(words):
|
||||
+ """Segments as Parakeet's decoding config makes them: a segment ends after a
|
||||
+ word ending in '.', '?' or '!' (segment_seperators, no gap threshold)."""
|
||||
+ segments, current = [], []
|
||||
+ for w in words:
|
||||
+ current.append(w)
|
||||
+ if w["word"].endswith((".", "?", "!")):
|
||||
+ segments.append(_segment_of(current))
|
||||
+ current = []
|
||||
+ if current:
|
||||
+ segments.append(_segment_of(current))
|
||||
+ return segments
|
||||
+
|
||||
+
|
||||
+def _segment_of(words):
|
||||
+ return {"segment": " ".join(w["word"] for w in words),
|
||||
+ "start_offset": words[0]["start_offset"], "end_offset": words[-1]["end_offset"],
|
||||
+ "start": words[0]["start"], "end": words[-1]["end"]}
|
||||
+
|
||||
+
|
||||
+def transcribe_chunk(asr_model, audio, sr, start_time, path):
|
||||
+ """Transcribe one chunk via a WAV file; word and segment times become absolute."""
|
||||
+ sf.write(path, audio, sr)
|
||||
+ try:
|
||||
+ result = asr_model.transcribe([path], batch_size=1, timestamps=True)[0]
|
||||
+ finally:
|
||||
+ if os.path.exists(path):
|
||||
+ os.remove(path)
|
||||
+ words, segments = [], []
|
||||
+ if getattr(result, "timestamp", None):
|
||||
+ words = [dict(w, start=w["start"] + start_time, end=w["end"] + start_time)
|
||||
+ for w in result.timestamp.get("word", [])]
|
||||
+ segments = [dict(g, start=g["start"] + start_time, end=g["end"] + start_time)
|
||||
+ for g in result.timestamp.get("segment", [])]
|
||||
+ return result.text, words, segments
|
||||
+
|
||||
+
|
||||
def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0):
|
||||
"""Split audio file into chunks of at most chunk_duration_secs."""
|
||||
audio, sr = librosa.load(audio_path, sr=None, mono=True)
|
||||
@@ -173,7 +300,7 @@ def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, sear
|
||||
'duration': len(chunk_audio) / sr
|
||||
})
|
||||
|
||||
- return chunks, sr, [cut / sr for cut in cuts]
|
||||
+ return chunks, sr, [cut / sr for cut in cuts], audio
|
||||
|
||||
|
||||
def transcribe_buffered(
|
||||
@@ -182,16 +309,13 @@ def transcribe_buffered(
|
||||
chunk_duration_secs: float = 300, # 5 minutes default
|
||||
overlap_secs: float = DEFAULT_OVERLAP_SECS,
|
||||
pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS,
|
||||
+ retry_gap_secs: float = DEFAULT_RETRY_GAP_SECS,
|
||||
):
|
||||
"""
|
||||
Transcribe long audio by splitting into chunks and merging results.
|
||||
"""
|
||||
import nemo.collections.asr as nemo_asr
|
||||
|
||||
- # Determine model path
|
||||
- model_filename = "parakeet-tdt-0.6b-v3.nemo"
|
||||
- model_path = None
|
||||
-
|
||||
# Locate project root: derived from VIRTUAL_ENV, which is set by `uv run` to path/.venv
|
||||
virtual_env = os.environ.get("VIRTUAL_ENV")
|
||||
if not virtual_env:
|
||||
@@ -199,10 +323,11 @@ def transcribe_buffered(
|
||||
sys.exit(1)
|
||||
|
||||
project_root = os.path.dirname(virtual_env)
|
||||
- model_path = os.path.join(project_root, model_filename)
|
||||
+ model_path = resolve_model_path(project_root)
|
||||
+ model_name = os.path.splitext(os.path.basename(model_path))[0]
|
||||
|
||||
if not os.path.exists(model_path):
|
||||
- print(f"Error during transcription: Can't find {model_filename} in project root: {project_root}")
|
||||
+ print(f"Error during transcription: Can't find model file: {model_path}")
|
||||
sys.exit(1)
|
||||
|
||||
print(f"Loading NVIDIA Parakeet model from: {model_path}")
|
||||
@@ -226,58 +351,25 @@ def transcribe_buffered(
|
||||
|
||||
print(f"Splitting audio into chunks of at most {chunk_duration_secs}s "
|
||||
f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...")
|
||||
- chunks, sr, cut_times = split_audio_file(
|
||||
+ chunks, sr, cut_times, audio = split_audio_file(
|
||||
audio_path, chunk_duration_secs, overlap_secs, pause_search_secs
|
||||
)
|
||||
print(f"Created {len(chunks)} chunks")
|
||||
|
||||
slice_results = []
|
||||
chunk_texts = []
|
||||
+ retried = 0
|
||||
+ workdir = tempfile.mkdtemp(prefix="parakeet-") # per run: concurrent jobs share /tmp
|
||||
|
||||
for i, chunk_info in enumerate(chunks):
|
||||
print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...")
|
||||
-
|
||||
- # Save chunk to temporary file
|
||||
- chunk_path = f"/tmp/chunk_{i}.wav"
|
||||
- sf.write(chunk_path, chunk_info['audio'], sr)
|
||||
-
|
||||
- try:
|
||||
- # Transcribe chunk
|
||||
- output = asr_model.transcribe(
|
||||
- [chunk_path],
|
||||
- batch_size=1,
|
||||
- timestamps=True,
|
||||
- )
|
||||
-
|
||||
- result_data = output[0]
|
||||
- chunk_text = result_data.text
|
||||
- chunk_texts.append(chunk_text)
|
||||
- chunk_words = []
|
||||
- chunk_segments = []
|
||||
-
|
||||
- # Extract and adjust timestamps
|
||||
- if hasattr(result_data, 'timestamp') and result_data.timestamp:
|
||||
- # Adjust timestamps by chunk start time
|
||||
- for word in result_data.timestamp.get("word", []):
|
||||
- word_copy = dict(word)
|
||||
- word_copy['start'] += chunk_info['start_time']
|
||||
- word_copy['end'] += chunk_info['start_time']
|
||||
- chunk_words.append(word_copy)
|
||||
-
|
||||
- for segment in result_data.timestamp.get("segment", []):
|
||||
- seg_copy = dict(segment)
|
||||
- seg_copy['start'] += chunk_info['start_time']
|
||||
- seg_copy['end'] += chunk_info['start_time']
|
||||
- chunk_segments.append(seg_copy)
|
||||
-
|
||||
- slice_results.append((chunk_words, chunk_segments))
|
||||
-
|
||||
- print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
|
||||
-
|
||||
- finally:
|
||||
- # Clean up temp file
|
||||
- if os.path.exists(chunk_path):
|
||||
- os.remove(chunk_path)
|
||||
+ chunk_text, chunk_words, chunk_segments = transcribe_chunk(
|
||||
+ asr_model, chunk_info['audio'], sr, chunk_info['start_time'],
|
||||
+ os.path.join(workdir, f"chunk_{i}.wav"),
|
||||
+ )
|
||||
+ chunk_texts.append(chunk_text)
|
||||
+ slice_results.append((chunk_words, chunk_segments))
|
||||
+ print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
|
||||
|
||||
chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks]
|
||||
all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans)
|
||||
@@ -287,7 +379,20 @@ def transcribe_buffered(
|
||||
print("Warning: a chunk has text but no word timestamps; joining chunk texts")
|
||||
final_text = " ".join(chunk_texts)
|
||||
else:
|
||||
+ if retry_gap_secs > 0:
|
||||
+ def transcribe_span(start_sample, end_sample):
|
||||
+ return transcribe_chunk(asr_model, audio[start_sample:end_sample], sr,
|
||||
+ start_sample / sr, os.path.join(workdir, "retry.wav"))
|
||||
+ try:
|
||||
+ all_words, all_segments, retried = retry_speech_gaps(
|
||||
+ all_words, all_segments, audio, sr, retry_gap_secs, transcribe_span
|
||||
+ )
|
||||
+ except Exception as err: # best effort: never lose a finished transcription
|
||||
+ print(f"Warning: gap retry failed ({err}); keeping the first pass")
|
||||
+ if retried:
|
||||
+ print(f"Re-transcribed {retried} stretch(es) of speech that got no words")
|
||||
final_text = " ".join(w["word"] for w in all_words)
|
||||
+ shutil.rmtree(workdir, ignore_errors=True)
|
||||
print(f"Transcription complete: {len(final_text)} characters total")
|
||||
|
||||
output_data = {
|
||||
@@ -296,13 +401,14 @@ def transcribe_buffered(
|
||||
"word_timestamps": all_words,
|
||||
"segment_timestamps": all_segments,
|
||||
"audio_file": audio_path,
|
||||
- "model": "parakeet-tdt-0.6b-v3",
|
||||
+ "model": model_name,
|
||||
"buffered": True,
|
||||
"chunk_duration_secs": chunk_duration_secs,
|
||||
"num_chunks": len(chunks),
|
||||
"overlap_secs": overlap_secs,
|
||||
"pause_search_secs": pause_search_secs,
|
||||
"cut_times": cut_times,
|
||||
+ "retried_gaps": retried,
|
||||
}
|
||||
|
||||
if output_file:
|
||||
@@ -328,6 +434,12 @@ def main():
|
||||
help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len "
|
||||
f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)"
|
||||
)
|
||||
+ parser.add_argument(
|
||||
+ "--retry-gaps", type=float, default=DEFAULT_RETRY_GAP_SECS,
|
||||
+ help=f"Re-transcribe on its own any stretch of at least this many seconds where "
|
||||
+ f"the audio holds speech but the model produced no words "
|
||||
+ f"(default: {DEFAULT_RETRY_GAP_SECS}; 0 disables)"
|
||||
+ )
|
||||
parser.add_argument(
|
||||
"--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS,
|
||||
help=f"Seconds before each chunk limit searched for the quietest point to cut at, "
|
||||
@@ -346,6 +458,7 @@ def main():
|
||||
chunk_duration_secs=args.chunk_len,
|
||||
overlap_secs=args.overlap,
|
||||
pause_search_secs=args.pause_search,
|
||||
+ retry_gap_secs=args.retry_gaps,
|
||||
)
|
||||
|
||||
|
||||
diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||
index 6a35947..c3d7537 100644
|
||||
--- a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||
+++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||
@@ -10,7 +10,10 @@ import numpy as np
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
-from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402
|
||||
+from parakeet_transcribe_buffered import ( # noqa: E402
|
||||
+ find_speech_gaps, plan_slices, resolve_model_path, retry_speech_gaps,
|
||||
+ segments_from_words, stitch_slices,
|
||||
+)
|
||||
|
||||
SR = 16000
|
||||
|
||||
@@ -327,3 +330,157 @@ def test_punctuation_alone_is_never_an_anchor():
|
||||
right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)]
|
||||
words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||
assert words[0] is left[0] and texts(words) == ["-", "no"]
|
||||
+
|
||||
+
|
||||
+# -- model and attention selection ----------------------------------------------
|
||||
+
|
||||
+
|
||||
+def test_the_bundled_v3_model_is_the_default(monkeypatch):
|
||||
+ monkeypatch.delenv("PARAKEET_MODEL_PATH", raising=False)
|
||||
+ assert resolve_model_path("/env") == "/env/parakeet-tdt-0.6b-v3.nemo"
|
||||
+
|
||||
+
|
||||
+def test_parakeet_model_path_overrides_the_model(monkeypatch):
|
||||
+ monkeypatch.setenv("PARAKEET_MODEL_PATH", "/models/parakeet-tdt-0.6b-v2.nemo")
|
||||
+ assert resolve_model_path("/env") == "/models/parakeet-tdt-0.6b-v2.nemo"
|
||||
+ monkeypatch.setenv("PARAKEET_MODEL_PATH", "other.nemo")
|
||||
+ assert resolve_model_path("/env") == "/env/other.nemo"
|
||||
+
|
||||
+
|
||||
+# -- retrying stretches where the model went quiet --------------------------------
|
||||
+
|
||||
+
|
||||
+def talk(seconds, level=0.1, seed=3):
|
||||
+ return (level * np.random.default_rng(seed).standard_normal(int(seconds * SR))).astype(np.float32)
|
||||
+
|
||||
+
|
||||
+def w_at(text, start, end):
|
||||
+ return {"word": text, "start_offset": 0, "end_offset": 0, "start": start, "end": end}
|
||||
+
|
||||
+
|
||||
+def test_a_long_wordless_stretch_over_speech_is_a_gap():
|
||||
+ audio = talk(40)
|
||||
+ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 38.0, 38.5)]
|
||||
+ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (10.0, 20.0), (20.5, 38.0)]
|
||||
+
|
||||
+
|
||||
+def test_a_wordless_stretch_over_silence_or_a_short_pause_is_not():
|
||||
+ audio = talk(40)
|
||||
+ audio[int(10 * SR):int(20 * SR)] = 0.0
|
||||
+ words = [w_at("a", 1.0, 1.5), w_at("b", 9.5, 10.0), w_at("c", 20.0, 20.5), w_at("d", 22.0, 22.5),
|
||||
+ w_at("e", 24.0, 24.5), w_at("f", 39.0, 39.5)]
|
||||
+ # 10-20 s is silent; 20.5-22 and 22.5-24 are too short; 1.5-9.5 and 24.5-39 are talk
|
||||
+ assert find_speech_gaps(words, audio, SR, min_gap=3.0) == [(1.5, 9.5), (24.5, 39.0)]
|
||||
+
|
||||
+
|
||||
+def test_speech_before_the_first_word_counts():
|
||||
+ audio = talk(20)
|
||||
+ assert find_speech_gaps([w_at("late", 15.0, 15.5)], audio, SR, min_gap=3.0) == [(0.0, 15.0), (15.5, 20.0)]
|
||||
+
|
||||
+
|
||||
+def test_no_words_at_all_over_speech_is_one_gap():
|
||||
+ assert find_speech_gaps([], talk(12), SR, min_gap=3.0) == [(0.0, 12.0)]
|
||||
+
|
||||
+
|
||||
+def test_retried_words_are_spliced_into_the_gap_and_segments_rebuilt():
|
||||
+ audio = talk(31)
|
||||
+ words = [w_at("Hello", 1.0, 1.5), w_at("there.", 2.0, 2.5), w_at("Bye.", 30.0, 30.5)]
|
||||
+ segments = [segment(words[:2]), segment(words[2:])]
|
||||
+ calls = []
|
||||
+
|
||||
+ def transcribe(s0, s1): # stands in for Parakeet: finds speech the main pass missed
|
||||
+ calls.append((s0 / SR, s1 / SR))
|
||||
+ found = [w_at("We", 5.0, 5.2), w_at("missed", 6.0, 6.3), w_at("this.", 7.0, 7.4),
|
||||
+ w_at("Again", 12.0, 12.4), w_at("outside", 40.5, 41.0)]
|
||||
+ return "", [w for w in found if s0 / SR <= w["start"] < s1 / SR], []
|
||||
+ new_words, new_segments, n = retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe)
|
||||
+ assert [w["word"] for w in new_words] == ["Hello", "there.", "We", "missed", "this.", "Again", "Bye."]
|
||||
+ assert n == 1 and calls == [(2.0, 30.5)] # one gap, 2.5-30.0 s, padded by 0.5 s
|
||||
+ assert [g["segment"] for g in new_segments] == ["Hello there.", "We missed this.", "Again Bye."]
|
||||
+ assert " ".join(g["segment"] for g in new_segments) == " ".join(w["word"] for w in new_words)
|
||||
+
|
||||
+
|
||||
+def test_without_gaps_nothing_is_retried_and_output_is_untouched():
|
||||
+ audio = talk(10)
|
||||
+ words = [w_at("One", 0.5, 1.0), w_at("two", 2.5, 3.0), w_at("three.", 5.0, 5.5), w_at("four", 8.0, 8.5)]
|
||||
+ segments = [segment(words[:3]), segment(words[3:])]
|
||||
+
|
||||
+ def transcribe(s0, s1):
|
||||
+ raise AssertionError("must not be called")
|
||||
+ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe) == (words, segments, 0)
|
||||
+
|
||||
+
|
||||
+def test_long_gaps_are_retried_in_pieces():
|
||||
+ audio = talk(191)
|
||||
+ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)]
|
||||
+ calls = []
|
||||
+
|
||||
+ def transcribe(s0, s1):
|
||||
+ calls.append(round((s1 - s0) / SR, 1))
|
||||
+ return "", [], []
|
||||
+ retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
|
||||
+ assert max(calls) <= 61.0 and len(calls) == 4 # 1.5-190 s in <= 60 s pieces (+0.5 s padding)
|
||||
+
|
||||
+
|
||||
+def test_segments_from_words_split_after_sentence_punctuation():
|
||||
+ ws = [w_at("Yes.", 0, 1), w_at("Is", 1, 2), w_at("it?", 2, 3), w_at("Go", 3, 4), w_at("now", 4, 5)]
|
||||
+ assert [g["segment"] for g in segments_from_words(ws)] == ["Yes.", "Is it?", "Go now"]
|
||||
+ assert segments_from_words([]) == []
|
||||
+
|
||||
+
|
||||
+def test_a_retried_copy_of_the_word_after_the_gap_is_not_kept_twice():
|
||||
+ # The retry piece is padded, so it re-hears the next kept word and may
|
||||
+ # timestamp it just inside the gap.
|
||||
+ audio = talk(31)
|
||||
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
|
||||
+
|
||||
+ def transcribe(s0, s1):
|
||||
+ return "", [w_at("We", 5.0, 5.2), w_at("left.", 6.0, 6.3), w_at("bye.", 29.7, 30.3)], []
|
||||
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
|
||||
+ assert [w["word"] for w in new_words] == ["Hello.", "We", "left.", "Bye."]
|
||||
+ assert new_words[-1] is words[-1]
|
||||
+
|
||||
+
|
||||
+def test_a_word_heard_by_two_retry_pieces_is_kept_once():
|
||||
+ audio = talk(191)
|
||||
+ words = [w_at("a", 1.0, 1.5), w_at("z", 190.0, 190.5)]
|
||||
+
|
||||
+ def transcribe(s0, s1): # both pieces around 61.5 s hear "boundary"
|
||||
+ t0, t1 = s0 / SR, s1 / SR
|
||||
+ out = [w_at("boundary", 61.3, 61.8)] if t0 <= 61.3 < t1 else []
|
||||
+ out += [w_at("boundary", 61.55, 61.9)] if t0 <= 61.55 < t1 and t0 > 60 else []
|
||||
+ return "", out, []
|
||||
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
|
||||
+ assert [w["word"] for w in new_words] == ["a", "boundary", "z"]
|
||||
+
|
||||
+
|
||||
+def test_a_failing_retry_keeps_the_first_pass():
|
||||
+ audio = talk(31)
|
||||
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
|
||||
+ segments = [segment(words[:1]), segment(words[1:])]
|
||||
+
|
||||
+ def transcribe(s0, s1):
|
||||
+ raise RuntimeError("CUDA out of memory")
|
||||
+ assert retry_speech_gaps(words, segments, audio, SR, 3.0, transcribe)[:2] == (words, segments)
|
||||
+
|
||||
+
|
||||
+def test_a_looping_retry_is_not_spliced_in():
|
||||
+ audio = talk(31)
|
||||
+ words = [w_at("Hello.", 1.0, 1.5), w_at("Bye.", 30.0, 30.5)]
|
||||
+
|
||||
+ def transcribe(s0, s1): # a hallucination loop over music-like audio
|
||||
+ return "", [w_at("la", 3.0 + k, 3.2 + k) for k in range(20)], []
|
||||
+ new_words, _, _ = retry_speech_gaps(words, [], audio, SR, 3.0, transcribe)
|
||||
+ assert [w["word"] for w in new_words] == ["Hello.", "Bye."]
|
||||
+
|
||||
+
|
||||
+def test_gap_levels_line_up_with_time_at_sample_rates_other_than_16k():
|
||||
+ # At 11025 Hz a 10 ms hop is 110 samples (9.977 ms); indexing frames by
|
||||
+ # time/10ms would drift ~2 s by 900 s and read the silence before the gap.
|
||||
+ sr = 11025
|
||||
+ audio = (0.1 * np.random.default_rng(5).standard_normal(1000 * sr)).astype(np.float32)
|
||||
+ audio[int(896.9 * sr):int(900.0 * sr)] = 0.0
|
||||
+ words = [w_at("w", t, t + 0.5) for t in np.arange(0.0, 1000.0, 1.0) if not 900.0 <= t < 903.0]
|
||||
+ words = [w for w in words if not (899.9 < w["start"] < 900.5)] + [w_at("w", 899.5, 900.0)]
|
||||
+ words.sort(key=lambda w: w["start"])
|
||||
+ assert (900.0, 903.0) in find_speech_gaps(words, audio, sr, 3.0)
|
||||
--
|
||||
2.39.5
|
||||
|
||||
Reference in New Issue
Block a user