Carry patches/0001 on our Scriberr build (upstream a353078): adjacent buffered chunks overlap by 4 s inside --chunk-len and hand over at a word both chunks transcribed alike, instead of cutting at fixed marks with no overlap. Pause-aware cutting is included as an opt-in (--pause-search); it measured neutral once the stitch was right. The Go<->Python CLI and JSON seam is unchanged. Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut whole-file reference; metrics only, private audio stays on fv-ml1): cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184) against a 19% background; floor +-0.08. Positive control: upstream's cutter +0.33 over background. A-vs-A byte-identical in-process and across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3). Also found: Parakeet skips runs of >=10 words mid-chunk with any slicer, upstream's included; not addressed here. scripts/scriberr-rebuild clones a pinned upstream sha into a new /opt/docker/src dir, git-apply-checks the patches, builds a distinct tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py) and the memory budget on idle GPU 3. Deploy stays manual. The upstream PR is prepared under patches/upstream-pr/ and not opened.
6.0 KiB
Upstream PR — prepared, NOT opened
Status: ready to submit to rishikanthc/Scriberr, held for Prime's explicit
yes (opening a PR is outward-facing). Nothing has been pushed to GitHub.
- Diff:
../0001-parakeet-pause-aware-slicer.patch, agit format-patchof one commit authored by Vuong Hoang against upstreama353078(current HEAD, 2026-09-20). It is the same file we carry, so the PR and our build cannot drift. - To open it (after the yes): fork on GitHub, then
git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 && git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch, and open the PR with the title and body below. - Before sending, decide: keep the opt-in
--pause-searchin the upstream version, or drop it for a smaller diff (about 40 lines with its tests). It measured neutral; see the body's last paragraph.
Title
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
Body
Problem
For audio longer than PARAKEET_CHUNK_THRESHOLD_SECS, parakeet_transcribe_buffered.py
cuts the file at fixed --chunk-len marks with no overlap. A word that straddles a
mark is chopped in two, lost, or transcribed twice. We measured it on four
recordings (118 minutes, see below): 52 % of the cuts had a transcription error
within ±3 s of them, against 19 % at points midway between cuts.
Lowering PARAKEET_CHUNK_THRESHOLD_SECS to save GPU memory, which is what the
knob is for on smaller cards, makes this worse, because there are more cuts.
Change
One Python file, plus a new test file. No Go changes; the CLI and JSON that
parakeet_adapter.go reads are unchanged.
-
Overlap. Adjacent chunks share
--overlapseconds (default 4), half on each side of the cut. The overlap counts inside--chunk-len, so no chunk gets longer and peak GPU memory is unchanged. -
Stitch at an agreed word. In each overlap, the two chunks hand over at the word nearest the cut that both transcribed alike: same text after lowercasing and stripping punctuation, start times within 0.5 s. The left chunk keeps the words before it and the right chunk the rest, with the anchor taken from whichever chunk keeps word starts in time order. With no agreed word, both split at the cut. Segments are trimmed to match, and
transcriptionis the stitched words joined by spaces (which is what it already equals for Parakeet).The obvious rule, "keep each word from the chunk whose half its start time falls in", is not enough. A word that follows a pause can be timestamped anywhere in the pause, so the two chunks often put the same word on opposite sides of the cut, one frame apart, and it comes out twice (or not at all).
-
--pause-search N(opt-in, off by default): move each cut back to the quietest 0.3 s within the last N seconds before the limit. -
NeMo is imported inside
transcribe_buffered()so the pure helpers (plan_slices,stitch_slices) can be unit-tested without a GPU.
New flags are optional with defaults, because the Go side does not pass them. The
JSON gains overlap_secs, pause_search_secs and cut_times. --overlap 0
reproduces the previous output exactly (words, segments and text were
byte-identical on all four test files).
Measurements
Setup: RTX PRO 6000 Blackwell, the Dockerfile.cuda.12.9 image, parakeet-tdt-0.6b-v3,
16 kHz mono input. Each run was scored against a no-cut reference: the same
model over the whole file in one pass with local attention (parakeet_transcribe.py --context-left 255 --context-right 255). Words were aligned after lowercasing and
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
and the same test at points midway between cuts gives the background rate. Each
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
different places. Decoding is deterministic, so repeats are identical.
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451, public domain), section 1 of the LibriVox dramatic reading The Trial of Oscar Wilde (public domain), and two private conversational recordings (22 and 35 min; numbers only).
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|---|---|---|---|---|
| current (fixed, no overlap) | 93/179 = 52 % | 19 % | 102 | 17 |
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
| overlap, agreed-word stitch (this PR) | 41/184 = 22 % | 18 % | 47 | 2 |
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
The 2-standard-error band on a difference of these rates is about ±0.08, so the
last three rows are statistically tied and all beat the current slicer by a wide
margin. The public files alone tell the same story: current 52/86 damaged cuts;
this PR 22/89. Peak GPU memory for a 35-minute file with --chunk-len 120 is
unchanged (5,496 MiB before and after, n=3, expandable_segments:True).
Pause-aware cutting is included but off by default: once the stitch was right, it did not measurably help, and without an overlap it drops words just before its cuts. Happy to drop it from this PR if you prefer the smaller diff.
Tests
cd internal/transcription/adapters/py/nvidia
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
a pause at the very start or end, audio with no pause, the chunk limit with the
overlap included, and stitching with known word lists, including a word the two
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
match at a different time, an anchor that would reverse time order, and
punctuation-only tokens. The existing test_parakeet_transcribe_buffered.py
still passes, since its --chunk-len 10 run now also exercises the overlap.