Files
esh-pfi-infrastructure/stacks/scriberr/patches/upstream-pr/PR.md
T
vh ee3db68db1 feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
2026-09-30 12:10:32 -07:00

6.0 KiB

Upstream PR — prepared, NOT opened

Status: ready to submit to rishikanthc/Scriberr, held for Prime's explicit yes (opening a PR is outward-facing). Nothing has been pushed to GitHub.

  • Diff: ../0001-parakeet-pause-aware-slicer.patch, a git format-patch of one commit authored by Vuong Hoang against upstream a353078 (current HEAD, 2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
  • To open it (after the yes): fork on GitHub, then git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 && git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch, and open the PR with the title and body below.
  • Before sending, decide: keep the opt-in --pause-search in the upstream version, or drop it for a smaller diff (about 40 lines with its tests). It measured neutral; see the body's last paragraph.

Title

Parakeet buffered transcription: overlap chunks and stitch at an agreed word

Body

Problem

For audio longer than PARAKEET_CHUNK_THRESHOLD_SECS, parakeet_transcribe_buffered.py cuts the file at fixed --chunk-len marks with no overlap. A word that straddles a mark is chopped in two, lost, or transcribed twice. We measured it on four recordings (118 minutes, see below): 52 % of the cuts had a transcription error within ±3 s of them, against 19 % at points midway between cuts.

Lowering PARAKEET_CHUNK_THRESHOLD_SECS to save GPU memory, which is what the knob is for on smaller cards, makes this worse, because there are more cuts.

Change

One Python file, plus a new test file. No Go changes; the CLI and JSON that parakeet_adapter.go reads are unchanged.

  1. Overlap. Adjacent chunks share --overlap seconds (default 4), half on each side of the cut. The overlap counts inside --chunk-len, so no chunk gets longer and peak GPU memory is unchanged.

  2. Stitch at an agreed word. In each overlap, the two chunks hand over at the word nearest the cut that both transcribed alike: same text after lowercasing and stripping punctuation, start times within 0.5 s. The left chunk keeps the words before it and the right chunk the rest, with the anchor taken from whichever chunk keeps word starts in time order. With no agreed word, both split at the cut. Segments are trimmed to match, and transcription is the stitched words joined by spaces (which is what it already equals for Parakeet).

    The obvious rule, "keep each word from the chunk whose half its start time falls in", is not enough. A word that follows a pause can be timestamped anywhere in the pause, so the two chunks often put the same word on opposite sides of the cut, one frame apart, and it comes out twice (or not at all).

  3. --pause-search N (opt-in, off by default): move each cut back to the quietest 0.3 s within the last N seconds before the limit.

  4. NeMo is imported inside transcribe_buffered() so the pure helpers (plan_slices, stitch_slices) can be unit-tested without a GPU.

New flags are optional with defaults, because the Go side does not pass them. The JSON gains overlap_secs, pause_search_secs and cut_times. --overlap 0 reproduces the previous output exactly (words, segments and text were byte-identical on all four test files).

Measurements

Setup: RTX PRO 6000 Blackwell, the Dockerfile.cuda.12.9 image, parakeet-tdt-0.6b-v3, 16 kHz mono input. Each run was scored against a no-cut reference: the same model over the whole file in one pass with local attention (parakeet_transcribe.py --context-left 255 --context-right 255). Words were aligned after lowercasing and stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it, and the same test at points midway between cuts gives the background rate. Each variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in different places. Decoding is deterministic, so repeats are identical.

Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451, public domain), section 1 of the LibriVox dramatic reading The Trial of Oscar Wilde (public domain), and two private conversational recordings (22 and 35 min; numbers only).

slicer damaged cuts background error events near cuts words duplicated at cuts
current (fixed, no overlap) 93/179 = 52 % 19 % 102 17
overlap, start-time stitch 52/184 = 28 % 18 % 60 18
overlap, agreed-word stitch (this PR) 41/184 = 22 % 18 % 47 2
pause-aware cut, no overlap 53/201 = 26 % 18 % 59 0
pause-aware + overlap, agreed-word stitch 52/207 = 25 % 18 % 54 2

The 2-standard-error band on a difference of these rates is about ±0.08, so the last three rows are statistically tied and all beat the current slicer by a wide margin. The public files alone tell the same story: current 52/86 damaged cuts; this PR 22/89. Peak GPU memory for a 35-minute file with --chunk-len 120 is unchanged (5,496 MiB before and after, n=3, expandable_segments:True).

Pause-aware cutting is included but off by default: once the stitch was right, it did not measurably help, and without an overlap it drops words just before its cuts. Happy to drop it from this PR if you prefer the smaller diff.

Tests

cd internal/transcription/adapters/py/nvidia
python -m pytest tests/test_parakeet_slicing.py   # needs numpy, librosa, soundfile; no GPU

The tests cover synthetic waveforms with known pauses, audio shorter than one chunk, a pause at the very start or end, audio with no pause, the chunk limit with the overlap included, and stitching with known word lists, including a word the two chunks timestamp on either side of the cut, a word one chunk missed, a longer text match at a different time, an anchor that would reverse time order, and punctuation-only tokens. The existing test_parakeet_transcribe_buffered.py still passes, since its --chunk-len 10 run now also exercises the overlap.