# Upstream PR — prepared, NOT opened **Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub. - **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one commit authored by Vuong Hoang against upstream `a353078` (current HEAD, 2026-09-20). It is the same file we carry, so the PR and our build cannot drift. - **To open it (after the yes):** fork on GitHub, then `git clone && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 && git am /0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`, and open the PR with the title and body below. - **Before sending, decide:** keep the opt-in `--pause-search` in the upstream version, or drop it for a smaller diff (about 40 lines with its tests). It measured neutral; see the body's last paragraph. --- ## Title Parakeet buffered transcription: overlap chunks and stitch at an agreed word ## Body ### Problem For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py` cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a mark is chopped in two, lost, or transcribed twice. We measured it on four recordings (118 minutes, see below): **52 % of the cuts had a transcription error within ±3 s of them, against 19 % at points midway between cuts.** Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the knob is for on smaller cards, makes this worse, because there are more cuts. ### Change One Python file, plus a new test file. No Go changes; the CLI and JSON that `parakeet_adapter.go` reads are unchanged. 1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk gets longer and peak GPU memory is unchanged. 2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the word nearest the cut that both transcribed alike: same text after lowercasing and stripping punctuation, start times within 0.5 s. The left chunk keeps the words before it and the right chunk the rest, with the anchor taken from whichever chunk keeps word starts in time order. With no agreed word, both split at the cut. Segments are trimmed to match, and `transcription` is the stitched words joined by spaces (which is what it already equals for Parakeet). The obvious rule, "keep each word from the chunk whose half its start time falls in", is not enough. A word that follows a pause can be timestamped anywhere in the pause, so the two chunks often put the *same* word on opposite sides of the cut, one frame apart, and it comes out twice (or not at all). 3. `--pause-search N` (opt-in, off by default): move each cut back to the quietest 0.3 s within the last N seconds before the limit. 4. NeMo is imported inside `transcribe_buffered()` so the pure helpers (`plan_slices`, `stitch_slices`) can be unit-tested without a GPU. New flags are optional with defaults, because the Go side does not pass them. The JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0` reproduces the previous output exactly (words, segments and text were byte-identical on all four test files). ### Measurements Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3, 16 kHz mono input. Each run was scored against a **no-cut reference**: the same model over the whole file in one pass with local attention (`parakeet_transcribe.py --context-left 255 --context-right 255`). Words were aligned after lowercasing and stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it, and the same test at points midway between cuts gives the background rate. Each variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in different places. Decoding is deterministic, so repeats are identical. Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451, public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar Wilde* (public domain), and two private conversational recordings (22 and 35 min; numbers only). | slicer | damaged cuts | background | error events near cuts | words duplicated at cuts | |---|---|---|---|---| | current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 | | overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 | | **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 | | pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 | | pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 | The 2-standard-error band on a difference of these rates is about ±0.08, so the last three rows are statistically tied and all beat the current slicer by a wide margin. The public files alone tell the same story: current 52/86 damaged cuts; this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`). Pause-aware cutting is included but off by default: once the stitch was right, it did not measurably help, and without an overlap it drops words just before its cuts. Happy to drop it from this PR if you prefer the smaller diff. ### Tests ``` cd internal/transcription/adapters/py/nvidia python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU ``` The tests cover synthetic waveforms with known pauses, audio shorter than one chunk, a pause at the very start or end, audio with no pause, the chunk limit with the overlap included, and stitching with known word lists, including a word the two chunks timestamp on either side of the cut, a word one chunk missed, a longer text match at a different time, an anchor that would reverse time order, and punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py` still passes, since its `--chunk-len 10` run now also exercises the overlap.