feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench

Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
This commit is contained in:
vh
2026-09-30 12:10:32 -07:00
parent ab62644315
commit ee3db68db1
8 changed files with 1556 additions and 7 deletions
+113
View File
@@ -0,0 +1,113 @@
# Upstream PR — prepared, NOT opened
**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit
yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub.
- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one
commit authored by Vuong Hoang against upstream `a353078` (current HEAD,
2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
- **To open it (after the yes):** fork on GitHub, then
`git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 &&
git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`,
and open the PR with the title and body below.
- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream
version, or drop it for a smaller diff (about 40 lines with its tests). It measured
neutral; see the body's last paragraph.
---
## Title
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
## Body
### Problem
For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py`
cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a
mark is chopped in two, lost, or transcribed twice. We measured it on four
recordings (118 minutes, see below): **52 % of the cuts had a transcription error
within ±3 s of them, against 19 % at points midway between cuts.**
Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the
knob is for on smaller cards, makes this worse, because there are more cuts.
### Change
One Python file, plus a new test file. No Go changes; the CLI and JSON that
`parakeet_adapter.go` reads are unchanged.
1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on
each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk
gets longer and peak GPU memory is unchanged.
2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the
word nearest the cut that both transcribed alike: same text after lowercasing
and stripping punctuation, start times within 0.5 s. The left chunk keeps the
words before it and the right chunk the rest, with the anchor taken from
whichever chunk keeps word starts in time order. With no agreed word, both
split at the cut. Segments are trimmed to match, and `transcription` is the
stitched words joined by spaces (which is what it already equals for Parakeet).
The obvious rule, "keep each word from the chunk whose half its start time falls
in", is not enough. A word that follows a pause can be timestamped anywhere in
the pause, so the two chunks often put the *same* word on opposite sides of the
cut, one frame apart, and it comes out twice (or not at all).
3. `--pause-search N` (opt-in, off by default): move each cut back to the
quietest 0.3 s within the last N seconds before the limit.
4. NeMo is imported inside `transcribe_buffered()` so the pure helpers
(`plan_slices`, `stitch_slices`) can be unit-tested without a GPU.
New flags are optional with defaults, because the Go side does not pass them. The
JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0`
reproduces the previous output exactly (words, segments and text were
byte-identical on all four test files).
### Measurements
Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3,
16 kHz mono input. Each run was scored against a **no-cut reference**: the same
model over the whole file in one pass with local attention (`parakeet_transcribe.py
--context-left 255 --context-right 255`). Words were aligned after lowercasing and
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
and the same test at points midway between cuts gives the background rate. Each
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
different places. Decoding is deterministic, so repeats are identical.
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451,
public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar
Wilde* (public domain), and two private conversational recordings (22 and 35 min;
numbers only).
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|---|---|---|---|---|
| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 |
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 |
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
The 2-standard-error band on a difference of these rates is about ±0.08, so the
last three rows are statistically tied and all beat the current slicer by a wide
margin. The public files alone tell the same story: current 52/86 damaged cuts;
this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is
unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`).
Pause-aware cutting is included but off by default: once the stitch was right, it
did not measurably help, and without an overlap it drops words just before its
cuts. Happy to drop it from this PR if you prefer the smaller diff.
### Tests
```
cd internal/transcription/adapters/py/nvidia
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
```
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
a pause at the very start or end, audio with no pause, the chunk limit with the
overlap included, and stitching with known word lists, including a word the two
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
match at a different time, an anchor that would reverse time order, and
punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py`
still passes, since its `--chunk-len 10` run now also exercises the overlap.