diff --git a/stacks/scriberr/patches/README.md b/stacks/scriberr/patches/README.md index 53cc8b0..7c217c7 100644 --- a/stacks/scriberr/patches/README.md +++ b/stacks/scriberr/patches/README.md @@ -7,7 +7,7 @@ distinctly tagged image, and proves it before anyone deploys it. | patch | against | status | |---|---|---| -| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) | +| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; superseded in production by 0002 on top of it; **no upstream PR (Prime, 2026-10-01)**. Both patches are carried locally only | | `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` | @@ -15,8 +15,7 @@ Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`, then `sudo -n docker compose up -d scriberr`. -Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is -outward-facing and needs Prime's explicit yes. +Ruling: Prime, 2026-09-30, "build the slicer". **No upstream contribution (Prime, 2026-10-01: "no upstream for scriberr").** The prepared PR text was removed (it remains in git history, ee3db68). ## 0001 — pause-aware Parakeet slicer diff --git a/stacks/scriberr/patches/upstream-pr/PR.md b/stacks/scriberr/patches/upstream-pr/PR.md deleted file mode 100644 index a537621..0000000 --- a/stacks/scriberr/patches/upstream-pr/PR.md +++ /dev/null @@ -1,113 +0,0 @@ -# Upstream PR — prepared, NOT opened - -**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit -yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub. - -- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one - commit authored by Vuong Hoang against upstream `a353078` (current HEAD, - 2026-09-20). It is the same file we carry, so the PR and our build cannot drift. -- **To open it (after the yes):** fork on GitHub, then - `git clone && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 && - git am /0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`, - and open the PR with the title and body below. -- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream - version, or drop it for a smaller diff (about 40 lines with its tests). It measured - neutral; see the body's last paragraph. - ---- - -## Title - -Parakeet buffered transcription: overlap chunks and stitch at an agreed word - -## Body - -### Problem - -For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py` -cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a -mark is chopped in two, lost, or transcribed twice. We measured it on four -recordings (118 minutes, see below): **52 % of the cuts had a transcription error -within ±3 s of them, against 19 % at points midway between cuts.** - -Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the -knob is for on smaller cards, makes this worse, because there are more cuts. - -### Change - -One Python file, plus a new test file. No Go changes; the CLI and JSON that -`parakeet_adapter.go` reads are unchanged. - -1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on - each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk - gets longer and peak GPU memory is unchanged. -2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the - word nearest the cut that both transcribed alike: same text after lowercasing - and stripping punctuation, start times within 0.5 s. The left chunk keeps the - words before it and the right chunk the rest, with the anchor taken from - whichever chunk keeps word starts in time order. With no agreed word, both - split at the cut. Segments are trimmed to match, and `transcription` is the - stitched words joined by spaces (which is what it already equals for Parakeet). - - The obvious rule, "keep each word from the chunk whose half its start time falls - in", is not enough. A word that follows a pause can be timestamped anywhere in - the pause, so the two chunks often put the *same* word on opposite sides of the - cut, one frame apart, and it comes out twice (or not at all). -3. `--pause-search N` (opt-in, off by default): move each cut back to the - quietest 0.3 s within the last N seconds before the limit. -4. NeMo is imported inside `transcribe_buffered()` so the pure helpers - (`plan_slices`, `stitch_slices`) can be unit-tested without a GPU. - -New flags are optional with defaults, because the Go side does not pass them. The -JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0` -reproduces the previous output exactly (words, segments and text were -byte-identical on all four test files). - -### Measurements - -Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3, -16 kHz mono input. Each run was scored against a **no-cut reference**: the same -model over the whole file in one pass with local attention (`parakeet_transcribe.py ---context-left 255 --context-right 255`). Words were aligned after lowercasing and -stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it, -and the same test at points midway between cuts gives the background rate. Each -variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in -different places. Decoding is deterministic, so repeats are identical. - -Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451, -public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar -Wilde* (public domain), and two private conversational recordings (22 and 35 min; -numbers only). - -| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts | -|---|---|---|---|---| -| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 | -| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 | -| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 | -| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 | -| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 | - -The 2-standard-error band on a difference of these rates is about ±0.08, so the -last three rows are statistically tied and all beat the current slicer by a wide -margin. The public files alone tell the same story: current 52/86 damaged cuts; -this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is -unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`). - -Pause-aware cutting is included but off by default: once the stitch was right, it -did not measurably help, and without an overlap it drops words just before its -cuts. Happy to drop it from this PR if you prefer the smaller diff. - -### Tests - -``` -cd internal/transcription/adapters/py/nvidia -python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU -``` - -The tests cover synthetic waveforms with known pauses, audio shorter than one chunk, -a pause at the very start or end, audio with no pause, the chunk limit with the -overlap included, and stitching with known word lists, including a word the two -chunks timestamp on either side of the cut, a word one chunk missed, a longer text -match at a different time, an anchor that would reverse time order, and -punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py` -still passes, since its `--chunk-len 10` run now also exercises the overlap.