scriberr: no upstream contribution (Prime) — drop prepared PR text; patches carried locally

This commit is contained in:
vh
2026-10-01 04:20:17 -07:00
parent 03826d2029
commit 6b66207b0c
2 changed files with 2 additions and 116 deletions
+2 -3
View File
@@ -7,7 +7,7 @@ distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; superseded in production by 0002 on top of it; **no upstream PR (Prime, 2026-10-01)**. Both patches are carried locally only |
| `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
@@ -15,8 +15,7 @@ Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
then `sudo -n docker compose up -d scriberr`.
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
outward-facing and needs Prime's explicit yes.
Ruling: Prime, 2026-09-30, "build the slicer". **No upstream contribution (Prime, 2026-10-01: "no upstream for scriberr").** The prepared PR text was removed (it remains in git history, ee3db68).
## 0001 — pause-aware Parakeet slicer
-113
View File
@@ -1,113 +0,0 @@
# Upstream PR — prepared, NOT opened
**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit
yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub.
- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one
commit authored by Vuong Hoang against upstream `a353078` (current HEAD,
2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
- **To open it (after the yes):** fork on GitHub, then
`git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 &&
git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`,
and open the PR with the title and body below.
- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream
version, or drop it for a smaller diff (about 40 lines with its tests). It measured
neutral; see the body's last paragraph.
---
## Title
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
## Body
### Problem
For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py`
cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a
mark is chopped in two, lost, or transcribed twice. We measured it on four
recordings (118 minutes, see below): **52 % of the cuts had a transcription error
within ±3 s of them, against 19 % at points midway between cuts.**
Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the
knob is for on smaller cards, makes this worse, because there are more cuts.
### Change
One Python file, plus a new test file. No Go changes; the CLI and JSON that
`parakeet_adapter.go` reads are unchanged.
1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on
each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk
gets longer and peak GPU memory is unchanged.
2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the
word nearest the cut that both transcribed alike: same text after lowercasing
and stripping punctuation, start times within 0.5 s. The left chunk keeps the
words before it and the right chunk the rest, with the anchor taken from
whichever chunk keeps word starts in time order. With no agreed word, both
split at the cut. Segments are trimmed to match, and `transcription` is the
stitched words joined by spaces (which is what it already equals for Parakeet).
The obvious rule, "keep each word from the chunk whose half its start time falls
in", is not enough. A word that follows a pause can be timestamped anywhere in
the pause, so the two chunks often put the *same* word on opposite sides of the
cut, one frame apart, and it comes out twice (or not at all).
3. `--pause-search N` (opt-in, off by default): move each cut back to the
quietest 0.3 s within the last N seconds before the limit.
4. NeMo is imported inside `transcribe_buffered()` so the pure helpers
(`plan_slices`, `stitch_slices`) can be unit-tested without a GPU.
New flags are optional with defaults, because the Go side does not pass them. The
JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0`
reproduces the previous output exactly (words, segments and text were
byte-identical on all four test files).
### Measurements
Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3,
16 kHz mono input. Each run was scored against a **no-cut reference**: the same
model over the whole file in one pass with local attention (`parakeet_transcribe.py
--context-left 255 --context-right 255`). Words were aligned after lowercasing and
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
and the same test at points midway between cuts gives the background rate. Each
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
different places. Decoding is deterministic, so repeats are identical.
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451,
public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar
Wilde* (public domain), and two private conversational recordings (22 and 35 min;
numbers only).
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|---|---|---|---|---|
| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 |
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 |
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
The 2-standard-error band on a difference of these rates is about ±0.08, so the
last three rows are statistically tied and all beat the current slicer by a wide
margin. The public files alone tell the same story: current 52/86 damaged cuts;
this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is
unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`).
Pause-aware cutting is included but off by default: once the stitch was right, it
did not measurably help, and without an overlap it drops words just before its
cuts. Happy to drop it from this PR if you prefer the smaller diff.
### Tests
```
cd internal/transcription/adapters/py/nvidia
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
```
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
a pause at the very start or end, audio with no pause, the chunk limit with the
overlap included, and stitching with known word lists, including a word the two
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
match at a different time, an anchor that would reverse time order, and
punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py`
still passes, since its `--chunk-len 10` run now also exercises the overlap.