scriberr: no upstream contribution (Prime) — drop prepared PR text; patches carried locally
This commit is contained in:
@@ -7,7 +7,7 @@ distinctly tagged image, and proves it before anyone deploys it.
|
||||
|
||||
| patch | against | status |
|
||||
|---|---|---|
|
||||
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
|
||||
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; superseded in production by 0002 on top of it; **no upstream PR (Prime, 2026-10-01)**. Both patches are carried locally only |
|
||||
| `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
|
||||
|
||||
|
||||
@@ -15,8 +15,7 @@ Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
|
||||
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
|
||||
then `sudo -n docker compose up -d scriberr`.
|
||||
|
||||
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
|
||||
outward-facing and needs Prime's explicit yes.
|
||||
Ruling: Prime, 2026-09-30, "build the slicer". **No upstream contribution (Prime, 2026-10-01: "no upstream for scriberr").** The prepared PR text was removed (it remains in git history, ee3db68).
|
||||
|
||||
## 0001 — pause-aware Parakeet slicer
|
||||
|
||||
|
||||
@@ -1,113 +0,0 @@
|
||||
# Upstream PR — prepared, NOT opened
|
||||
|
||||
**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit
|
||||
yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub.
|
||||
|
||||
- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one
|
||||
commit authored by Vuong Hoang against upstream `a353078` (current HEAD,
|
||||
2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
|
||||
- **To open it (after the yes):** fork on GitHub, then
|
||||
`git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 &&
|
||||
git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`,
|
||||
and open the PR with the title and body below.
|
||||
- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream
|
||||
version, or drop it for a smaller diff (about 40 lines with its tests). It measured
|
||||
neutral; see the body's last paragraph.
|
||||
|
||||
---
|
||||
|
||||
## Title
|
||||
|
||||
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
|
||||
|
||||
## Body
|
||||
|
||||
### Problem
|
||||
|
||||
For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py`
|
||||
cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a
|
||||
mark is chopped in two, lost, or transcribed twice. We measured it on four
|
||||
recordings (118 minutes, see below): **52 % of the cuts had a transcription error
|
||||
within ±3 s of them, against 19 % at points midway between cuts.**
|
||||
|
||||
Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the
|
||||
knob is for on smaller cards, makes this worse, because there are more cuts.
|
||||
|
||||
### Change
|
||||
|
||||
One Python file, plus a new test file. No Go changes; the CLI and JSON that
|
||||
`parakeet_adapter.go` reads are unchanged.
|
||||
|
||||
1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on
|
||||
each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk
|
||||
gets longer and peak GPU memory is unchanged.
|
||||
2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the
|
||||
word nearest the cut that both transcribed alike: same text after lowercasing
|
||||
and stripping punctuation, start times within 0.5 s. The left chunk keeps the
|
||||
words before it and the right chunk the rest, with the anchor taken from
|
||||
whichever chunk keeps word starts in time order. With no agreed word, both
|
||||
split at the cut. Segments are trimmed to match, and `transcription` is the
|
||||
stitched words joined by spaces (which is what it already equals for Parakeet).
|
||||
|
||||
The obvious rule, "keep each word from the chunk whose half its start time falls
|
||||
in", is not enough. A word that follows a pause can be timestamped anywhere in
|
||||
the pause, so the two chunks often put the *same* word on opposite sides of the
|
||||
cut, one frame apart, and it comes out twice (or not at all).
|
||||
3. `--pause-search N` (opt-in, off by default): move each cut back to the
|
||||
quietest 0.3 s within the last N seconds before the limit.
|
||||
4. NeMo is imported inside `transcribe_buffered()` so the pure helpers
|
||||
(`plan_slices`, `stitch_slices`) can be unit-tested without a GPU.
|
||||
|
||||
New flags are optional with defaults, because the Go side does not pass them. The
|
||||
JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0`
|
||||
reproduces the previous output exactly (words, segments and text were
|
||||
byte-identical on all four test files).
|
||||
|
||||
### Measurements
|
||||
|
||||
Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3,
|
||||
16 kHz mono input. Each run was scored against a **no-cut reference**: the same
|
||||
model over the whole file in one pass with local attention (`parakeet_transcribe.py
|
||||
--context-left 255 --context-right 255`). Words were aligned after lowercasing and
|
||||
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
|
||||
and the same test at points midway between cuts gives the background rate. Each
|
||||
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
|
||||
different places. Decoding is deterministic, so repeats are identical.
|
||||
|
||||
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451,
|
||||
public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar
|
||||
Wilde* (public domain), and two private conversational recordings (22 and 35 min;
|
||||
numbers only).
|
||||
|
||||
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|
||||
|---|---|---|---|---|
|
||||
| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 |
|
||||
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
|
||||
| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 |
|
||||
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
|
||||
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
|
||||
|
||||
The 2-standard-error band on a difference of these rates is about ±0.08, so the
|
||||
last three rows are statistically tied and all beat the current slicer by a wide
|
||||
margin. The public files alone tell the same story: current 52/86 damaged cuts;
|
||||
this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is
|
||||
unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`).
|
||||
|
||||
Pause-aware cutting is included but off by default: once the stitch was right, it
|
||||
did not measurably help, and without an overlap it drops words just before its
|
||||
cuts. Happy to drop it from this PR if you prefer the smaller diff.
|
||||
|
||||
### Tests
|
||||
|
||||
```
|
||||
cd internal/transcription/adapters/py/nvidia
|
||||
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
|
||||
```
|
||||
|
||||
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
|
||||
a pause at the very start or end, audio with no pause, the chunk limit with the
|
||||
overlap included, and stitching with known word lists, including a word the two
|
||||
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
|
||||
match at a different time, an anchor that would reverse time order, and
|
||||
punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py`
|
||||
still passes, since its `--chunk-len 10` run now also exercises the overlap.
|
||||
Reference in New Issue
Block a user