Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
181 lines
10 KiB
Markdown
181 lines
10 KiB
Markdown
# Scriberr local patches — contract
|
||
|
||
We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so
|
||
we can carry patches on that build. This directory holds them, and
|
||
`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a
|
||
distinctly tagged image, and proves it before anyone deploys it.
|
||
|
||
| patch | against | status |
|
||
|---|---|---|
|
||
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
|
||
| `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
|
||
|
||
|
||
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
|
||
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
|
||
then `sudo -n docker compose up -d scriberr`.
|
||
|
||
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
|
||
outward-facing and needs Prime's explicit yes.
|
||
|
||
## 0001 — pause-aware Parakeet slicer
|
||
|
||
### What it changes
|
||
|
||
One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
|
||
plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change.
|
||
|
||
Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a
|
||
word that straddles a mark is chopped in two, lost, or transcribed twice. The
|
||
patch:
|
||
|
||
1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on
|
||
each side of the cut, counted *inside* `--chunk-len`.
|
||
2. **Hands over at an agreed word.** In each overlap, the chunks switch at the
|
||
word nearest the cut that both transcribed alike: the same text after
|
||
lowercasing and stripping punctuation (punctuation alone never counts), with
|
||
start times within 0.5 s. The left chunk keeps the words before it and the
|
||
right chunk keeps the rest, the anchor taken from whichever chunk keeps the
|
||
words in time order. With no agreed word, both split at the cut by start time.
|
||
Segments are trimmed to the words their chunk keeps, and `transcription` is
|
||
the stitched words joined by spaces (upstream's text already equals that).
|
||
3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut
|
||
moves back to the middle of the quietest 0.3 s within the last N seconds
|
||
before the limit. It measured neutral once the stitch was right, so it is not
|
||
the default; Go never passes the flag.
|
||
4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers
|
||
(`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo.
|
||
|
||
`--overlap 0` (with pause search off, the default) reproduces upstream's output
|
||
exactly: words, segments and text were byte-identical on all four test recordings.
|
||
|
||
Why the handover is by agreed word and not simply "each word goes to the chunk
|
||
its start time falls in" (the first design): at a quarter of the stitches the
|
||
two chunks put the *same* word on opposite sides of the cut, one frame apart,
|
||
so it was kept twice. Parakeet timestamps a word that follows a pause anywhere
|
||
inside the pause. Details in the bench doc.
|
||
|
||
### The seam it must keep (Go ↔ Python)
|
||
|
||
`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen:
|
||
|
||
- **Invocation** (Go, `buildBufferedArgs`):
|
||
`uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>`.
|
||
Go never passes `--overlap` or `--pause-search`, so **their defaults are the
|
||
shipped behaviour**. New flags must stay optional.
|
||
- **JSON** (Go, `parseResult`): `transcription` (str), `language` (str),
|
||
`word_timestamps` [{`word` str, `start_offset` **int**, `end_offset` **int**,
|
||
`start` float, `end` float}], `segment_timestamps` (same, with `segment`),
|
||
`audio_file`, `model`, `buffered`, `chunk_duration_secs`, `num_chunks`. An int
|
||
field that becomes a float fails the Go unmarshal. Extra keys are fine; the
|
||
patch adds `overlap_secs`, `pause_search_secs` and `cut_times`.
|
||
- `start_offset`/`end_offset` stay chunk-relative frame indices, as upstream
|
||
leaves them; Go does not read them.
|
||
- `scripts/scriberr-seam-check.py` asserts all of the above plus the stitch
|
||
invariants (word starts never go backwards; segments tile the words once).
|
||
|
||
### The memory bound
|
||
|
||
Scriberr shares fv-ml1 GPU 1 with intern-decision. **Parakeet's per-process
|
||
peak must stay ≤ 5,496 MiB** (nvidia-smi used_memory, sampled every 0.2 s), the
|
||
measured peak of upstream's 120 s slicer with `expandable_segments`. The patch
|
||
keeps it because the overlap counts **inside** `--chunk-len`: consecutive cuts
|
||
are at most `chunk-len − overlap` apart, so no chunk ever exceeds `--chunk-len`
|
||
(120 s in our compose). The pause search runs on the CPU copy of the waveform.
|
||
|
||
### Measured (2026-09-30, `docs/pfi/scriberr-slicer-bench-2026-09-30.md`)
|
||
|
||
Four recordings, 118 min in total: Prime's two uploads (private, metrics only),
|
||
the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic
|
||
reading. Each was scored against a no-cut whole-file transcript, with every
|
||
variant at three cut placements.
|
||
|
||
| | upstream (fixed 120 s) | **patch (overlap 4 s, agreed-word handover)** |
|
||
|---|---|---|
|
||
| cuts with an error within ±3 s | 52 % (93/179) | **22 % (41/184)** |
|
||
| background: same test midway between cuts | 19 % | 18 % |
|
||
| near-cut error events | 102 | 47 |
|
||
| words duplicated at cuts | 17 | 2 |
|
||
| GPU peak, 35-min file, n=3 | 5,496 MiB | **5,496 MiB** (budget 5,496) |
|
||
|
||
Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with
|
||
the handover, and pause + overlap at 4 or 8 s all land inside that floor of
|
||
each other; the default is the simplest of them.
|
||
|
||
⚠ **Separate finding, not fixed by this patch:** Parakeet sometimes skips a run
|
||
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
|
||
transcripts, for upstream's slicer too). See the bench doc.
|
||
|
||
### 0002 (live since 2026-09-30 1602) — gap retry and an explicit model path
|
||
|
||
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
|
||
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
|
||
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
|
||
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
|
||
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
|
||
the Go adapter record it as `ModelUsed`. It is carried (in this directory) since 2026-09-30, and
|
||
`scripts/scriberr-rebuild` applies 0001+0002 by default (suffix `dropout2`, memory budget 5,600 MiB
|
||
as a regression guard). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell-a353078-slicer1`
|
||
(`.env.bak-20260930-pre-dropout2`). To test-build a future proposed patch: `--patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
|
||
|
||
### Upgrading upstream
|
||
|
||
```bash
|
||
# on nh3-dev, from this repo
|
||
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1
|
||
```
|
||
|
||
It clones that sha into a new `/opt/docker/src/scriberr-<sha7>-<suffix>`,
|
||
`git apply --check`s each patch (a conflict stops the run and names the patch),
|
||
builds `scriberr:local-blackwell-<sha7>-<suffix>` without touching older tags,
|
||
then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table.
|
||
Memory runs on GPU 3, which Scriberr itself now occupies (2026-09-30); the stage needs ≥ 20 GB free and counts only its own container's PIDs, so a Scriberr job on the card neither blocks nor pollutes it. A conflict means the
|
||
patch needs rebasing: in a checkout of the new sha, `git am -3` the old patch,
|
||
resolve, run the unit tests, then `git format-patch -1 --stdout >
|
||
0001-parakeet-pause-aware-slicer.patch` and re-run the rebuild.
|
||
|
||
Also re-measure after an upgrade that touches NeMo, torch or the slicer
|
||
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md` has the harness).
|
||
|
||
**Disk.** Everything a rebuild writes lands on fv-ml1's root pool (zroot), not
|
||
`/tank`: the build dir under `/opt/docker/src` (~125 MB) and the image, which
|
||
shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its
|
||
own (measured 2026-09-30). The script refuses to build with less than 20 GB free
|
||
under Docker's root dir. Once a deploy has soaked, remove superseded builds by
|
||
their literal names: `docker rmi scriberr:local-blackwell-<sha7>-<suffix>` and
|
||
`sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>`, keeping the running
|
||
tag and the one before it for rollback.
|
||
|
||
⚠ `Dockerfile.cuda.12.9` installs the **latest** `uv`, `yt-dlp` and `deno` at
|
||
build time, so a rebuild changes those too, not just our patch. The seam stage
|
||
is what catches a `uv run` behaviour change.
|
||
|
||
### Deploy (manual; the rebuild script never does this)
|
||
|
||
`SCRIBERR_IMAGE` in `/opt/docker/compose/scriberr/.env` on fv-ml1 selects the
|
||
image (`.env.example` documents it). That `.env` is lkraven-owned mode 600,
|
||
kept out of git.
|
||
|
||
```bash
|
||
ssh infra-ops@10.251.50.54
|
||
cd /opt/docker/compose/scriberr
|
||
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat # backup first
|
||
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat # current value
|
||
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
|
||
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job" # must be 0
|
||
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat
|
||
```
|
||
|
||
Then prove the embed path live: the env's rewritten copy must match the patch.
|
||
|
||
```bash
|
||
docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
|
||
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||
```
|
||
|
||
Record the deploy with `scripts/ops-log record`.
|
||
|
||
**Rollback:** set `SCRIBERR_IMAGE` back to the previous tag (or restore the
|
||
`.env` backup) and `docker compose up -d scriberr`. The old image is never
|
||
deleted by the rebuild.
|