Files
esh-pfi-infrastructure/stacks/scriberr/patches/README.md
T
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00

181 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Scriberr local patches — contract
We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so
we can carry patches on that build. This directory holds them, and
`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a
distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
then `sudo -n docker compose up -d scriberr`.
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
outward-facing and needs Prime's explicit yes.
## 0001 — pause-aware Parakeet slicer
### What it changes
One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change.
Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a
word that straddles a mark is chopped in two, lost, or transcribed twice. The
patch:
1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on
each side of the cut, counted *inside* `--chunk-len`.
2. **Hands over at an agreed word.** In each overlap, the chunks switch at the
word nearest the cut that both transcribed alike: the same text after
lowercasing and stripping punctuation (punctuation alone never counts), with
start times within 0.5 s. The left chunk keeps the words before it and the
right chunk keeps the rest, the anchor taken from whichever chunk keeps the
words in time order. With no agreed word, both split at the cut by start time.
Segments are trimmed to the words their chunk keeps, and `transcription` is
the stitched words joined by spaces (upstream's text already equals that).
3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut
moves back to the middle of the quietest 0.3 s within the last N seconds
before the limit. It measured neutral once the stitch was right, so it is not
the default; Go never passes the flag.
4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers
(`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo.
`--overlap 0` (with pause search off, the default) reproduces upstream's output
exactly: words, segments and text were byte-identical on all four test recordings.
Why the handover is by agreed word and not simply "each word goes to the chunk
its start time falls in" (the first design): at a quarter of the stitches the
two chunks put the *same* word on opposite sides of the cut, one frame apart,
so it was kept twice. Parakeet timestamps a word that follows a pause anywhere
inside the pause. Details in the bench doc.
### The seam it must keep (Go ↔ Python)
`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen:
- **Invocation** (Go, `buildBufferedArgs`):
`uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>`.
Go never passes `--overlap` or `--pause-search`, so **their defaults are the
shipped behaviour**. New flags must stay optional.
- **JSON** (Go, `parseResult`): `transcription` (str), `language` (str),
`word_timestamps` [{`word` str, `start_offset` **int**, `end_offset` **int**,
`start` float, `end` float}], `segment_timestamps` (same, with `segment`),
`audio_file`, `model`, `buffered`, `chunk_duration_secs`, `num_chunks`. An int
field that becomes a float fails the Go unmarshal. Extra keys are fine; the
patch adds `overlap_secs`, `pause_search_secs` and `cut_times`.
- `start_offset`/`end_offset` stay chunk-relative frame indices, as upstream
leaves them; Go does not read them.
- `scripts/scriberr-seam-check.py` asserts all of the above plus the stitch
invariants (word starts never go backwards; segments tile the words once).
### The memory bound
Scriberr shares fv-ml1 GPU 1 with intern-decision. **Parakeet's per-process
peak must stay ≤ 5,496 MiB** (nvidia-smi used_memory, sampled every 0.2 s), the
measured peak of upstream's 120 s slicer with `expandable_segments`. The patch
keeps it because the overlap counts **inside** `--chunk-len`: consecutive cuts
are at most `chunk-len − overlap` apart, so no chunk ever exceeds `--chunk-len`
(120 s in our compose). The pause search runs on the CPU copy of the waveform.
### Measured (2026-09-30, `docs/pfi/scriberr-slicer-bench-2026-09-30.md`)
Four recordings, 118 min in total: Prime's two uploads (private, metrics only),
the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic
reading. Each was scored against a no-cut whole-file transcript, with every
variant at three cut placements.
| | upstream (fixed 120 s) | **patch (overlap 4 s, agreed-word handover)** |
|---|---|---|
| cuts with an error within ±3 s | 52 % (93/179) | **22 % (41/184)** |
| background: same test midway between cuts | 19 % | 18 % |
| near-cut error events | 102 | 47 |
| words duplicated at cuts | 17 | 2 |
| GPU peak, 35-min file, n=3 | 5,496 MiB | **5,496 MiB** (budget 5,496) |
Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with
the handover, and pause + overlap at 4 or 8 s all land inside that floor of
each other; the default is the simplest of them.
⚠ **Separate finding, not fixed by this patch:** Parakeet sometimes skips a run
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc.
### 0002 (live since 2026-09-30 1602) — gap retry and an explicit model path
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
the Go adapter record it as `ModelUsed`. It is carried (in this directory) since 2026-09-30, and
`scripts/scriberr-rebuild` applies 0001+0002 by default (suffix `dropout2`, memory budget 5,600 MiB
as a regression guard). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell-a353078-slicer1`
(`.env.bak-20260930-pre-dropout2`). To test-build a future proposed patch: `--patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
### Upgrading upstream
```bash
# on nh3-dev, from this repo
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1
```
It clones that sha into a new `/opt/docker/src/scriberr-<sha7>-<suffix>`,
`git apply --check`s each patch (a conflict stops the run and names the patch),
builds `scriberr:local-blackwell-<sha7>-<suffix>` without touching older tags,
then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table.
Memory runs on GPU 3, which Scriberr itself now occupies (2026-09-30); the stage needs ≥ 20 GB free and counts only its own container's PIDs, so a Scriberr job on the card neither blocks nor pollutes it. A conflict means the
patch needs rebasing: in a checkout of the new sha, `git am -3` the old patch,
resolve, run the unit tests, then `git format-patch -1 --stdout >
0001-parakeet-pause-aware-slicer.patch` and re-run the rebuild.
Also re-measure after an upgrade that touches NeMo, torch or the slicer
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md` has the harness).
**Disk.** Everything a rebuild writes lands on fv-ml1's root pool (zroot), not
`/tank`: the build dir under `/opt/docker/src` (~125 MB) and the image, which
shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its
own (measured 2026-09-30). The script refuses to build with less than 20 GB free
under Docker's root dir. Once a deploy has soaked, remove superseded builds by
their literal names: `docker rmi scriberr:local-blackwell-<sha7>-<suffix>` and
`sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>`, keeping the running
tag and the one before it for rollback.
⚠ `Dockerfile.cuda.12.9` installs the **latest** `uv`, `yt-dlp` and `deno` at
build time, so a rebuild changes those too, not just our patch. The seam stage
is what catches a `uv run` behaviour change.
### Deploy (manual; the rebuild script never does this)
`SCRIBERR_IMAGE` in `/opt/docker/compose/scriberr/.env` on fv-ml1 selects the
image (`.env.example` documents it). That `.env` is lkraven-owned mode 600,
kept out of git.
```bash
ssh infra-ops@10.251.50.54
cd /opt/docker/compose/scriberr
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat # backup first
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat # current value
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job" # must be 0
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat
```
Then prove the embed path live: the env's rewritten copy must match the patch.
```bash
docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
```
Record the deploy with `scripts/ops-log record`.
**Rollback:** set `SCRIBERR_IMAGE` back to the previous tag (or restore the
`.env` backup) and `docker compose up -d scriberr`. The old image is never
deleted by the rebuild.