Files
esh-pfi-infrastructure/stacks/scriberr/patches/README.md
T
vh ee3db68db1 feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
2026-09-30 12:10:32 -07:00

163 lines
8.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Scriberr local patches — contract
We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so
we can carry patches on that build. This directory holds them, and
`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a
distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) |
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
outward-facing and needs Prime's explicit yes.
## 0001 — pause-aware Parakeet slicer
### What it changes
One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change.
Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a
word that straddles a mark is chopped in two, lost, or transcribed twice. The
patch:
1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on
each side of the cut, counted *inside* `--chunk-len`.
2. **Hands over at an agreed word.** In each overlap, the chunks switch at the
word nearest the cut that both transcribed alike: the same text after
lowercasing and stripping punctuation (punctuation alone never counts), with
start times within 0.5 s. The left chunk keeps the words before it and the
right chunk keeps the rest, the anchor taken from whichever chunk keeps the
words in time order. With no agreed word, both split at the cut by start time.
Segments are trimmed to the words their chunk keeps, and `transcription` is
the stitched words joined by spaces (upstream's text already equals that).
3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut
moves back to the middle of the quietest 0.3 s within the last N seconds
before the limit. It measured neutral once the stitch was right, so it is not
the default; Go never passes the flag.
4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers
(`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo.
`--overlap 0` (with pause search off, the default) reproduces upstream's output
exactly: words, segments and text were byte-identical on all four test recordings.
Why the handover is by agreed word and not simply "each word goes to the chunk
its start time falls in" (the first design): at a quarter of the stitches the
two chunks put the *same* word on opposite sides of the cut, one frame apart,
so it was kept twice. Parakeet timestamps a word that follows a pause anywhere
inside the pause. Details in the bench doc.
### The seam it must keep (Go ↔ Python)
`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen:
- **Invocation** (Go, `buildBufferedArgs`):
`uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>`.
Go never passes `--overlap` or `--pause-search`, so **their defaults are the
shipped behaviour**. New flags must stay optional.
- **JSON** (Go, `parseResult`): `transcription` (str), `language` (str),
`word_timestamps` [{`word` str, `start_offset` **int**, `end_offset` **int**,
`start` float, `end` float}], `segment_timestamps` (same, with `segment`),
`audio_file`, `model`, `buffered`, `chunk_duration_secs`, `num_chunks`. An int
field that becomes a float fails the Go unmarshal. Extra keys are fine; the
patch adds `overlap_secs`, `pause_search_secs` and `cut_times`.
- `start_offset`/`end_offset` stay chunk-relative frame indices, as upstream
leaves them; Go does not read them.
- `scripts/scriberr-seam-check.py` asserts all of the above plus the stitch
invariants (word starts never go backwards; segments tile the words once).
### The memory bound
Scriberr shares fv-ml1 GPU 1 with intern-decision. **Parakeet's per-process
peak must stay ≤ 5,496 MiB** (nvidia-smi used_memory, sampled every 0.2 s), the
measured peak of upstream's 120 s slicer with `expandable_segments`. The patch
keeps it because the overlap counts **inside** `--chunk-len`: consecutive cuts
are at most `chunk-len − overlap` apart, so no chunk ever exceeds `--chunk-len`
(120 s in our compose). The pause search runs on the CPU copy of the waveform.
### Measured (2026-09-30, `docs/pfi/scriberr-slicer-bench-2026-09-30.md`)
Four recordings, 118 min in total: Prime's two uploads (private, metrics only),
the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic
reading. Each was scored against a no-cut whole-file transcript, with every
variant at three cut placements.
| | upstream (fixed 120 s) | **patch (overlap 4 s, agreed-word handover)** |
|---|---|---|
| cuts with an error within ±3 s | 52 % (93/179) | **22 % (41/184)** |
| background: same test midway between cuts | 19 % | 18 % |
| near-cut error events | 102 | 47 |
| words duplicated at cuts | 17 | 2 |
| GPU peak, 35-min file, n=3 | 5,496 MiB | **5,496 MiB** (budget 5,496) |
Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with
the handover, and pause + overlap at 4 or 8 s all land inside that floor of
each other; the default is the simplest of them.
⚠ **Separate finding, not fixed by this patch:** Parakeet sometimes skips a run
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc.
### Upgrading upstream
```bash
# on nh3-dev, from this repo
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1
```
It clones that sha into a new `/opt/docker/src/scriberr-<sha7>-<suffix>`,
`git apply --check`s each patch (a conflict stops the run and names the patch),
builds `scriberr:local-blackwell-<sha7>-<suffix>` without touching older tags,
then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table.
Memory runs on GPU 3 and refuses a GPU that is not idle. A conflict means the
patch needs rebasing: in a checkout of the new sha, `git am -3` the old patch,
resolve, run the unit tests, then `git format-patch -1 --stdout >
0001-parakeet-pause-aware-slicer.patch` and re-run the rebuild.
Also re-measure after an upgrade that touches NeMo, torch or the slicer
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md` has the harness).
**Disk.** Everything a rebuild writes lands on fv-ml1's root pool (zroot), not
`/tank`: the build dir under `/opt/docker/src` (~125 MB) and the image, which
shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its
own (measured 2026-09-30). The script refuses to build with less than 20 GB free
under Docker's root dir. Once a deploy has soaked, remove superseded builds by
their literal names: `docker rmi scriberr:local-blackwell-<sha7>-<suffix>` and
`sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>`, keeping the running
tag and the one before it for rollback.
⚠ `Dockerfile.cuda.12.9` installs the **latest** `uv`, `yt-dlp` and `deno` at
build time, so a rebuild changes those too, not just our patch. The seam stage
is what catches a `uv run` behaviour change.
### Deploy (manual; the rebuild script never does this)
`SCRIBERR_IMAGE` in `/opt/docker/compose/scriberr/.env` on fv-ml1 selects the
image (`.env.example` documents it). That `.env` is lkraven-owned mode 600,
kept out of git.
```bash
ssh infra-ops@10.251.50.54
cd /opt/docker/compose/scriberr
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat # backup first
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat # current value
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job" # must be 0
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat
```
Then prove the embed path live: the env's rewritten copy must match the patch.
```bash
docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
```
Record the deploy with `scripts/ops-log record`.
**Rollback:** set `SCRIBERR_IMAGE` back to the previous tag (or restore the
`.env` backup) and `docker compose up -d scriberr`. The old image is never
deleted by the rebuild.