Files
esh-pfi-infrastructure/stacks/scriberr/patches/README.md
T
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00

10 KiB
Raw Blame History

Scriberr local patches — contract

We build Scriberr from source (no upstream sm_120 image; see ../README.md), so we can carry patches on that build. This directory holds them, and scripts/scriberr-rebuild applies them to a pinned upstream sha, builds a distinctly tagged image, and proves it before anyone deploys it.

patch against status
0001-parakeet-pause-aware-slicer.patch upstream a353078 (HEAD 2026-09-20) LIVE on fv-ml1 since 2026-09-30 1211 PT as scriberr:local-blackwell-a353078-slicer1; upstream PR prepared, not opened (upstream-pr/)
0002-parakeet-model-path-and-gap-retry.patch 0001 LIVE on fv-ml1 since 2026-09-30 1602 PT as scriberr:local-blackwell-a353078-dropout2 (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, retried_gaps: 1, +29 words against the slicer1 image. Why: docs/pfi/parakeet-dropout-investigation-2026-09-30.md

Rollback for the live deploy: SCRIBERR_IMAGE=scriberr:local-blackwell (the unpatched image, kept), or restore /opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1, then sudo -n docker compose up -d scriberr.

Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is outward-facing and needs Prime's explicit yes.

0001 — pause-aware Parakeet slicer

What it changes

One file of product code, internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py, plus one new test file beside it (tests/test_parakeet_slicing.py). No Go change.

Upstream cuts long audio at fixed --chunk-len marks with no overlap, so a word that straddles a mark is chopped in two, lost, or transcribed twice. The patch:

  1. Overlaps adjacent chunks by --overlap seconds (default 4), half on each side of the cut, counted inside --chunk-len.
  2. Hands over at an agreed word. In each overlap, the chunks switch at the word nearest the cut that both transcribed alike: the same text after lowercasing and stripping punctuation (punctuation alone never counts), with start times within 0.5 s. The left chunk keeps the words before it and the right chunk keeps the rest, the anchor taken from whichever chunk keeps the words in time order. With no agreed word, both split at the cut by start time. Segments are trimmed to the words their chunk keeps, and transcription is the stitched words joined by spaces (upstream's text already equals that).
  3. Optional pause-aware cuts (--pause-search N, default off): each cut moves back to the middle of the quietest 0.3 s within the last N seconds before the limit. It measured neutral once the stitch was right, so it is not the default; Go never passes the flag.
  4. Imports NeMo inside transcribe_buffered() so the pure helpers (plan_slices, stitch_slices) import and test without a GPU or NeMo.

--overlap 0 (with pause search off, the default) reproduces upstream's output exactly: words, segments and text were byte-identical on all four test recordings.

Why the handover is by agreed word and not simply "each word goes to the chunk its start time falls in" (the first design): at a quarter of the stitches the two chunks put the same word on opposite sides of the cut, one frame apart, so it was kept twice. Parakeet timestamps a word that follows a pause anywhere inside the pause. Details in the bench doc.

The seam it must keep (Go ↔ Python)

parakeet_adapter.go is not patched, so the script's CLI and JSON are frozen:

  • Invocation (Go, buildBufferedArgs): uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>. Go never passes --overlap or --pause-search, so their defaults are the shipped behaviour. New flags must stay optional.
  • JSON (Go, parseResult): transcription (str), language (str), word_timestamps [{word str, start_offset int, end_offset int, start float, end float}], segment_timestamps (same, with segment), audio_file, model, buffered, chunk_duration_secs, num_chunks. An int field that becomes a float fails the Go unmarshal. Extra keys are fine; the patch adds overlap_secs, pause_search_secs and cut_times.
  • start_offset/end_offset stay chunk-relative frame indices, as upstream leaves them; Go does not read them.
  • scripts/scriberr-seam-check.py asserts all of the above plus the stitch invariants (word starts never go backwards; segments tile the words once).

The memory bound

Scriberr shares fv-ml1 GPU 1 with intern-decision. Parakeet's per-process peak must stay ≤ 5,496 MiB (nvidia-smi used_memory, sampled every 0.2 s), the measured peak of upstream's 120 s slicer with expandable_segments. The patch keeps it because the overlap counts inside --chunk-len: consecutive cuts are at most chunk-len − overlap apart, so no chunk ever exceeds --chunk-len (120 s in our compose). The pause search runs on the CPU copy of the waveform.

Measured (2026-09-30, docs/pfi/scriberr-slicer-bench-2026-09-30.md)

Four recordings, 118 min in total: Prime's two uploads (private, metrics only), the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic reading. Each was scored against a no-cut whole-file transcript, with every variant at three cut placements.

upstream (fixed 120 s) patch (overlap 4 s, agreed-word handover)
cuts with an error within ±3 s 52 % (93/179) 22 % (41/184)
background: same test midway between cuts 19 % 18 %
near-cut error events 102 47
words duplicated at cuts 17 2
GPU peak, 35-min file, n=3 5,496 MiB 5,496 MiB (budget 5,496)

Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with the handover, and pause + overlap at 4 or 8 s all land inside that floor of each other; the default is the simplest of them.

⚠ Separate finding, not fixed by this patch: Parakeet sometimes skips a run of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12 transcripts, for upstream's slicer too). See the bench doc.

0002 (live since 2026-09-30 1602) — gap retry and an explicit model path

Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the audio holds speech but no word came out (--retry-gaps, default 3; 0 off), which cut those losses 80–90 % on four recordings. It also adds PARAKEET_MODEL_PATH (the .nemo to load; default unchanged), reports the loaded model in the JSON, and makes the Go adapter record it as ModelUsed. It is carried (in this directory) since 2026-09-30, and scripts/scriberr-rebuild applies 0001+0002 by default (suffix dropout2, memory budget 5,600 MiB as a regression guard). Rollback: SCRIBERR_IMAGE=scriberr:local-blackwell-a353078-slicer1 (.env.bak-20260930-pre-dropout2). To test-build a future proposed patch: --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed.

Upgrading upstream

# on nh3-dev, from this repo
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1

It clones that sha into a new /opt/docker/src/scriberr-<sha7>-<suffix>, git apply --checks each patch (a conflict stops the run and names the patch), builds scriberr:local-blackwell-<sha7>-<suffix> without touching older tags, then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table. Memory runs on GPU 3, which Scriberr itself now occupies (2026-09-30); the stage needs ≥ 20 GB free and counts only its own container's PIDs, so a Scriberr job on the card neither blocks nor pollutes it. A conflict means the patch needs rebasing: in a checkout of the new sha, git am -3 the old patch, resolve, run the unit tests, then git format-patch -1 --stdout > 0001-parakeet-pause-aware-slicer.patch and re-run the rebuild.

Also re-measure after an upgrade that touches NeMo, torch or the slicer (docs/pfi/scriberr-slicer-bench-2026-09-30.md has the harness).

Disk. Everything a rebuild writes lands on fv-ml1's root pool (zroot), not /tank: the build dir under /opt/docker/src (~125 MB) and the image, which shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its own (measured 2026-09-30). The script refuses to build with less than 20 GB free under Docker's root dir. Once a deploy has soaked, remove superseded builds by their literal names: docker rmi scriberr:local-blackwell-<sha7>-<suffix> and sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>, keeping the running tag and the one before it for rollback.

⚠ Dockerfile.cuda.12.9 installs the latest uv, yt-dlp and deno at build time, so a rebuild changes those too, not just our patch. The seam stage is what catches a uv run behaviour change.

Deploy (manual; the rebuild script never does this)

SCRIBERR_IMAGE in /opt/docker/compose/scriberr/.env on fv-ml1 selects the image (.env.example documents it). That .env is lkraven-owned mode 600, kept out of git.

ssh infra-ops@10.251.50.54
cd /opt/docker/compose/scriberr
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat        # backup first
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat                   # current value
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job"   # must be 0
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat

Then prove the embed path live: the env's rewritten copy must match the patch.

docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py

Record the deploy with scripts/ops-log record.

Rollback: set SCRIBERR_IMAGE back to the previous tag (or restore the .env backup) and docker compose up -d scriberr. The old image is never deleted by the rebuild.