docs(scriberr): slicer patch live on fv-ml1 as local-blackwell-a353078-slicer1; live GPU 1 peak 5,496 MiB

Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
This commit is contained in:
vh
2026-09-30 12:14:09 -07:00
parent ee3db68db1
commit 3c5f1ea803
3 changed files with 10 additions and 6 deletions
@@ -191,6 +191,7 @@ opt-in `--pause-search`, off by default.
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
| **live, GPU 1**, deployed container, `docker exec` as appuser, 120 s, p1 | 1 | **5,496 MiB** (2026-09-30 1212 PT, beside intern-decision; 49 s for 35 min of audio; seam check OK) |
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside
it. The scotus file reads the same peak as p1, so the rebuild script uses it
+4 -5
View File
@@ -179,11 +179,10 @@ _As of 2026-09-30 ~0120 PT._
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
- **Scriberr pause-aware slicer: BUILD IN FLIGHT (Prime, 2026-09-30: "build the slicer").** A background agent owns it.
- The design: pause-aware cuts plus a small overlap with timestamp stitching, carried in `stacks/scriberr/patches/` and rebuilt by `scripts/scriberr-rebuild` (pinned sha, `git apply --check`, a distinct image tag, then a seam, memory and quality check).
- Constraints: the peak must stay ≤ 5,496 MiB (the GPU 1 budget beside intern-decision); the Go↔Python seam (CLI + JSON) is unchanged; it is judged on ≥3 recordings against a no-cut reference. Prime's transcripts stay private and off git.
- An upstream PR is PREPARED, not opened: that needs Prime's yes.
- Upstream cadence: 1 commit in 90 days, and the maintainer is restarting (a353078, 2026-09-20), so the patch burden is low.
- **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`.
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
- **OPEN, not fixed:** Parakeet skips runs of ≥10 words mid-chunk with ANY slicer, upstream's included (12–17 runs, 500–720 words per 12 transcripts; p2 lost 85 words at today's old setting). It is chaotic with cut placement. The investigation (decoder, chunk length, model) is Prime's call.
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).**
- Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`.
- Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16.
+5 -1
View File
@@ -7,7 +7,11 @@ distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
unpatched image, kept), or restore `/opt/docker/compose/scriberr/.env.bak-20260930-pre-slicer1`,
then `sudo -n docker compose up -d scriberr`.
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
outward-facing and needs Prime's explicit yes.