From af450f4ef725903519f294c65d262b5f5c4c5722 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Wed, 30 Sep 2026 15:41:55 -0700 Subject: [PATCH] docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up --- persistent-memory.md | 5 +++-- stacks/intern-decision/README.md | 12 ++++++++++-- 2 files changed, 13 insertions(+), 4 deletions(-) diff --git a/persistent-memory.md b/persistent-memory.md index 2f3eeff..ef10ede 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -181,12 +181,13 @@ _As of 2026-09-30 ~0120 PT._ - **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness. - **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha `. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. - **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`). - - **OPEN, not fixed:** Parakeet skips runs of ≥10 words mid-chunk with ANY slicer, upstream's included (12–17 runs, 500–720 words per 12 transcripts; p2 lost 85 words at today's old setting). It is chaotic with cut placement. The investigation (decoder, chunk length, model) is Prime's call. + - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout1`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. + - Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`). - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. - **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1. - **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs. - - ⚠ The first call in a new length bucket after a restart costs ~6.5 s (kernel autotune per shape bucket). A startup warm-up across the buckets would fix it; not done. + - ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. Fix, not done: a named volume for the cache plus a ~2 min warm-up after an image change. - **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).** - Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`. - Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16. diff --git a/stacks/intern-decision/README.md b/stacks/intern-decision/README.md index e7c05b5..cbf6ac3 100644 --- a/stacks/intern-decision/README.md +++ b/stacks/intern-decision/README.md @@ -104,8 +104,16 @@ probability diffs against the bench's r1..r4. `/decide` is unchanged. ⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra (3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape -bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would -remove it; that is not done yet. +bucket. Warm calls are as tabled. + - **How long the warm state lasts (measured 1541):** it lives in the Triton cache at `/tmp/triton-cache`, including + fla's `*.autotune.json`, in the container's writable layer. It SURVIVES `docker restart` (and so host reboots): + after a restart, a seen 32k bucket answered in 2.0 s. It is LOST on recreate (`compose up -d` after a change, a + deploy, an image upgrade). An unseen bucket after the restart still cost +6.4 s (12,187 tokens), which is the + positive control. + - **Buckets:** every observation so far fits buckets of 2,048 tokens (inferred from ~20 calls, not proven), so ~16 + buckets up to 32,768. A full warm-up is estimated at ~2 min, once per container. + - **Durable fix (not done):** put `/tmp/triton-cache` on a named volume so that recreates keep it, and run one + ~2 min warm-up across the buckets after an image change (a new triton/fla version keys new entries). The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.