memory: warm-up cache audit passed; parakeet seat A/B in flight
This commit is contained in:
@@ -184,11 +184,12 @@ _As of 2026-09-30 ~0120 PT._
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
|
||||
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
|
||||
- **Parakeet SEAT A/B IN FLIGHT (Prime 1558: "a/b the one on fv-ml1's general seat against the unified new one"; latency is load-bearing, even 50 ms).** A background agent is comparing the seat (`parakeet`, GPU 0, sherpa-onnx int8 v3, :8300, LiteLLM `ext-stt`/`whisper-1`) with `nvidia/parakeet-unified-en-0.6b`. It runs under NeMo, AND in the seat's own ONNX runtime if an export exists or can be made, with v2 int8 as a cheap English control. Candidates run on GPU 3 only. ⚠ GPU 0 has ~101 MiB free, so an in-place replacement must fit about 1,790 MiB. No deploy.
|
||||
- **Scriberr gap-retry fix (patch 0002) LIVE 1602 2026-09-30** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "basic fix, no surgery for the new toolkit"; v3 stays). The investigation (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`) found the v3 drops are real against ground truth (50–140 words per transcript); the retry cuts them 80–90%. Live check: 5,502 MiB, `retried_gaps` reported. Rollback: `.env.bak-20260930-pre-dropout2` / slicer1. Leftovers kept pending Prime: 26 GB of candidate weights, `envs/nemo300`, and the private bench data in `/tank/spikes/scriberr-slicer/`.
|
||||
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
|
||||
- **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.
|
||||
- **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs.
|
||||
- ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. **Prime 1545: build it. TASKED to infra-hermes (a named volume for the cache plus `scripts/intern-decision-warmup`, proving the bucket width with random sizes); infra-ops AUDITS when it lands.**
|
||||
- ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. **DONE as 0.1.3 (infra-hermes f65e27b), and my AUDIT PASSED 1604:** the named volume `intern-decision_triton-cache` survives a force-recreate (6 random sizes from 5.5k to 32.6k all warm), and the 16 × 2,048-token buckets are PROVEN. Run `scripts/intern-decision-warmup` after an IMAGE change only; cold it takes 109 s, warm 17 s. JevBench still has 0 diffs.
|
||||
- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).**
|
||||
- Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`.
|
||||
- Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16.
|
||||
|
||||
Reference in New Issue
Block a user