diff --git a/persistent-memory.md b/persistent-memory.md index 4300fe6..0c08de6 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -119,17 +119,12 @@ _As of 2026-10-01 ~0420 PT._ ### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01) -- **⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:** - - `vllm-gen-small`'s EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache. - - vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23. - - I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB. - - **infra-hermes is tasked with nemo-0.1.1:** `empty_cache` after windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth. - - **Awaiting Prime:** trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3. - - Until fixed, a long transcription can re-grow the seat's cache and starve gen-small. -- **LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, `local/parakeet-nemo:nemo-0.1.0`, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLM `ext-stt`/`whisper-1` unchanged. **infra-ops audit PASSED 0137.** - - p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat. - - WER: LibriSpeech clean 1.965, other 3.026. - - Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word. +- **INCIDENT 04:21 PT 2026-10-01: FIXED by nemo-0.1.1 (infra-hermes 963f9ed, live ~0441, infra-ops spot audit passed).** + - Cause: gen-small hit CUDA OOM on a 394 MiB lazily allocated workspace while the seat was parked at its 3,582 MiB window cache (GPU 0 Free 388). + - The fix: `empty_cache` around every window; `MEM_CAP_MIB=3840` (a hard process ceiling); `CUDA_GRAPHS=0`, because empty_cache poisons the graph pool (an illegal memory access on the next request). + - Result: the seat rests at 2,108 MiB and peaks at 3,028 during a 12-min file. gen-small is steady at 36,116 (its runtime growth landed at init, with zero movement on later requests). GPU 0 Free is ~1.05 GB at rest. + - Cost: graphs-off is ~2–8 ms slower at short clips and ~27 ms at 20–60 s (35 / 40 / 50 / 98 ms against 33 / 36 / 42 / 71), still 4–15× faster than the old seat. + - **Lesson: a shared-card tenant must carry a HARD cap, whether or not it returns memory.** - **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept. - **GPU 0 is FULL:** - The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**. @@ -294,6 +289,7 @@ _As of 2026-10-01 ~0420 PT._ ## Recent decisions +- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them. - `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it. - `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. - `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).