From 212b736836470fdad6fc4aa7f40b859152138adb Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Wed, 30 Sep 2026 18:53:04 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20parakeet=20seat=20A/B=20done=20?= =?UTF-8?q?=E2=80=94=20runtime=20is=20the=20bottleneck;=20unified-en=20NeM?= =?UTF-8?q?o=20bf16=20wins=20speed+WER;=20seat=20defects?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- persistent-memory.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/persistent-memory.md b/persistent-memory.md index 9cbe669..0be55a6 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -184,7 +184,14 @@ _As of 2026-09-30 ~0120 PT._ - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. - Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`). - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. -- **Parakeet SEAT A/B IN FLIGHT (Prime 1558: "a/b the one on fv-ml1's general seat against the unified new one"; latency is load-bearing, even 50 ms).** A background agent is comparing the seat (`parakeet`, GPU 0, sherpa-onnx int8 v3, :8300, LiteLLM `ext-stt`/`whisper-1`) with `nvidia/parakeet-unified-en-0.6b`. It runs under NeMo, AND in the seat's own ONNX runtime if an export exists or can be made, with v2 int8 as a cheap English control. Candidates run on GPU 3 only. ⚠ GPU 0 has ~101 MiB free, but **Prime 1610: room can be freed by trimming KV concurrency on a big GPU 0 seat.** The best donor is `vllm-gen-small` (util 0.48, NVFP4, 670k-token fp8 KV, 2.56× at 262k). Each 0.01 util is about 0.95 GB, or about 35k tokens (~5% of its KV); freeing 2 GB takes it to roughly 2.3×. `vllm-cyberprev` (0.40, 359k KV, 1.37×) would lose ~9% per GB. Do NOT touch `vllm-voices` (0.11). Changing util means a gen-small restart (~2–3 min). The KV figures are from each seat's boot log (09-14); the per-token cost is my estimate, and the restart log will print the real numbers. No deploy. +- **Parakeet SEAT A/B DONE 1745 2026-09-30** (`docs/pfi/parakeet-seat-ab-2026-09-30.md`, a6c1d3c). **The seat's latency is its RUNTIME, not its model:** the sherpa-onnx int8 graph runs on ONE CPU thread with the GPU at 2–9%. + - End-to-end p50 for 1–3 / 3–8 / 8–20 s clips: seat 144 / 260 / 565 ms; `parakeet-unified-en-0.6b` under NeMo with bf16 weights 23 / 27 / 33 ms. The floor is ≤ 6 ms, and a +50 ms positive control read +52. + - Unified also wins English WER everywhere: LS-clean 1.97 against 2.70, LS-other 3.09 against 4.56, AMI 8.30 against 12.69. + - Unified int8 in the seat's runtime is SLOWER, so the runtime has to change. + - **Live-seat defects:** HTTP 500 above ~400 s (the ONNX position table is fixed at 5,000 frames); long-form dropouts of 320–1,676 of 3,580 words on 6-minute files; and after a 1.5 s digital-silence pause it can drop the rest of the utterance (6 of 40; `talk` is a caller). + - **Fit:** unified bf16 needs +1.1 GB serving (+1.5 at load) over the seat's 1,690 on GPU 0. ⚠ GPU 1's 6,625 free is NOT spare: it is intern-decision's 32k headroom. + - Switch kit: NeMo 3.0.0 plus `services/parakeet-ab-2026-09-30/code/serve_nemo.py` (same endpoints and text); the image is NOT built; it needs a warm-up, a bf16 cast before moving to the GPU, and local attention for long files. Licence: NVIDIA Open Model License. The spike dir is fv-ml1 `/tank/spikes/parakeet-ab` (~25 GB). + - **Awaiting Prime:** the switch, the gen-small KV trim (~2 GB, util 0.48→0.46), the licence, and a `NUM_THREADS=16` stopgap (~40% faster). - **Scriberr gap-retry fix (patch 0002) LIVE 1602 2026-09-30** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "basic fix, no surgery for the new toolkit"; v3 stays). The investigation (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`) found the v3 drops are real against ground truth (50–140 words per transcript); the retry cuts them 80–90%. Live check: 5,502 MiB, `retried_gaps` reported. Rollback: `.env.bak-20260930-pre-dropout2` / slicer1. Leftovers kept pending Prime: 26 GB of candidate weights, `envs/nemo300`, and the private bench data in `/tank/spikes/scriberr-slicer/`. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. - **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.