memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state

This commit is contained in:
vh
2026-10-01 01:37:49 -07:00
parent 392660bd5f
commit cb28d8c951
+11 -3
View File
@@ -117,9 +117,17 @@ no longer deployed sidecars here. See Recent decisions.)
_As of 2026-09-30 ~1800 PT._ _As of 2026-09-30 ~1800 PT._
### NEXT: switch the Parakeet speech seat (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch") ### Parakeet speech seat switched to unified-en under NeMo (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
- **TASKED to infra-hermes at 2354 2026-09-30** (Prime corrected me: "send the switch to infra-hermes"). infra-hermes builds and deploys; **infra-ops AUDITS when it reports**, the same way as the Jev endpoint: re-run the latency bins, the WER spot-check, the defect checks and memory myself. The details below are the brief it got. - **DONE: LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, image `local/parakeet-nemo:nemo-0.1.0`, infra-hermes de6ea32 + 41d2014) on :8300; LiteLLM untouched. **infra-ops AUDIT PASSED 0137.**
- Latency on GPU 0, p50 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms before. Through LiteLLM a 2.5 s clip takes 80–98 ms.
- WER: LibriSpeech clean 1.965, other 3.026.
- Fixed: a 714–726 s file returns 200 (long files go in 360 s windows, because NeMo builds the full T×T attention mask even under local attention); no drop after a pause.
- Rollback: `docker stop parakeet-nemo && docker start parakeet` (the old container is stopped, not removed).
- **gen-small: util 0.48 → 0.36** (0.46 and 0.40 failed its boot check). Its KV is byte-pinned (`--kv-cache-memory 8 GiB`), still 670,142 tokens / 2.56×, so there was no KV cost. It was down ~34 min (0012–0046 PDT) while the util was iterated; zero LiteLLM errors.
- ⚠ **Steady state is 3,582 MiB for the seat (its cached window peak); GPU 0 Free is 385 MiB.** That leaves gen-small's restart boot-check margin at only ~0.45 GiB. infra-hermes was asked to set gen-small util to 0.33 in the .env WITHOUT restarting, so it applies at the next restart. **Nothing else fits on GPU 0.**
- Seat invariants are in its README: cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11`, because httptools 0.8.0 writes `HTTP/1.1 200\x00OK` and LiteLLM/httpx rejects it; the 360 s window.
- The original brief, for reference:
- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use. - **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use.
- The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`. - The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`.
@@ -305,7 +313,7 @@ _As of 2026-09-30 ~1800 PT._
## Recent decisions ## Recent decisions
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. Implementation was TASKED to infra-hermes (2354); infra-ops audits. Tracked by the in-flight "NEXT" section and a6c1d3c.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` - `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` - `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md` - `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
- `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` - `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`