memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state
This commit is contained in:
+11
-3
@@ -117,9 +117,17 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
_As of 2026-09-30 ~1800 PT._
|
||||
|
||||
### NEXT: switch the Parakeet speech seat (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
|
||||
### Parakeet speech seat switched to unified-en under NeMo (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
|
||||
|
||||
- **TASKED to infra-hermes at 2354 2026-09-30** (Prime corrected me: "send the switch to infra-hermes"). infra-hermes builds and deploys; **infra-ops AUDITS when it reports**, the same way as the Jev endpoint: re-run the latency bins, the WER spot-check, the defect checks and memory myself. The details below are the brief it got.
|
||||
- **DONE: LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, image `local/parakeet-nemo:nemo-0.1.0`, infra-hermes de6ea32 + 41d2014) on :8300; LiteLLM untouched. **infra-ops AUDIT PASSED 0137.**
|
||||
- Latency on GPU 0, p50 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms before. Through LiteLLM a 2.5 s clip takes 80–98 ms.
|
||||
- WER: LibriSpeech clean 1.965, other 3.026.
|
||||
- Fixed: a 714–726 s file returns 200 (long files go in 360 s windows, because NeMo builds the full T×T attention mask even under local attention); no drop after a pause.
|
||||
- Rollback: `docker stop parakeet-nemo && docker start parakeet` (the old container is stopped, not removed).
|
||||
- **gen-small: util 0.48 → 0.36** (0.46 and 0.40 failed its boot check). Its KV is byte-pinned (`--kv-cache-memory 8 GiB`), still 670,142 tokens / 2.56×, so there was no KV cost. It was down ~34 min (0012–0046 PDT) while the util was iterated; zero LiteLLM errors.
|
||||
- ⚠ **Steady state is 3,582 MiB for the seat (its cached window peak); GPU 0 Free is 385 MiB.** That leaves gen-small's restart boot-check margin at only ~0.45 GiB. infra-hermes was asked to set gen-small util to 0.33 in the .env WITHOUT restarting, so it applies at the next restart. **Nothing else fits on GPU 0.**
|
||||
- Seat invariants are in its README: cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11`, because httptools 0.8.0 writes `HTTP/1.1 200\x00OK` and LiteLLM/httpx rejects it; the 360 s window.
|
||||
- The original brief, for reference:
|
||||
|
||||
- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use.
|
||||
- The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`.
|
||||
@@ -305,7 +313,7 @@ _As of 2026-09-30 ~1800 PT._
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. Implementation was TASKED to infra-hermes (2354); infra-ops audits. Tracked by the in-flight "NEXT" section and a6c1d3c.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
|
||||
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
|
||||
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
|
||||
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
|
||||
- `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
|
||||
|
||||
Reference in New Issue
Block a user