memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits

This commit is contained in:
vh
2026-09-30 23:54:54 -07:00
parent 1ae324d576
commit aa7eeff445
2 changed files with 4 additions and 2 deletions
@@ -13,7 +13,7 @@
- Unified int8 in the seat's runtime was SLOWER than the seat. v3 fp32 ONNX was 4–12× faster in the same image (the fallback if NeMo is blocked). - Unified int8 in the seat's runtime was SLOWER than the seat. v3 fp32 ONNX was 4–12× faster in the same image (the fallback if NeMo is blocked).
- Seat defects: HTTP 500 above ~400 s; long-form dropouts; the rest of an utterance dropped after a 1.5 s digital-silence pause. - Seat defects: HTTP 500 above ~400 s; long-form dropouts; the rest of an utterance dropped after a 1.5 s digital-silence pause.
**Implementation plan (tracked by the in-flight "NEXT" section and the /tmp handoff):** **Implementation: TASKED to infra-hermes at 2354 2026-09-30 (Prime: "send the switch to infra-hermes"); infra-ops audits on report. The plan as briefed:**
1. Build an image from `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, with a warm-up, a bf16 cast before `.to(cuda)`, and local attention for long files. 1. Build an image from `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, with a warm-up, a bf16 cast before `.to(cuda)`, and local attention for long files.
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM. 2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
3. Cut over with the old seat kept as the rollback. 3. Cut over with the old seat kept as the rollback.
+3 -1
View File
@@ -119,6 +119,8 @@ _As of 2026-09-30 ~1800 PT._
### NEXT: switch the Parakeet speech seat (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch") ### NEXT: switch the Parakeet speech seat (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
- **TASKED to infra-hermes at 2354 2026-09-30** (Prime corrected me: "send the switch to infra-hermes"). infra-hermes builds and deploys; **infra-ops AUDITS when it reports**, the same way as the Jev endpoint: re-run the latency bins, the WER spot-check, the defect checks and memory myself. The details below are the brief it got.
- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use. - **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use.
- The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`. - The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`.
- Evidence, `docs/pfi/parakeet-seat-ab-2026-09-30.md` (a6c1d3c): end-to-end p50 for 1–3 / 3–8 / 8–20 s clips goes 144 / 260 / 565 → 23 / 27 / 33 ms, and WER is lower on every set. - Evidence, `docs/pfi/parakeet-seat-ab-2026-09-30.md` (a6c1d3c): end-to-end p50 for 1–3 / 3–8 / 8–20 s clips goes 144 / 260 / 565 → 23 / 27 / 33 ms, and WER is lower on every set.
@@ -308,7 +310,7 @@ _As of 2026-09-30 ~1800 PT._
## Recent decisions ## Recent decisions
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. Implementation is deferred to the next session, tracked by the in-flight "NEXT" section and a6c1d3c.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` - `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. Implementation was TASKED to infra-hermes (2354); infra-ops audits. Tracked by the in-flight "NEXT" section and a6c1d3c.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` - `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md` - `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
- `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` - `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`