memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits

This commit is contained in:
vh
2026-09-30 23:54:54 -07:00
parent 1ae324d576
commit aa7eeff445
2 changed files with 4 additions and 2 deletions
@@ -13,7 +13,7 @@
- Unified int8 in the seat's runtime was SLOWER than the seat. v3 fp32 ONNX was 4–12× faster in the same image (the fallback if NeMo is blocked).
- Seat defects: HTTP 500 above ~400 s; long-form dropouts; the rest of an utterance dropped after a 1.5 s digital-silence pause.
**Implementation plan (tracked by the in-flight "NEXT" section and the /tmp handoff):**
**Implementation: TASKED to infra-hermes at 2354 2026-09-30 (Prime: "send the switch to infra-hermes"); infra-ops audits on report. The plan as briefed:**
1. Build an image from `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, with a warm-up, a bf16 cast before `.to(cuda)`, and local attention for long files.
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
3. Cut over with the old seat kept as the rollback.