Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md
T

21 lines
1.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Parakeet speech seat → parakeet-unified-en-0.6b under NeMo: APPROVED, implementation next session (2026-09-30)
**Rulings:**
- Prime ~1558: "a/b the one on fv-ml1's general seat against the unified new one in jun for speed and accuracy for english… any win, even 50ms, is load-bearing."
- Prime ~1755, after the A/B and a licence summary: "reasonable terms, ship the switch."
- **The NVIDIA Open Model License is accepted for internal use.** It allows commercial use. NVIDIA may revise the terms. The licence terminates on IP litigation over the model or on bypassing guardrails. We indemnify NVIDIA. Redistribution needs a NOTICE.
**A/B** (`docs/pfi/parakeet-seat-ab-2026-09-30.md`, a6c1d3c, b38ec6d):
- The seat's latency is its RUNTIME: the sherpa-onnx int8 graph runs on one CPU thread, with the GPU at 2–9%.
- End-to-end p50 for 1–3 / 3–8 / 8–20 s clips: the seat 144 / 260 / 565 ms; unified-en under NeMo with bf16 weights 23 / 27 / 33 ms.
- Floor ≤ 6 ms; a +50 ms positive control read +52.
- WER: LibriSpeech clean 2.70 → 1.97, other 4.56 → 3.09, AMI 12.69 → 8.30.
- Unified int8 in the seat's runtime was SLOWER than the seat. v3 fp32 ONNX was 4–12× faster in the same image (the fallback if NeMo is blocked).
- Seat defects: HTTP 500 above ~400 s; long-form dropouts; the rest of an utterance dropped after a 1.5 s digital-silence pause.
**Implementation: TASKED to infra-hermes at 2354 2026-09-30 (Prime: "send the switch to infra-hermes"); infra-ops audits on report. The plan as briefed:**
1. Build an image from `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, with a warm-up, a bf16 cast before `.to(cuda)`, and local attention for long files.
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
3. Cut over with the old seat kept as the rollback.
4. Re-measure live on GPU 0.