memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived
This commit is contained in:
@@ -0,0 +1,20 @@
|
||||
# Parakeet speech seat → parakeet-unified-en-0.6b under NeMo: APPROVED, implementation next session (2026-09-30)
|
||||
|
||||
**Rulings:**
|
||||
- Prime ~1558: "a/b the one on fv-ml1's general seat against the unified new one in jun for speed and accuracy for english… any win, even 50ms, is load-bearing."
|
||||
- Prime ~1755, after the A/B and a licence summary: "reasonable terms, ship the switch."
|
||||
- **The NVIDIA Open Model License is accepted for internal use.** It allows commercial use. NVIDIA may revise the terms. The licence terminates on IP litigation over the model or on bypassing guardrails. We indemnify NVIDIA. Redistribution needs a NOTICE.
|
||||
|
||||
**A/B** (`docs/pfi/parakeet-seat-ab-2026-09-30.md`, a6c1d3c, b38ec6d):
|
||||
- The seat's latency is its RUNTIME: the sherpa-onnx int8 graph runs on one CPU thread, with the GPU at 2–9%.
|
||||
- End-to-end p50 for 1–3 / 3–8 / 8–20 s clips: the seat 144 / 260 / 565 ms; unified-en under NeMo with bf16 weights 23 / 27 / 33 ms.
|
||||
- Floor ≤ 6 ms; a +50 ms positive control read +52.
|
||||
- WER: LibriSpeech clean 2.70 → 1.97, other 4.56 → 3.09, AMI 12.69 → 8.30.
|
||||
- Unified int8 in the seat's runtime was SLOWER than the seat. v3 fp32 ONNX was 4–12× faster in the same image (the fallback if NeMo is blocked).
|
||||
- Seat defects: HTTP 500 above ~400 s; long-form dropouts; the rest of an utterance dropped after a 1.5 s digital-silence pause.
|
||||
|
||||
**Implementation plan (tracked by the in-flight "NEXT" section and the /tmp handoff):**
|
||||
1. Build an image from `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, with a warm-up, a bf16 cast before `.to(cuda)`, and local attention for long files.
|
||||
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
|
||||
3. Cut over with the old seat kept as the rollback.
|
||||
4. Re-measure live on GPU 0.
|
||||
Reference in New Issue
Block a user