feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/ whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the 2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026 vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s (windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own long-form method), no pause truncation, no long-form dropout. GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40 refuse their boot check; cyberprev+voices hold the card). Its KV is byte- pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free. Two runtime landmines documented in the README: NeMo's attention mask is materialised T x T even under local attention (hence the window), and httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e. LiteLLM — rejects, so the image ships plain uvicorn with --http h11. Old seat stopped, not removed: docker stop parakeet-nemo && docker start parakeet is the rollback. License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in stacks/parakeet-nemo/README.md.
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
|
||||
|
||||
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
|
||||
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
|
||||
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
|
||||
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
|
||||
|
||||
## Why this runtime
|
||||
|
||||
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
|
||||
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
|
||||
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
|
||||
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
|
||||
long-audio mode (±128), which is what removes all three defects.
|
||||
|
||||
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
|
||||
|
||||
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
|
||||
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
|
||||
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
|
||||
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
|
||||
(`WARMUP_SECONDS`, default 1,8,60).
|
||||
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
|
||||
grows linearly in length instead of quadratically.
|
||||
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
|
||||
differently call to call.
|
||||
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
|
||||
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
|
||||
auto-selected by uvicorn when importable) writes a NUL into the response status line —
|
||||
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
|
||||
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
|
||||
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
|
||||
installs plain `uvicorn==0.53.0` for exactly this reason.
|
||||
|
||||
## License
|
||||
|
||||
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
|
||||
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
|
||||
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
|
||||
required; keep this section as the license note. Weights pinned at HF revision
|
||||
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
|
||||
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
|
||||
|
||||
## GPU 0 room
|
||||
|
||||
The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over;
|
||||
see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46
|
||||
(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT
|
||||
predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before
|
||||
believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
|
||||
|
||||
## Rollback
|
||||
|
||||
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
|
||||
restores the sherpa seat on :8300 exactly as before (container and image both kept).
|
||||
|
||||
## Deploy
|
||||
|
||||
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
|
||||
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
|
||||
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).
|
||||
Reference in New Issue
Block a user