Files
esh-pfi-infrastructure/stacks/parakeet-nemo/README.md
T
vh de6ea32f34 feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
2026-10-01 01:32:48 -07:00

62 lines
3.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
## Why this runtime
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
long-audio mode (±128), which is what removes all three defects.
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
(`WARMUP_SECONDS`, default 1,8,60).
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
grows linearly in length instead of quadratically.
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
differently call to call.
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
auto-selected by uvicorn when importable) writes a NUL into the response status line —
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
installs plain `uvicorn==0.53.0` for exactly this reason.
## License
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
required; keep this section as the license note. Weights pinned at HF revision
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
## GPU 0 room
The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over;
see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46
(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT
predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before
believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
## Rollback
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
restores the sherpa seat on :8300 exactly as before (container and image both kept).
## Deploy
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).