Files
esh-pfi-infrastructure/stacks/parakeet-nemo/README.md
T

70 lines
4.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
## Why this runtime
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
long-audio mode (±128), which is what removes all three defects.
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
(`WARMUP_SECONDS`, default 1,8,60).
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
grows linearly in length instead of quadratically.
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
differently call to call.
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
auto-selected by uvicorn when importable) writes a NUL into the response status line —
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
installs plain `uvicorn==0.53.0` for exactly this reason.
## License
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
required; keep this section as the license note. Weights pinned at HF revision
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
## GPU 0 room
The seat rests ~2.1 GB, serves to ~2.8 GB, loads under ~3.0 GB (see the ops log). But its
**steady state after ANY long (windowed) request is 3,582 MiB** — torch caches the window peak
and does not return it (measured flat across repeated 12-min requests, audit 2026-10-01). Plan
GPU 0 against 3,582 MiB, not the at-rest figure: in steady state the card sits at ~385 MiB Free.
Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.36 at cut-over → 0.33
after the audit (config-only; applies at its NEXT restart, and gives that restart ~3 GiB of
boot-check margin against the steady-state figure). Its KV is byte-pinned
(`--kv-cache-memory`), so the util number costs it nothing — boot log identical at 670,142
tokens / 2.56×. ⚠ Before restarting ANY vLLM seat on this card, do the boot-check arithmetic
against measured `nvidia-smi` Free: required = util × total, available = Free + that seat's own
resident memory. The cut-over iteration (0.46, 0.40 both refusing boot before 0.36 booted) took
gen-small down ~34 minutes for want of that one line of arithmetic. util does NOT predict
resident VRAM. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
## Rollback
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
restores the sherpa seat on :8300 exactly as before (container and image both kept).
## Deploy
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).