The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left no room for vLLM's runtime workspace. Seat-side fix, three controls: - windowed path wraps every window in torch.cuda.empty_cache(), so the seat returns to ~2,108 MiB rest after a 12-min file instead of parking at the peak (measured: peak 3,028 MiB during, rest after, restarts=0); - MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests answer 503 with the seat alive (proved at cap=2000), so the failure lands on us, never on a neighbour; - CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must free (illegal-memory-access wedge when both were on first try). Cost: 12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x under the sherpa seat. Measured, not computed: gen-small moved ZERO from 36,116 MiB across three realistic requests (1,351 in / ~180 out) -- its workspace lands at engine init; the growth window is restart-relative, matching infra-ops's observation. WINDOW_S is now a real compose tunable. README memory section rewritten.
80 lines
5.2 KiB
Markdown
80 lines
5.2 KiB
Markdown
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
|
||
|
||
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
|
||
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
|
||
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
|
||
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
|
||
|
||
## Why this runtime
|
||
|
||
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
|
||
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
|
||
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
|
||
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
|
||
long-audio mode (±128), which is what removes all three defects.
|
||
|
||
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
|
||
|
||
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
|
||
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
|
||
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
|
||
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
|
||
(`WARMUP_SECONDS`, default 1,8,60).
|
||
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
|
||
grows linearly in length instead of quadratically.
|
||
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
|
||
differently call to call.
|
||
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
|
||
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
|
||
auto-selected by uvicorn when importable) writes a NUL into the response status line —
|
||
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
|
||
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
|
||
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
|
||
installs plain `uvicorn==0.53.0` for exactly this reason.
|
||
|
||
## License
|
||
|
||
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
|
||
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
|
||
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
|
||
required; keep this section as the license note. Weights pinned at HF revision
|
||
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
|
||
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
|
||
|
||
## GPU 0 room
|
||
|
||
The seat rests ~2.1 GB. **Since nemo-0.1.1 it returns to rest after long requests**: the
|
||
windowed path calls `torch.cuda.empty_cache()` around each window, so a 12-min file peaks at
|
||
~3,028 MiB during the request and falls back to ~2,108 after (measured 2026-10-01, restarts=0).
|
||
Before 0.1.1 the seat PARKED at the window peak (3,582 MiB steady), and that cached peak left
|
||
gen-small no room for its runtime workspace — an EngineCore CUDA-OOM incident at 04:21 PT.
|
||
Three seat-side controls, all load-bearing:
|
||
- **`MEM_CAP_MIB=3840`** (env, default): a hard `set_per_process_memory_fraction` ceiling. An
|
||
over-cap request answers **503** with the seat still alive (proved at cap=2000: two 503s, then
|
||
short requests fine) — the failure lands on us, never on a neighbour's allocation.
|
||
- **`CUDA_GRAPHS=0`** (default): the CUDA-graph greedy decoder pins memory in torch's cache and
|
||
died with an illegal-memory-access the first time `empty_cache` freed a graph-pool block.
|
||
Graphs off costs latency (12-min file 3.0 s vs 1.2 s; short bins 35-62 ms vs 33-42 ms — still
|
||
4-15x faster than the sherpa seat) and buys a lower, honestly-returned footprint.
|
||
- **`WINDOW_S=360`** (default): see the windowing note above; mask is T×T even under local attention.
|
||
Room came from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.33 (config-only, applies at
|
||
its next restart; KV byte-pinned, boot log identical at 670,142 tokens / 2.56×). gen-small's
|
||
runtime growth was measured nvidia-smi-per-process, not computed: 3 realistic requests
|
||
(1,351 prompt / ~180 completion tokens) moved it ZERO from its 36,116 MiB — the workspace
|
||
allocation lands at engine init, right after restart (infra-ops saw the growth window at
|
||
restart+3-requests). ⚠ Before restarting ANY vLLM seat on this card, do the boot-check
|
||
arithmetic against measured `nvidia-smi` Free: required = util × total, available = Free + that
|
||
seat's own resident memory. GPU 1 is NOT an option: its free memory is intern-decision's 32k
|
||
headroom.
|
||
|
||
## Rollback
|
||
|
||
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
|
||
restores the sherpa seat on :8300 exactly as before (container and image both kept).
|
||
|
||
## Deploy
|
||
|
||
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
|
||
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
|
||
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).
|