Files
esh-pfi-infrastructure/stacks/parakeet-nemo/README.md
T
vh 963f9ed8c0 fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:

- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
  returns to ~2,108 MiB rest after a 12-min file instead of parking at the
  peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
  answer 503 with the seat alive (proved at cap=2000), so the failure lands
  on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
  free (illegal-memory-access wedge when both were on first try). Cost:
  12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
  under the sherpa seat.

Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.
2026-10-01 04:39:39 -07:00

80 lines
5.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
## Why this runtime
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
long-audio mode (±128), which is what removes all three defects.
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
(`WARMUP_SECONDS`, default 1,8,60).
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
grows linearly in length instead of quadratically.
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
differently call to call.
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
auto-selected by uvicorn when importable) writes a NUL into the response status line —
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
installs plain `uvicorn==0.53.0` for exactly this reason.
## License
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
required; keep this section as the license note. Weights pinned at HF revision
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
## GPU 0 room
The seat rests ~2.1 GB. **Since nemo-0.1.1 it returns to rest after long requests**: the
windowed path calls `torch.cuda.empty_cache()` around each window, so a 12-min file peaks at
~3,028 MiB during the request and falls back to ~2,108 after (measured 2026-10-01, restarts=0).
Before 0.1.1 the seat PARKED at the window peak (3,582 MiB steady), and that cached peak left
gen-small no room for its runtime workspace — an EngineCore CUDA-OOM incident at 04:21 PT.
Three seat-side controls, all load-bearing:
- **`MEM_CAP_MIB=3840`** (env, default): a hard `set_per_process_memory_fraction` ceiling. An
over-cap request answers **503** with the seat still alive (proved at cap=2000: two 503s, then
short requests fine) — the failure lands on us, never on a neighbour's allocation.
- **`CUDA_GRAPHS=0`** (default): the CUDA-graph greedy decoder pins memory in torch's cache and
died with an illegal-memory-access the first time `empty_cache` freed a graph-pool block.
Graphs off costs latency (12-min file 3.0 s vs 1.2 s; short bins 35-62 ms vs 33-42 ms — still
4-15x faster than the sherpa seat) and buys a lower, honestly-returned footprint.
- **`WINDOW_S=360`** (default): see the windowing note above; mask is T×T even under local attention.
Room came from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.33 (config-only, applies at
its next restart; KV byte-pinned, boot log identical at 670,142 tokens / 2.56×). gen-small's
runtime growth was measured nvidia-smi-per-process, not computed: 3 realistic requests
(1,351 prompt / ~180 completion tokens) moved it ZERO from its 36,116 MiB — the workspace
allocation lands at engine init, right after restart (infra-ops saw the growth window at
restart+3-requests). ⚠ Before restarting ANY vLLM seat on this card, do the boot-check
arithmetic against measured `nvidia-smi` Free: required = util × total, available = Free + that
seat's own resident memory. GPU 1 is NOT an option: its free memory is intern-decision's 32k
headroom.
## Rollback
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
restores the sherpa seat on :8300 exactly as before (container and image both kept).
## Deploy
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).