Files
esh-pfi-infrastructure/stacks/parakeet-nemo

parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)

STT seat on fv-ml1, port 8300, behind LiteLLM as ext-stt / whisper-1; the caller is talk. Switched over from the sherpa-onnx int8 seat (stacks/parakeet) on 2026-09-30 on Prime's order, after the A/B in docs/pfi/parakeet-seat-ab-2026-09-30.md (arm B-bf16w won: p50 23/27/33 ms vs the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).

Why this runtime

The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention long-audio mode (±128), which is what removes all three defects.

Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")

  • bf16 cast BEFORE .to("cuda") — restoring fp32 onto the GPU and casting there spikes the load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
  • Warm-up at the longest served length — the CUDA-graph greedy decoder costs ~330 ms extra on the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips (WARMUP_SECONDS, default 1,8,60).
  • rel_pos_local_attn ±128 — a 30-minute file transcribes in ~2.6 s in ONE request; memory grows linearly in length instead of quadratically.
  • dither = 0.0 — dither is a training-time augmentation; it makes identical files decode differently call to call.
  • NeMo 3.0.0 is a floor — released 2.7.3 lacks this encoder's att_chunk_context_size.
  • No httptools; --http h11 is explicit. httptools 0.8.0 (pulled by uvicorn[standard], and auto-selected by uvicorn when importable) writes a NUL into the response status line — HTTP/1.1 200\x00OK — that h11/httpx reject with RemoteProtocolError: illegal status line. curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven A/B on the same image: --http h11 clean, --http httptools dirty (2026-10-01). The image installs plain uvicorn==0.53.0 for exactly this reason.

License

nvidia/parakeet-unified-en-0.6b is distributed under the NVIDIA Open Model License Agreement (commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file required; keep this section as the license note. Weights pinned at HF revision fe53cd885760c96b6a5f51a0bfd362cb4584a98b (sha256 ec23ed91…), mounted read-only from /tank/aimodels/huggingface, HF_HUB_OFFLINE=1.

GPU 0 room

The seat rests ~2.1 GB, serves to ~2.8 GB, loads under ~3.0 GB (see the ops log). But its steady state after ANY long (windowed) request is 3,582 MiB — torch caches the window peak and does not return it (measured flat across repeated 12-min requests, audit 2026-10-01). Plan GPU 0 against 3,582 MiB, not the at-rest figure: in steady state the card sits at ~385 MiB Free. Room was taken from vllm-gen-small: --gpu-memory-utilization 0.48 → 0.36 at cut-over → 0.33 after the audit (config-only; applies at its NEXT restart, and gives that restart ~3 GiB of boot-check margin against the steady-state figure). Its KV is byte-pinned (--kv-cache-memory), so the util number costs it nothing — boot log identical at 670,142 tokens / 2.56×. ⚠ Before restarting ANY vLLM seat on this card, do the boot-check arithmetic against measured nvidia-smi Free: required = util × total, available = Free + that seat's own resident memory. The cut-over iteration (0.46, 0.40 both refusing boot before 0.36 booted) took gen-small down ~34 minutes for want of that one line of arithmetic. util does NOT predict resident VRAM. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.

Rollback

The old seat was STOPPED, not removed: docker stop parakeet-nemo && docker start parakeet restores the sherpa seat on :8300 exactly as before (container and image both kept).

Deploy

Build on fv-ml1 in a versioned dir under /opt/docker/src/ (house convention), tag local/parakeet-nemo:nemo-X.Y.Z, point .env at it, docker compose up -d. Acceptance harness and the A/B's paired latency/WER tooling: /tank/spikes/parakeet-ab on fv-ml1 (do not delete).