Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/ whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the 2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026 vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s (windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own long-form method), no pause truncation, no long-form dropout. GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40 refuse their boot check; cyberprev+voices hold the card). Its KV is byte- pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free. Two runtime landmines documented in the README: NeMo's attention mask is materialised T x T even under local attention (hence the window), and httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e. LiteLLM — rejects, so the image ships plain uvicorn with --http h11. Old seat stopped, not removed: docker stop parakeet-nemo && docker start parakeet is the rollback. License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in stacks/parakeet-nemo/README.md.
12 lines
596 B
Bash
12 lines
596 B
Bash
# Copy to .env next to compose.yaml on the host.
|
|
PARAKEET_NEMO_TAG=nemo-0.1.0
|
|
# Port the seat listens on. 8300 is the seat port LiteLLM's ext-stt/whisper-1 point at;
|
|
# run acceptance on a temporary port first, then cut over by changing this line.
|
|
PARAKEET_NEMO_PORT=8300
|
|
# PARAKEET_NEMO_BIND=0.0.0.0
|
|
# PARAKEET_NEMO_GPU=0
|
|
# Pinned HF revision of nvidia/parakeet-unified-en-0.6b (sha256 ec23ed91... of the .nemo).
|
|
PARAKEET_NEMO_REV=fe53cd885760c96b6a5f51a0bfd362cb4584a98b
|
|
# Ascending silent warm-up clips in seconds (CUDA-graph capture + longest-shape kernel warm).
|
|
# PARAKEET_NEMO_WARMUP=1,8,60
|