feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)

Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
This commit is contained in:
vh
2026-10-01 01:32:48 -07:00
parent dbd583d6ca
commit de6ea32f34
5 changed files with 324 additions and 0 deletions
+63
View File
@@ -0,0 +1,63 @@
# Parakeet ASR via NeMo torch (unified-en-0.6b, bf16 weights) + our own thin FastAPI wrapper.
#
# Replacement seat for stacks/parakeet (sherpa-onnx int8). Same port (:8300), same endpoints,
# same body — LiteLLM and `talk` need no change. Rollback: stop this container, start the old
# `parakeet` one (kept; container and image both intact).
#
# HOST: fv-ml1, GPU 0.
#
# ⚠ GPU 0 room came from gen-small's KV: its --gpu-memory-utilization dropped 0.48 -> 0.46
# (measured boot, 2026-09-30: the seat rests ~2.5 GB, served peak ~2.8 GB, load peak ~3.0 GB;
# the old seat held 1,690 MiB). Do not raise that util back without re-measuring nvidia-smi Free
# on GPU 0 — util does not predict resident VRAM (see the 09-15 note in gen-small-seat/.env).
#
# ⚠ Weights: /tank/aimodels/huggingface mounted READ-ONLY. nvidia/parakeet-unified-en-0.6b @
# fe53cd885760c96b6a5f51a0bfd362cb4584a98b. HF_HUB_OFFLINE=1 in the image: the seat never phones home.
#
# API (identical to the replaced seat):
# POST /transcribe — multipart file upload, returns {"text": "..."}
# POST /v1/audio/transcriptions — same body, OpenAI-compatible path alias
# GET /healthz
#
# All tunables live in .env — edit that, not this file.
services:
parakeet-nemo:
image: local/parakeet-nemo:${PARAKEET_NEMO_TAG}
container_name: parakeet-nemo
restart: unless-stopped
ports:
- "${PARAKEET_NEMO_BIND:-0.0.0.0}:${PARAKEET_NEMO_PORT}:8000"
environment:
- MODEL_PATH=/hf/hub/models--nvidia--parakeet-unified-en-0.6b/snapshots/${PARAKEET_NEMO_REV}/parakeet-unified-en-0.6b.nemo
- WARMUP_SECONDS=${PARAKEET_NEMO_WARMUP:-1,8,60}
- LOG_LEVEL=${PARAKEET_NEMO_LOG_LEVEL:-INFO}
volumes:
- /tank/aimodels/huggingface:/hf:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${PARAKEET_NEMO_GPU:-0}"]
capabilities: [gpu]
networks:
- tnet
healthcheck:
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8000/healthz || exit 1"]
interval: 30s
timeout: 10s
retries: 3
# Import-time model load + three warm-up decodes; no download (weights are mounted).
start_period: 240s
labels:
- homepage.group=AI - Audio Tools
- homepage.name=Parakeet ASR (NeMo)
- homepage.icon=mdi-microphone
- homepage.description=Parakeet-unified-en speech-to-text via NeMo bf16 (fv-ml1 GPU 0)
- homepage.href=http://10.251.50.54:${PARAKEET_NEMO_PORT}
networks:
tnet:
name: traefik-net
external: true