feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)

Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
This commit is contained in:
vh
2026-10-01 01:32:48 -07:00
parent dbd583d6ca
commit de6ea32f34
5 changed files with 324 additions and 0 deletions
+11
View File
@@ -0,0 +1,11 @@
# Copy to .env next to compose.yaml on the host.
PARAKEET_NEMO_TAG=nemo-0.1.0
# Port the seat listens on. 8300 is the seat port LiteLLM's ext-stt/whisper-1 point at;
# run acceptance on a temporary port first, then cut over by changing this line.
PARAKEET_NEMO_PORT=8300
# PARAKEET_NEMO_BIND=0.0.0.0
# PARAKEET_NEMO_GPU=0
# Pinned HF revision of nvidia/parakeet-unified-en-0.6b (sha256 ec23ed91... of the .nemo).
PARAKEET_NEMO_REV=fe53cd885760c96b6a5f51a0bfd362cb4584a98b
# Ascending silent warm-up clips in seconds (CUDA-graph capture + longest-shape kernel warm).
# PARAKEET_NEMO_WARMUP=1,8,60
+35
View File
@@ -0,0 +1,35 @@
# Parakeet ASR seat: parakeet-unified-en-0.6b under NeMo torch, bf16 weights.
# CUDA 12.8 runtime base + a uv-managed venv pinned to the A/B's proven stack
# (torch 2.8 cu128, nemo_toolkit[asr]==3.0.0; the A/B found NeMo 2.7.3 lacks this encoder's
# att_chunk_context_size, so 3.0.0 is a floor, not a preference).
# Weights are NOT baked in: /tank/aimodels/huggingface is bind-mounted read-only (see compose).
FROM nvidia/cuda:12.8.1-base-ubuntu24.04
ENV DEBIAN_FRONTEND=noninteractive \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
PYTHONUNBUFFERED=1 \
HF_HUB_OFFLINE=1
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-venv python3-pip wget libsndfile1 ca-certificates \
&& rm -rf /var/lib/apt/lists/*
RUN python3 -m venv /opt/venv \
&& /opt/venv/bin/pip install -q uv \
&& UV_LINK_MODE=copy /opt/venv/bin/uv pip install -q --python /opt/venv/bin/python \
--index-url https://download.pytorch.org/whl/cu128 \
--extra-index-url https://pypi.org/simple \
"torch==2.8.*" "torchaudio==2.8.*" "nemo_toolkit[asr]==3.0.0" \
fastapi "uvicorn==0.53.0" python-multipart soundfile \
&& /opt/venv/bin/python -c "import nemo, torch; print('nemo', nemo.__version__, 'torch', torch.__version__, 'cuda_ok', torch.cuda.is_available())"
WORKDIR /app
COPY app.py /app/app.py
# The seat's only writable need is NeMo/HF scratch; keep it off the rootfs surprises.
ENV HOME=/tmp
EXPOSE 8000
# --http h11: the [standard] extra pulls httptools, and httptools 0.8.0 writes a NUL into the
# status line ("HTTP/1.1 200\x00OK") that h11/httpx reject. uvicorn auto-picks httptools when
# importable, so it must stay UNinstalled and the flag must stay explicit. See README.
CMD ["/opt/venv/bin/uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000", "--http", "h11"]
+61
View File
@@ -0,0 +1,61 @@
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
## Why this runtime
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
long-audio mode (±128), which is what removes all three defects.
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
(`WARMUP_SECONDS`, default 1,8,60).
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
grows linearly in length instead of quadratically.
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
differently call to call.
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
auto-selected by uvicorn when importable) writes a NUL into the response status line —
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
installs plain `uvicorn==0.53.0` for exactly this reason.
## License
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
required; keep this section as the license note. Weights pinned at HF revision
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
## GPU 0 room
The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over;
see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46
(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT
predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before
believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
## Rollback
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
restores the sherpa seat on :8300 exactly as before (container and image both kept).
## Deploy
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).
+154
View File
@@ -0,0 +1,154 @@
"""Parakeet ASR seat: nvidia/parakeet-unified-en-0.6b under NeMo torch (bf16 weights).
Fork of the A/B winner (services/parakeet-ab-2026-09-30 arm B-bf16w), with the three shipping
changes that arm's doc called for and a longer warm-up. Same HTTP shape as the sherpa seat it
replaces: model loaded at import, warm-up before traffic, async handlers over a serialised
blocking decode, {"text": ...} responses on /transcribe and /v1/audio/transcriptions.
Hard-wired, because every one of these is load-bearing on GPU 0:
bf16 weights — the encoder/decoder/joint are cast to bfloat16 (mel front end stays fp32).
WER is identical to fp32 (399/400 utterances byte-equal in the A/B).
CPU-then-cast — the .nemo restores on CPU, weights are cast to bf16 there, and only then move
to the GPU. Restoring to CUDA spikes the load by ~1.5 GB; GPU 0 cannot absorb it.
local attn — rel_pos_local_attn ±128 (NeMo's documented long-audio mode). This is what fixes
the old seat's >400 s HTTP 500 and the long-form dropouts; memory grows linearly
in file length. A 30-min file transcribes in ~2.6 s in one request (A/B, 2026-09-30).
Warm-up runs ascending silent clips (1 s, 8 s, 60 s): the CUDA-graph greedy decoder and the
encoder kernels cost ~330 ms extra on the first call at a new maximum length.
"""
from __future__ import annotations
import io
import logging
import os
import time
import numpy as np
import soundfile as sf
import torch
from fastapi import FastAPI, File, HTTPException, UploadFile
from fastapi.responses import JSONResponse
MODEL_PATH = os.environ["MODEL_PATH"]
WARMUP_SECONDS = [int(x) for x in os.environ.get("WARMUP_SECONDS", "1,8,60").split(",")]
SR = 16000
logger = logging.getLogger("parakeet-nemo")
logging.basicConfig(level=os.environ.get("LOG_LEVEL", "INFO"))
def _load():
import nemo.collections.asr as nemo_asr
from omegaconf import open_dict
t0 = time.monotonic()
m = nemo_asr.models.ASRModel.restore_from(MODEL_PATH, map_location="cpu")
m.eval()
if m.cfg.get("validation_ds") is None: # the unified .nemo ships without it; transcribe() reads it
with open_dict(m.cfg):
m.cfg.validation_ds = {}
d = m.cfg.decoding
with open_dict(d):
d.strategy = "greedy_batch"
d.greedy["use_cuda_graph_decoder"] = True
m.change_decoding_strategy(d, verbose=False)
# transcribe() sets these on entry; the direct path must match, and must not dither (dither is
# a training-time augmentation and makes the same file decode differently on each call).
m.preprocessor.featurizer.dither = 0.0
m.preprocessor.featurizer.pad_to = 0
m.change_attention_model("rel_pos_local_attn", [128, 128])
# bf16 the serving modules BEFORE the H2D copy: halves the transfer and skips the GPU-side
# fp32->bf16 transient entirely (the measured load spike goes from 3,194 MiB to under 2,600).
# ⚠ Must run AFTER change_attention_model: that call rebuilds the attention modules in fp32,
# and casting first leaves fp32 islands behind (RuntimeError: mat1 and mat2 ... BFloat16/Float,
# hit live at the first boot of this image, 2026-10-01).
for mod in (m.encoder, m.decoder, m.joint):
mod.to(torch.bfloat16)
m = m.to("cuda")
logger.info("loaded %s (%s) bf16w local_att=128,128 in %.1fs", os.path.basename(MODEL_PATH),
type(m).__name__, time.monotonic() - t0)
return m
model = _load()
def _hyp_text(h) -> str:
if isinstance(h, str):
return h
t = getattr(h, "text", None)
if isinstance(t, str):
return t
return model.tokenizer.ids_to_text([int(i) for i in h.y_sequence])
@torch.inference_mode()
def _infer(samples: np.ndarray) -> str:
x = torch.from_numpy(samples).to("cuda", non_blocking=True).unsqueeze(0)
xl = torch.tensor([x.shape[1]], device="cuda", dtype=torch.long)
feats, fl = model.preprocessor(input_signal=x, length=xl)
feats = feats.to(torch.bfloat16)
enc, el = model.encoder(audio_signal=feats, length=fl)
hyps = model.decoding.rnnt_decoder_predictions_tensor(encoder_output=enc, encoded_lengths=el,
return_hypotheses=False)
if isinstance(hyps, tuple):
hyps = hyps[0]
return _hyp_text(hyps[0])
def _warm() -> None:
for secs in WARMUP_SECONDS:
t0 = time.monotonic()
_infer(np.zeros(SR * secs, dtype=np.float32))
torch.cuda.synchronize()
logger.info("warmup %ss decode complete in %.1fs", secs, time.monotonic() - t0)
_warm()
app = FastAPI(title="Parakeet ASR (NeMo torch, bf16w)")
def _decode(raw: bytes) -> str:
try:
samples, sample_rate = sf.read(io.BytesIO(raw), dtype="float32")
except Exception as exc:
raise HTTPException(400, f"Could not decode audio: {exc}") from exc
if samples.ndim > 1:
samples = samples.mean(axis=1).astype(np.float32)
if sample_rate != SR:
import torchaudio.functional as AF
samples = AF.resample(torch.from_numpy(samples), sample_rate, SR).numpy()
samples = np.ascontiguousarray(samples, dtype=np.float32)
# Windowed long-form (> WINDOW_S): NeMo's attention mask is materialised T×T even under
# rel_pos_local_attn, so one whole 12-min request wanted +1.1 GiB of scratch and OOMed on
# GPU 0 (measured at acceptance, 2026-10-01). The A/B's own long-form arm used ~6-min
# windows and lost zero clean speech on unified-en in 4/4 placements (doc § 5.4), so the
# seat chunks at the same size: bounded memory, any length, no API change for callers.
win = int(os.environ.get("WINDOW_S", "360")) * SR
if len(samples) <= win:
return _infer(samples)
parts = [_infer(samples[i:i + win]) for i in range(0, len(samples), win)]
return " ".join(p for p in parts if p)
def _timed(raw: bytes) -> JSONResponse:
t0 = time.perf_counter()
text = _decode(raw)
torch.cuda.synchronize()
ms = (time.perf_counter() - t0) * 1000.0
return JSONResponse({"text": text}, headers={"x-decode-ms": f"{ms:.3f}"})
@app.get("/healthz")
def healthz() -> dict[str, str]:
return {"status": "ok"}
@app.post("/transcribe")
async def transcribe(file: UploadFile = File(...)):
return _timed(await file.read())
@app.post("/v1/audio/transcriptions")
async def openai_transcriptions(file: UploadFile = File(...)):
return _timed(await file.read())
+63
View File
@@ -0,0 +1,63 @@
# Parakeet ASR via NeMo torch (unified-en-0.6b, bf16 weights) + our own thin FastAPI wrapper.
#
# Replacement seat for stacks/parakeet (sherpa-onnx int8). Same port (:8300), same endpoints,
# same body — LiteLLM and `talk` need no change. Rollback: stop this container, start the old
# `parakeet` one (kept; container and image both intact).
#
# HOST: fv-ml1, GPU 0.
#
# ⚠ GPU 0 room came from gen-small's KV: its --gpu-memory-utilization dropped 0.48 -> 0.46
# (measured boot, 2026-09-30: the seat rests ~2.5 GB, served peak ~2.8 GB, load peak ~3.0 GB;
# the old seat held 1,690 MiB). Do not raise that util back without re-measuring nvidia-smi Free
# on GPU 0 — util does not predict resident VRAM (see the 09-15 note in gen-small-seat/.env).
#
# ⚠ Weights: /tank/aimodels/huggingface mounted READ-ONLY. nvidia/parakeet-unified-en-0.6b @
# fe53cd885760c96b6a5f51a0bfd362cb4584a98b. HF_HUB_OFFLINE=1 in the image: the seat never phones home.
#
# API (identical to the replaced seat):
# POST /transcribe — multipart file upload, returns {"text": "..."}
# POST /v1/audio/transcriptions — same body, OpenAI-compatible path alias
# GET /healthz
#
# All tunables live in .env — edit that, not this file.
services:
parakeet-nemo:
image: local/parakeet-nemo:${PARAKEET_NEMO_TAG}
container_name: parakeet-nemo
restart: unless-stopped
ports:
- "${PARAKEET_NEMO_BIND:-0.0.0.0}:${PARAKEET_NEMO_PORT}:8000"
environment:
- MODEL_PATH=/hf/hub/models--nvidia--parakeet-unified-en-0.6b/snapshots/${PARAKEET_NEMO_REV}/parakeet-unified-en-0.6b.nemo
- WARMUP_SECONDS=${PARAKEET_NEMO_WARMUP:-1,8,60}
- LOG_LEVEL=${PARAKEET_NEMO_LOG_LEVEL:-INFO}
volumes:
- /tank/aimodels/huggingface:/hf:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${PARAKEET_NEMO_GPU:-0}"]
capabilities: [gpu]
networks:
- tnet
healthcheck:
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8000/healthz || exit 1"]
interval: 30s
timeout: 10s
retries: 3
# Import-time model load + three warm-up decodes; no download (weights are mounted).
start_period: 240s
labels:
- homepage.group=AI - Audio Tools
- homepage.name=Parakeet ASR (NeMo)
- homepage.icon=mdi-microphone
- homepage.description=Parakeet-unified-en speech-to-text via NeMo bf16 (fv-ml1 GPU 0)
- homepage.href=http://10.251.50.54:${PARAKEET_NEMO_PORT}
networks:
tnet:
name: traefik-net
external: true