feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/ whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the 2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026 vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s (windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own long-form method), no pause truncation, no long-form dropout. GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40 refuse their boot check; cyberprev+voices hold the card). Its KV is byte- pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free. Two runtime landmines documented in the README: NeMo's attention mask is materialised T x T even under local attention (hence the window), and httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e. LiteLLM — rejects, so the image ships plain uvicorn with --http h11. Old seat stopped, not removed: docker stop parakeet-nemo && docker start parakeet is the rollback. License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in stacks/parakeet-nemo/README.md.
This commit is contained in:
@@ -0,0 +1,11 @@
|
||||
# Copy to .env next to compose.yaml on the host.
|
||||
PARAKEET_NEMO_TAG=nemo-0.1.0
|
||||
# Port the seat listens on. 8300 is the seat port LiteLLM's ext-stt/whisper-1 point at;
|
||||
# run acceptance on a temporary port first, then cut over by changing this line.
|
||||
PARAKEET_NEMO_PORT=8300
|
||||
# PARAKEET_NEMO_BIND=0.0.0.0
|
||||
# PARAKEET_NEMO_GPU=0
|
||||
# Pinned HF revision of nvidia/parakeet-unified-en-0.6b (sha256 ec23ed91... of the .nemo).
|
||||
PARAKEET_NEMO_REV=fe53cd885760c96b6a5f51a0bfd362cb4584a98b
|
||||
# Ascending silent warm-up clips in seconds (CUDA-graph capture + longest-shape kernel warm).
|
||||
# PARAKEET_NEMO_WARMUP=1,8,60
|
||||
@@ -0,0 +1,35 @@
|
||||
# Parakeet ASR seat: parakeet-unified-en-0.6b under NeMo torch, bf16 weights.
|
||||
# CUDA 12.8 runtime base + a uv-managed venv pinned to the A/B's proven stack
|
||||
# (torch 2.8 cu128, nemo_toolkit[asr]==3.0.0; the A/B found NeMo 2.7.3 lacks this encoder's
|
||||
# att_chunk_context_size, so 3.0.0 is a floor, not a preference).
|
||||
# Weights are NOT baked in: /tank/aimodels/huggingface is bind-mounted read-only (see compose).
|
||||
FROM nvidia/cuda:12.8.1-base-ubuntu24.04
|
||||
|
||||
ENV DEBIAN_FRONTEND=noninteractive \
|
||||
PIP_DISABLE_PIP_VERSION_CHECK=1 \
|
||||
PYTHONUNBUFFERED=1 \
|
||||
HF_HUB_OFFLINE=1
|
||||
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
python3 python3-venv python3-pip wget libsndfile1 ca-certificates \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
RUN python3 -m venv /opt/venv \
|
||||
&& /opt/venv/bin/pip install -q uv \
|
||||
&& UV_LINK_MODE=copy /opt/venv/bin/uv pip install -q --python /opt/venv/bin/python \
|
||||
--index-url https://download.pytorch.org/whl/cu128 \
|
||||
--extra-index-url https://pypi.org/simple \
|
||||
"torch==2.8.*" "torchaudio==2.8.*" "nemo_toolkit[asr]==3.0.0" \
|
||||
fastapi "uvicorn==0.53.0" python-multipart soundfile \
|
||||
&& /opt/venv/bin/python -c "import nemo, torch; print('nemo', nemo.__version__, 'torch', torch.__version__, 'cuda_ok', torch.cuda.is_available())"
|
||||
|
||||
WORKDIR /app
|
||||
COPY app.py /app/app.py
|
||||
|
||||
# The seat's only writable need is NeMo/HF scratch; keep it off the rootfs surprises.
|
||||
ENV HOME=/tmp
|
||||
EXPOSE 8000
|
||||
# --http h11: the [standard] extra pulls httptools, and httptools 0.8.0 writes a NUL into the
|
||||
# status line ("HTTP/1.1 200\x00OK") that h11/httpx reject. uvicorn auto-picks httptools when
|
||||
# importable, so it must stay UNinstalled and the flag must stay explicit. See README.
|
||||
CMD ["/opt/venv/bin/uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000", "--http", "h11"]
|
||||
@@ -0,0 +1,61 @@
|
||||
# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo)
|
||||
|
||||
STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`.
|
||||
Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order,
|
||||
after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs
|
||||
the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set).
|
||||
|
||||
## Why this runtime
|
||||
|
||||
The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The
|
||||
defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation
|
||||
after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model
|
||||
under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention
|
||||
long-audio mode (±128), which is what removes all three defects.
|
||||
|
||||
## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up")
|
||||
|
||||
- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the
|
||||
load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats.
|
||||
- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on
|
||||
the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips
|
||||
(`WARMUP_SECONDS`, default 1,8,60).
|
||||
- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory
|
||||
grows linearly in length instead of quadratically.
|
||||
- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode
|
||||
differently call to call.
|
||||
- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`.
|
||||
- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and
|
||||
auto-selected by uvicorn when importable) writes a NUL into the response status line —
|
||||
`HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`.
|
||||
curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven
|
||||
A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image
|
||||
installs plain `uvicorn==0.53.0` for exactly this reason.
|
||||
|
||||
## License
|
||||
|
||||
`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement**
|
||||
(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the
|
||||
CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file
|
||||
required; keep this section as the license note. Weights pinned at HF revision
|
||||
`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from
|
||||
`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`.
|
||||
|
||||
## GPU 0 room
|
||||
|
||||
The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over;
|
||||
see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46
|
||||
(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT
|
||||
predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before
|
||||
believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
|
||||
|
||||
## Rollback
|
||||
|
||||
The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet`
|
||||
restores the sherpa seat on :8300 exactly as before (container and image both kept).
|
||||
|
||||
## Deploy
|
||||
|
||||
Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag
|
||||
`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness
|
||||
and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete).
|
||||
@@ -0,0 +1,154 @@
|
||||
"""Parakeet ASR seat: nvidia/parakeet-unified-en-0.6b under NeMo torch (bf16 weights).
|
||||
|
||||
Fork of the A/B winner (services/parakeet-ab-2026-09-30 arm B-bf16w), with the three shipping
|
||||
changes that arm's doc called for and a longer warm-up. Same HTTP shape as the sherpa seat it
|
||||
replaces: model loaded at import, warm-up before traffic, async handlers over a serialised
|
||||
blocking decode, {"text": ...} responses on /transcribe and /v1/audio/transcriptions.
|
||||
|
||||
Hard-wired, because every one of these is load-bearing on GPU 0:
|
||||
bf16 weights — the encoder/decoder/joint are cast to bfloat16 (mel front end stays fp32).
|
||||
WER is identical to fp32 (399/400 utterances byte-equal in the A/B).
|
||||
CPU-then-cast — the .nemo restores on CPU, weights are cast to bf16 there, and only then move
|
||||
to the GPU. Restoring to CUDA spikes the load by ~1.5 GB; GPU 0 cannot absorb it.
|
||||
local attn — rel_pos_local_attn ±128 (NeMo's documented long-audio mode). This is what fixes
|
||||
the old seat's >400 s HTTP 500 and the long-form dropouts; memory grows linearly
|
||||
in file length. A 30-min file transcribes in ~2.6 s in one request (A/B, 2026-09-30).
|
||||
Warm-up runs ascending silent clips (1 s, 8 s, 60 s): the CUDA-graph greedy decoder and the
|
||||
encoder kernels cost ~330 ms extra on the first call at a new maximum length.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
|
||||
import numpy as np
|
||||
import soundfile as sf
|
||||
import torch
|
||||
from fastapi import FastAPI, File, HTTPException, UploadFile
|
||||
from fastapi.responses import JSONResponse
|
||||
|
||||
MODEL_PATH = os.environ["MODEL_PATH"]
|
||||
WARMUP_SECONDS = [int(x) for x in os.environ.get("WARMUP_SECONDS", "1,8,60").split(",")]
|
||||
SR = 16000
|
||||
|
||||
logger = logging.getLogger("parakeet-nemo")
|
||||
logging.basicConfig(level=os.environ.get("LOG_LEVEL", "INFO"))
|
||||
|
||||
|
||||
def _load():
|
||||
import nemo.collections.asr as nemo_asr
|
||||
from omegaconf import open_dict
|
||||
|
||||
t0 = time.monotonic()
|
||||
m = nemo_asr.models.ASRModel.restore_from(MODEL_PATH, map_location="cpu")
|
||||
m.eval()
|
||||
if m.cfg.get("validation_ds") is None: # the unified .nemo ships without it; transcribe() reads it
|
||||
with open_dict(m.cfg):
|
||||
m.cfg.validation_ds = {}
|
||||
d = m.cfg.decoding
|
||||
with open_dict(d):
|
||||
d.strategy = "greedy_batch"
|
||||
d.greedy["use_cuda_graph_decoder"] = True
|
||||
m.change_decoding_strategy(d, verbose=False)
|
||||
# transcribe() sets these on entry; the direct path must match, and must not dither (dither is
|
||||
# a training-time augmentation and makes the same file decode differently on each call).
|
||||
m.preprocessor.featurizer.dither = 0.0
|
||||
m.preprocessor.featurizer.pad_to = 0
|
||||
m.change_attention_model("rel_pos_local_attn", [128, 128])
|
||||
# bf16 the serving modules BEFORE the H2D copy: halves the transfer and skips the GPU-side
|
||||
# fp32->bf16 transient entirely (the measured load spike goes from 3,194 MiB to under 2,600).
|
||||
# ⚠ Must run AFTER change_attention_model: that call rebuilds the attention modules in fp32,
|
||||
# and casting first leaves fp32 islands behind (RuntimeError: mat1 and mat2 ... BFloat16/Float,
|
||||
# hit live at the first boot of this image, 2026-10-01).
|
||||
for mod in (m.encoder, m.decoder, m.joint):
|
||||
mod.to(torch.bfloat16)
|
||||
m = m.to("cuda")
|
||||
logger.info("loaded %s (%s) bf16w local_att=128,128 in %.1fs", os.path.basename(MODEL_PATH),
|
||||
type(m).__name__, time.monotonic() - t0)
|
||||
return m
|
||||
|
||||
|
||||
model = _load()
|
||||
|
||||
|
||||
def _hyp_text(h) -> str:
|
||||
if isinstance(h, str):
|
||||
return h
|
||||
t = getattr(h, "text", None)
|
||||
if isinstance(t, str):
|
||||
return t
|
||||
return model.tokenizer.ids_to_text([int(i) for i in h.y_sequence])
|
||||
|
||||
|
||||
@torch.inference_mode()
|
||||
def _infer(samples: np.ndarray) -> str:
|
||||
x = torch.from_numpy(samples).to("cuda", non_blocking=True).unsqueeze(0)
|
||||
xl = torch.tensor([x.shape[1]], device="cuda", dtype=torch.long)
|
||||
feats, fl = model.preprocessor(input_signal=x, length=xl)
|
||||
feats = feats.to(torch.bfloat16)
|
||||
enc, el = model.encoder(audio_signal=feats, length=fl)
|
||||
hyps = model.decoding.rnnt_decoder_predictions_tensor(encoder_output=enc, encoded_lengths=el,
|
||||
return_hypotheses=False)
|
||||
if isinstance(hyps, tuple):
|
||||
hyps = hyps[0]
|
||||
return _hyp_text(hyps[0])
|
||||
|
||||
|
||||
def _warm() -> None:
|
||||
for secs in WARMUP_SECONDS:
|
||||
t0 = time.monotonic()
|
||||
_infer(np.zeros(SR * secs, dtype=np.float32))
|
||||
torch.cuda.synchronize()
|
||||
logger.info("warmup %ss decode complete in %.1fs", secs, time.monotonic() - t0)
|
||||
|
||||
|
||||
_warm()
|
||||
app = FastAPI(title="Parakeet ASR (NeMo torch, bf16w)")
|
||||
|
||||
|
||||
def _decode(raw: bytes) -> str:
|
||||
try:
|
||||
samples, sample_rate = sf.read(io.BytesIO(raw), dtype="float32")
|
||||
except Exception as exc:
|
||||
raise HTTPException(400, f"Could not decode audio: {exc}") from exc
|
||||
if samples.ndim > 1:
|
||||
samples = samples.mean(axis=1).astype(np.float32)
|
||||
if sample_rate != SR:
|
||||
import torchaudio.functional as AF
|
||||
samples = AF.resample(torch.from_numpy(samples), sample_rate, SR).numpy()
|
||||
samples = np.ascontiguousarray(samples, dtype=np.float32)
|
||||
# Windowed long-form (> WINDOW_S): NeMo's attention mask is materialised T×T even under
|
||||
# rel_pos_local_attn, so one whole 12-min request wanted +1.1 GiB of scratch and OOMed on
|
||||
# GPU 0 (measured at acceptance, 2026-10-01). The A/B's own long-form arm used ~6-min
|
||||
# windows and lost zero clean speech on unified-en in 4/4 placements (doc § 5.4), so the
|
||||
# seat chunks at the same size: bounded memory, any length, no API change for callers.
|
||||
win = int(os.environ.get("WINDOW_S", "360")) * SR
|
||||
if len(samples) <= win:
|
||||
return _infer(samples)
|
||||
parts = [_infer(samples[i:i + win]) for i in range(0, len(samples), win)]
|
||||
return " ".join(p for p in parts if p)
|
||||
|
||||
|
||||
def _timed(raw: bytes) -> JSONResponse:
|
||||
t0 = time.perf_counter()
|
||||
text = _decode(raw)
|
||||
torch.cuda.synchronize()
|
||||
ms = (time.perf_counter() - t0) * 1000.0
|
||||
return JSONResponse({"text": text}, headers={"x-decode-ms": f"{ms:.3f}"})
|
||||
|
||||
|
||||
@app.get("/healthz")
|
||||
def healthz() -> dict[str, str]:
|
||||
return {"status": "ok"}
|
||||
|
||||
|
||||
@app.post("/transcribe")
|
||||
async def transcribe(file: UploadFile = File(...)):
|
||||
return _timed(await file.read())
|
||||
|
||||
|
||||
@app.post("/v1/audio/transcriptions")
|
||||
async def openai_transcriptions(file: UploadFile = File(...)):
|
||||
return _timed(await file.read())
|
||||
@@ -0,0 +1,63 @@
|
||||
# Parakeet ASR via NeMo torch (unified-en-0.6b, bf16 weights) + our own thin FastAPI wrapper.
|
||||
#
|
||||
# Replacement seat for stacks/parakeet (sherpa-onnx int8). Same port (:8300), same endpoints,
|
||||
# same body — LiteLLM and `talk` need no change. Rollback: stop this container, start the old
|
||||
# `parakeet` one (kept; container and image both intact).
|
||||
#
|
||||
# HOST: fv-ml1, GPU 0.
|
||||
#
|
||||
# ⚠ GPU 0 room came from gen-small's KV: its --gpu-memory-utilization dropped 0.48 -> 0.46
|
||||
# (measured boot, 2026-09-30: the seat rests ~2.5 GB, served peak ~2.8 GB, load peak ~3.0 GB;
|
||||
# the old seat held 1,690 MiB). Do not raise that util back without re-measuring nvidia-smi Free
|
||||
# on GPU 0 — util does not predict resident VRAM (see the 09-15 note in gen-small-seat/.env).
|
||||
#
|
||||
# ⚠ Weights: /tank/aimodels/huggingface mounted READ-ONLY. nvidia/parakeet-unified-en-0.6b @
|
||||
# fe53cd885760c96b6a5f51a0bfd362cb4584a98b. HF_HUB_OFFLINE=1 in the image: the seat never phones home.
|
||||
#
|
||||
# API (identical to the replaced seat):
|
||||
# POST /transcribe — multipart file upload, returns {"text": "..."}
|
||||
# POST /v1/audio/transcriptions — same body, OpenAI-compatible path alias
|
||||
# GET /healthz
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
services:
|
||||
parakeet-nemo:
|
||||
image: local/parakeet-nemo:${PARAKEET_NEMO_TAG}
|
||||
container_name: parakeet-nemo
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${PARAKEET_NEMO_BIND:-0.0.0.0}:${PARAKEET_NEMO_PORT}:8000"
|
||||
environment:
|
||||
- MODEL_PATH=/hf/hub/models--nvidia--parakeet-unified-en-0.6b/snapshots/${PARAKEET_NEMO_REV}/parakeet-unified-en-0.6b.nemo
|
||||
- WARMUP_SECONDS=${PARAKEET_NEMO_WARMUP:-1,8,60}
|
||||
- LOG_LEVEL=${PARAKEET_NEMO_LOG_LEVEL:-INFO}
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/hf:ro
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${PARAKEET_NEMO_GPU:-0}"]
|
||||
capabilities: [gpu]
|
||||
networks:
|
||||
- tnet
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8000/healthz || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# Import-time model load + three warm-up decodes; no download (weights are mounted).
|
||||
start_period: 240s
|
||||
labels:
|
||||
- homepage.group=AI - Audio Tools
|
||||
- homepage.name=Parakeet ASR (NeMo)
|
||||
- homepage.icon=mdi-microphone
|
||||
- homepage.description=Parakeet-unified-en speech-to-text via NeMo bf16 (fv-ml1 GPU 0)
|
||||
- homepage.href=http://10.251.50.54:${PARAKEET_NEMO_PORT}
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
Reference in New Issue
Block a user