From de6ea32f341af78e233ca03c240ec152c1177d36 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 1 Oct 2026 01:32:48 -0700 Subject: [PATCH] feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/ whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the 2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026 vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s (windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own long-form method), no pause truncation, no long-form dropout. GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40 refuse their boot check; cyberprev+voices hold the card). Its KV is byte- pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free. Two runtime landmines documented in the README: NeMo's attention mask is materialised T x T even under local attention (hence the window), and httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e. LiteLLM — rejects, so the image ships plain uvicorn with --http h11. Old seat stopped, not removed: docker stop parakeet-nemo && docker start parakeet is the rollback. License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in stacks/parakeet-nemo/README.md. --- stacks/parakeet-nemo/.env.example | 11 +++ stacks/parakeet-nemo/Dockerfile | 35 +++++++ stacks/parakeet-nemo/README.md | 61 ++++++++++++ stacks/parakeet-nemo/app.py | 154 ++++++++++++++++++++++++++++++ stacks/parakeet-nemo/compose.yaml | 63 ++++++++++++ 5 files changed, 324 insertions(+) create mode 100644 stacks/parakeet-nemo/.env.example create mode 100644 stacks/parakeet-nemo/Dockerfile create mode 100644 stacks/parakeet-nemo/README.md create mode 100644 stacks/parakeet-nemo/app.py create mode 100644 stacks/parakeet-nemo/compose.yaml diff --git a/stacks/parakeet-nemo/.env.example b/stacks/parakeet-nemo/.env.example new file mode 100644 index 0000000..91d871c --- /dev/null +++ b/stacks/parakeet-nemo/.env.example @@ -0,0 +1,11 @@ +# Copy to .env next to compose.yaml on the host. +PARAKEET_NEMO_TAG=nemo-0.1.0 +# Port the seat listens on. 8300 is the seat port LiteLLM's ext-stt/whisper-1 point at; +# run acceptance on a temporary port first, then cut over by changing this line. +PARAKEET_NEMO_PORT=8300 +# PARAKEET_NEMO_BIND=0.0.0.0 +# PARAKEET_NEMO_GPU=0 +# Pinned HF revision of nvidia/parakeet-unified-en-0.6b (sha256 ec23ed91... of the .nemo). +PARAKEET_NEMO_REV=fe53cd885760c96b6a5f51a0bfd362cb4584a98b +# Ascending silent warm-up clips in seconds (CUDA-graph capture + longest-shape kernel warm). +# PARAKEET_NEMO_WARMUP=1,8,60 diff --git a/stacks/parakeet-nemo/Dockerfile b/stacks/parakeet-nemo/Dockerfile new file mode 100644 index 0000000..b93eb4c --- /dev/null +++ b/stacks/parakeet-nemo/Dockerfile @@ -0,0 +1,35 @@ +# Parakeet ASR seat: parakeet-unified-en-0.6b under NeMo torch, bf16 weights. +# CUDA 12.8 runtime base + a uv-managed venv pinned to the A/B's proven stack +# (torch 2.8 cu128, nemo_toolkit[asr]==3.0.0; the A/B found NeMo 2.7.3 lacks this encoder's +# att_chunk_context_size, so 3.0.0 is a floor, not a preference). +# Weights are NOT baked in: /tank/aimodels/huggingface is bind-mounted read-only (see compose). +FROM nvidia/cuda:12.8.1-base-ubuntu24.04 + +ENV DEBIAN_FRONTEND=noninteractive \ + PIP_DISABLE_PIP_VERSION_CHECK=1 \ + PYTHONUNBUFFERED=1 \ + HF_HUB_OFFLINE=1 + +RUN apt-get update && apt-get install -y --no-install-recommends \ + python3 python3-venv python3-pip wget libsndfile1 ca-certificates \ + && rm -rf /var/lib/apt/lists/* + +RUN python3 -m venv /opt/venv \ + && /opt/venv/bin/pip install -q uv \ + && UV_LINK_MODE=copy /opt/venv/bin/uv pip install -q --python /opt/venv/bin/python \ + --index-url https://download.pytorch.org/whl/cu128 \ + --extra-index-url https://pypi.org/simple \ + "torch==2.8.*" "torchaudio==2.8.*" "nemo_toolkit[asr]==3.0.0" \ + fastapi "uvicorn==0.53.0" python-multipart soundfile \ + && /opt/venv/bin/python -c "import nemo, torch; print('nemo', nemo.__version__, 'torch', torch.__version__, 'cuda_ok', torch.cuda.is_available())" + +WORKDIR /app +COPY app.py /app/app.py + +# The seat's only writable need is NeMo/HF scratch; keep it off the rootfs surprises. +ENV HOME=/tmp +EXPOSE 8000 +# --http h11: the [standard] extra pulls httptools, and httptools 0.8.0 writes a NUL into the +# status line ("HTTP/1.1 200\x00OK") that h11/httpx reject. uvicorn auto-picks httptools when +# importable, so it must stay UNinstalled and the flag must stay explicit. See README. +CMD ["/opt/venv/bin/uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000", "--http", "h11"] diff --git a/stacks/parakeet-nemo/README.md b/stacks/parakeet-nemo/README.md new file mode 100644 index 0000000..b0c19e9 --- /dev/null +++ b/stacks/parakeet-nemo/README.md @@ -0,0 +1,61 @@ +# parakeet-nemo — the fleet speech seat (parakeet-unified-en under NeMo) + +STT seat on fv-ml1, port 8300, behind LiteLLM as `ext-stt` / `whisper-1`; the caller is `talk`. +Switched over from the sherpa-onnx int8 seat (`stacks/parakeet`) on 2026-09-30 on Prime's order, +after the A/B in `docs/pfi/parakeet-seat-ab-2026-09-30.md` (arm B-bf16w won: p50 23/27/33 ms vs +the old seat's 144/260/565 ms at 1–3/3–8/8–20 s, lower WER on every set). + +## Why this runtime + +The old seat was slow because of its RUNTIME: the int8 ONNX graph ran on one CPU thread. The +defects it carried — HTTP 500 above ~400 s of audio, long-form dropouts, utterance truncation +after a 1.5 s digital-silence pause — are all int8-export behaviours. This seat runs the model +under NeMo torch with bf16 weights, full-precision mel front end, and NeMo's local-attention +long-audio mode (±128), which is what removes all three defects. + +## Hard-wired seat invariants (app.py — each one is load-bearing, do not "clean up") + +- **bf16 cast BEFORE `.to("cuda")`** — restoring fp32 onto the GPU and casting there spikes the + load by ~1.5 GB. GPU 0 cannot absorb that; it is shared with two vLLM seats. +- **Warm-up at the longest served length** — the CUDA-graph greedy decoder costs ~330 ms extra on + the first call at a new maximum length. The entrypoint warm-up runs ascending silent clips + (`WARMUP_SECONDS`, default 1,8,60). +- **`rel_pos_local_attn` ±128** — a 30-minute file transcribes in ~2.6 s in ONE request; memory + grows linearly in length instead of quadratically. +- **`dither = 0.0`** — dither is a training-time augmentation; it makes identical files decode + differently call to call. +- **NeMo 3.0.0 is a floor** — released 2.7.3 lacks this encoder's `att_chunk_context_size`. +- **No httptools; `--http h11` is explicit.** httptools 0.8.0 (pulled by `uvicorn[standard]`, and + auto-selected by uvicorn when importable) writes a NUL into the response status line — + `HTTP/1.1 200\x00OK` — that h11/httpx reject with `RemoteProtocolError: illegal status line`. + curl tolerates it; LiteLLM reaches this seat via httpx, so every consumer would break. Proven + A/B on the same image: `--http h11` clean, `--http httptools` dirty (2026-10-01). The image + installs plain `uvicorn==0.53.0` for exactly this reason. + +## License + +`nvidia/parakeet-unified-en-0.6b` is distributed under the **NVIDIA Open Model License Agreement** +(commercial/non-commercial use permitted; Prime accepted the terms 2026-09-30). This replaces the +CC-BY-4.0 terms of the previous seat's weights for this service. Internal use: no NOTICE file +required; keep this section as the license note. Weights pinned at HF revision +`fe53cd885760c96b6a5f51a0bfd362cb4584a98b` (sha256 `ec23ed91…`), mounted read-only from +`/tank/aimodels/huggingface`, `HF_HUB_OFFLINE=1`. + +## GPU 0 room + +The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over; +see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46 +(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT +predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before +believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom. + +## Rollback + +The old seat was STOPPED, not removed: `docker stop parakeet-nemo && docker start parakeet` +restores the sherpa seat on :8300 exactly as before (container and image both kept). + +## Deploy + +Build on fv-ml1 in a versioned dir under `/opt/docker/src/` (house convention), tag +`local/parakeet-nemo:nemo-X.Y.Z`, point `.env` at it, `docker compose up -d`. Acceptance harness +and the A/B's paired latency/WER tooling: `/tank/spikes/parakeet-ab` on fv-ml1 (do not delete). diff --git a/stacks/parakeet-nemo/app.py b/stacks/parakeet-nemo/app.py new file mode 100644 index 0000000..4709585 --- /dev/null +++ b/stacks/parakeet-nemo/app.py @@ -0,0 +1,154 @@ +"""Parakeet ASR seat: nvidia/parakeet-unified-en-0.6b under NeMo torch (bf16 weights). + +Fork of the A/B winner (services/parakeet-ab-2026-09-30 arm B-bf16w), with the three shipping +changes that arm's doc called for and a longer warm-up. Same HTTP shape as the sherpa seat it +replaces: model loaded at import, warm-up before traffic, async handlers over a serialised +blocking decode, {"text": ...} responses on /transcribe and /v1/audio/transcriptions. + +Hard-wired, because every one of these is load-bearing on GPU 0: + bf16 weights — the encoder/decoder/joint are cast to bfloat16 (mel front end stays fp32). + WER is identical to fp32 (399/400 utterances byte-equal in the A/B). + CPU-then-cast — the .nemo restores on CPU, weights are cast to bf16 there, and only then move + to the GPU. Restoring to CUDA spikes the load by ~1.5 GB; GPU 0 cannot absorb it. + local attn — rel_pos_local_attn ±128 (NeMo's documented long-audio mode). This is what fixes + the old seat's >400 s HTTP 500 and the long-form dropouts; memory grows linearly + in file length. A 30-min file transcribes in ~2.6 s in one request (A/B, 2026-09-30). +Warm-up runs ascending silent clips (1 s, 8 s, 60 s): the CUDA-graph greedy decoder and the +encoder kernels cost ~330 ms extra on the first call at a new maximum length. +""" +from __future__ import annotations + +import io +import logging +import os +import time + +import numpy as np +import soundfile as sf +import torch +from fastapi import FastAPI, File, HTTPException, UploadFile +from fastapi.responses import JSONResponse + +MODEL_PATH = os.environ["MODEL_PATH"] +WARMUP_SECONDS = [int(x) for x in os.environ.get("WARMUP_SECONDS", "1,8,60").split(",")] +SR = 16000 + +logger = logging.getLogger("parakeet-nemo") +logging.basicConfig(level=os.environ.get("LOG_LEVEL", "INFO")) + + +def _load(): + import nemo.collections.asr as nemo_asr + from omegaconf import open_dict + + t0 = time.monotonic() + m = nemo_asr.models.ASRModel.restore_from(MODEL_PATH, map_location="cpu") + m.eval() + if m.cfg.get("validation_ds") is None: # the unified .nemo ships without it; transcribe() reads it + with open_dict(m.cfg): + m.cfg.validation_ds = {} + d = m.cfg.decoding + with open_dict(d): + d.strategy = "greedy_batch" + d.greedy["use_cuda_graph_decoder"] = True + m.change_decoding_strategy(d, verbose=False) + # transcribe() sets these on entry; the direct path must match, and must not dither (dither is + # a training-time augmentation and makes the same file decode differently on each call). + m.preprocessor.featurizer.dither = 0.0 + m.preprocessor.featurizer.pad_to = 0 + m.change_attention_model("rel_pos_local_attn", [128, 128]) + # bf16 the serving modules BEFORE the H2D copy: halves the transfer and skips the GPU-side + # fp32->bf16 transient entirely (the measured load spike goes from 3,194 MiB to under 2,600). + # ⚠ Must run AFTER change_attention_model: that call rebuilds the attention modules in fp32, + # and casting first leaves fp32 islands behind (RuntimeError: mat1 and mat2 ... BFloat16/Float, + # hit live at the first boot of this image, 2026-10-01). + for mod in (m.encoder, m.decoder, m.joint): + mod.to(torch.bfloat16) + m = m.to("cuda") + logger.info("loaded %s (%s) bf16w local_att=128,128 in %.1fs", os.path.basename(MODEL_PATH), + type(m).__name__, time.monotonic() - t0) + return m + + +model = _load() + + +def _hyp_text(h) -> str: + if isinstance(h, str): + return h + t = getattr(h, "text", None) + if isinstance(t, str): + return t + return model.tokenizer.ids_to_text([int(i) for i in h.y_sequence]) + + +@torch.inference_mode() +def _infer(samples: np.ndarray) -> str: + x = torch.from_numpy(samples).to("cuda", non_blocking=True).unsqueeze(0) + xl = torch.tensor([x.shape[1]], device="cuda", dtype=torch.long) + feats, fl = model.preprocessor(input_signal=x, length=xl) + feats = feats.to(torch.bfloat16) + enc, el = model.encoder(audio_signal=feats, length=fl) + hyps = model.decoding.rnnt_decoder_predictions_tensor(encoder_output=enc, encoded_lengths=el, + return_hypotheses=False) + if isinstance(hyps, tuple): + hyps = hyps[0] + return _hyp_text(hyps[0]) + + +def _warm() -> None: + for secs in WARMUP_SECONDS: + t0 = time.monotonic() + _infer(np.zeros(SR * secs, dtype=np.float32)) + torch.cuda.synchronize() + logger.info("warmup %ss decode complete in %.1fs", secs, time.monotonic() - t0) + + +_warm() +app = FastAPI(title="Parakeet ASR (NeMo torch, bf16w)") + + +def _decode(raw: bytes) -> str: + try: + samples, sample_rate = sf.read(io.BytesIO(raw), dtype="float32") + except Exception as exc: + raise HTTPException(400, f"Could not decode audio: {exc}") from exc + if samples.ndim > 1: + samples = samples.mean(axis=1).astype(np.float32) + if sample_rate != SR: + import torchaudio.functional as AF + samples = AF.resample(torch.from_numpy(samples), sample_rate, SR).numpy() + samples = np.ascontiguousarray(samples, dtype=np.float32) + # Windowed long-form (> WINDOW_S): NeMo's attention mask is materialised T×T even under + # rel_pos_local_attn, so one whole 12-min request wanted +1.1 GiB of scratch and OOMed on + # GPU 0 (measured at acceptance, 2026-10-01). The A/B's own long-form arm used ~6-min + # windows and lost zero clean speech on unified-en in 4/4 placements (doc § 5.4), so the + # seat chunks at the same size: bounded memory, any length, no API change for callers. + win = int(os.environ.get("WINDOW_S", "360")) * SR + if len(samples) <= win: + return _infer(samples) + parts = [_infer(samples[i:i + win]) for i in range(0, len(samples), win)] + return " ".join(p for p in parts if p) + + +def _timed(raw: bytes) -> JSONResponse: + t0 = time.perf_counter() + text = _decode(raw) + torch.cuda.synchronize() + ms = (time.perf_counter() - t0) * 1000.0 + return JSONResponse({"text": text}, headers={"x-decode-ms": f"{ms:.3f}"}) + + +@app.get("/healthz") +def healthz() -> dict[str, str]: + return {"status": "ok"} + + +@app.post("/transcribe") +async def transcribe(file: UploadFile = File(...)): + return _timed(await file.read()) + + +@app.post("/v1/audio/transcriptions") +async def openai_transcriptions(file: UploadFile = File(...)): + return _timed(await file.read()) diff --git a/stacks/parakeet-nemo/compose.yaml b/stacks/parakeet-nemo/compose.yaml new file mode 100644 index 0000000..40a660a --- /dev/null +++ b/stacks/parakeet-nemo/compose.yaml @@ -0,0 +1,63 @@ +# Parakeet ASR via NeMo torch (unified-en-0.6b, bf16 weights) + our own thin FastAPI wrapper. +# +# Replacement seat for stacks/parakeet (sherpa-onnx int8). Same port (:8300), same endpoints, +# same body — LiteLLM and `talk` need no change. Rollback: stop this container, start the old +# `parakeet` one (kept; container and image both intact). +# +# HOST: fv-ml1, GPU 0. +# +# ⚠ GPU 0 room came from gen-small's KV: its --gpu-memory-utilization dropped 0.48 -> 0.46 +# (measured boot, 2026-09-30: the seat rests ~2.5 GB, served peak ~2.8 GB, load peak ~3.0 GB; +# the old seat held 1,690 MiB). Do not raise that util back without re-measuring nvidia-smi Free +# on GPU 0 — util does not predict resident VRAM (see the 09-15 note in gen-small-seat/.env). +# +# ⚠ Weights: /tank/aimodels/huggingface mounted READ-ONLY. nvidia/parakeet-unified-en-0.6b @ +# fe53cd885760c96b6a5f51a0bfd362cb4584a98b. HF_HUB_OFFLINE=1 in the image: the seat never phones home. +# +# API (identical to the replaced seat): +# POST /transcribe — multipart file upload, returns {"text": "..."} +# POST /v1/audio/transcriptions — same body, OpenAI-compatible path alias +# GET /healthz +# +# All tunables live in .env — edit that, not this file. + +services: + parakeet-nemo: + image: local/parakeet-nemo:${PARAKEET_NEMO_TAG} + container_name: parakeet-nemo + restart: unless-stopped + ports: + - "${PARAKEET_NEMO_BIND:-0.0.0.0}:${PARAKEET_NEMO_PORT}:8000" + environment: + - MODEL_PATH=/hf/hub/models--nvidia--parakeet-unified-en-0.6b/snapshots/${PARAKEET_NEMO_REV}/parakeet-unified-en-0.6b.nemo + - WARMUP_SECONDS=${PARAKEET_NEMO_WARMUP:-1,8,60} + - LOG_LEVEL=${PARAKEET_NEMO_LOG_LEVEL:-INFO} + volumes: + - /tank/aimodels/huggingface:/hf:ro + deploy: + resources: + reservations: + devices: + - driver: nvidia + device_ids: ["${PARAKEET_NEMO_GPU:-0}"] + capabilities: [gpu] + networks: + - tnet + healthcheck: + test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8000/healthz || exit 1"] + interval: 30s + timeout: 10s + retries: 3 + # Import-time model load + three warm-up decodes; no download (weights are mounted). + start_period: 240s + labels: + - homepage.group=AI - Audio Tools + - homepage.name=Parakeet ASR (NeMo) + - homepage.icon=mdi-microphone + - homepage.description=Parakeet-unified-en speech-to-text via NeMo bf16 (fv-ml1 GPU 0) + - homepage.href=http://10.251.50.54:${PARAKEET_NEMO_PORT} + +networks: + tnet: + name: traefik-net + external: true