feat(parakeet): stand up Parakeet STT on fv-ml1 GPU 3 + LiteLLM ext-stt/whisper-1
Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry the vLLM seats at 84-95.5 GB of 96. Changes: - compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used `count: all`, which would have handed a 0.6B ASR seat all four cards); join traefik-net; port 8300; homepage href to the live FV address. - .env.example: default to the v3 int8 model (25 European languages, 464 MiB) rather than English-only v2; models to /tank/parakeet/models. - app.py: warm the recognizer at startup before uvicorn accepts traffic. The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at 45.1s on a second container) against ~0.50s warm. A 45s first request is indistinguishable from a hang and LiteLLM's default timeout abandons it long before it returns. Decoding 1s of silence at load moves the cost inside the healthcheck's 300s start_period; first real request after restart is now 0.65s. Verification, because "provider=cuda" in the log is only an echo of the env var: ORT falls back to CPU silently and still returns correct text, so the service being up and the transcript being right establishes nothing. The discriminator is a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS sentence transcribes near-exactly (positive), 3s of digital silence returns empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread 0.47-0.65s, single-stream, one clip: a smoke measurement with its harness stated, not a benchmark. Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1` (OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's Postgres store where the ext-tts family already lives — no gateway restart, and config.yaml is consequently not a complete picture of what the gateway serves. Both verified end to end. The aliases use a raw IP deliberately: ana-docker resolves no .internal names at all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a hand-pinned extra_hosts entry. A second hosts entry would mean recreating the container and bouncing the gateway for every consumer. Also records the svos_miranda plugin validation pass and its structural findings, and notes that the irv-ml1 parakeet is still running — there are two now, and retiring the old one is the operator's call.
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
# Parakeet ASR stack tunables. Copy to `.env` on irv-ml1 before deploying.
|
||||
# Parakeet ASR stack tunables. Copy to `.env` on fv-ml1 before deploying.
|
||||
#
|
||||
# cp .env.example .env
|
||||
# # edit as needed
|
||||
@@ -7,29 +7,42 @@
|
||||
|
||||
# Image tag. Bump when you change the Dockerfile / app.py so docker caches
|
||||
# cleanly.
|
||||
PARAKEET_TAG=sherpa-onnx-v2
|
||||
PARAKEET_TAG=sherpa-onnx-v4
|
||||
|
||||
# Host port for the FastAPI server (container listens on 8000)
|
||||
PARAKEET_PORT=8765
|
||||
# Which GPU to pin. fv-ml1 GPU 3 is the utility card — 0/1/2 carry the vLLM
|
||||
# serving seats and sit at 85-98% VRAM, so this is the only one with room.
|
||||
# The container sees whichever card this names as cuda:0 internally.
|
||||
PARAKEET_GPU=3
|
||||
|
||||
# Bind address. 0.0.0.0 exposes on all interfaces including the WG tunnel IP
|
||||
# (10.100.79.3). Use 127.0.0.1 to restrict to local-only.
|
||||
# Host port for the FastAPI server (container listens on 8000). 8300 is
|
||||
# fv-ml1's established parakeet port; the 80xx range belongs to the vLLM seats.
|
||||
PARAKEET_PORT=8300
|
||||
|
||||
# Bind address. 0.0.0.0 exposes on all interfaces. Use 127.0.0.1 to restrict
|
||||
# to local-only — but LiteLLM on ana-docker reaches this over the LAN, so it
|
||||
# has to be 0.0.0.0 for the gateway alias to work.
|
||||
PARAKEET_BIND=0.0.0.0
|
||||
|
||||
# Host path for the ONNX model files — encoder/decoder/joiner/tokens.txt.
|
||||
# Downloaded by the entrypoint on first run if absent. Must exist before
|
||||
# first `up` (directory, not files).
|
||||
PARAKEET_MODELS_DIR=/worktank/parakeet/models
|
||||
# first `up` (directory, not files). Regenerable — exclude from restic.
|
||||
PARAKEET_MODELS_DIR=/tank/parakeet/models
|
||||
|
||||
# Which sherpa-onnx release tarball to fetch on first boot. Default is the
|
||||
# int8-quantized English-only v2 (~400 MB). Switch to the v3 tarball below
|
||||
# to cover 25 European languages at a similar size:
|
||||
# https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8.tar.bz2
|
||||
PARAKEET_MODEL_URL=https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8.tar.bz2
|
||||
# Which sherpa-onnx release tarball to fetch on first boot.
|
||||
# v3 (default, 464 MiB) — 25 European languages
|
||||
# v2 — English only, swap the URL below
|
||||
# https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8.tar.bz2
|
||||
PARAKEET_MODEL_URL=https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8.tar.bz2
|
||||
|
||||
# ONNX Runtime execution provider. `cuda` uses the GPU (requires nvidia
|
||||
# runtime + matching CUDA/cuDNN in the image). `cpu` falls back to CPU —
|
||||
# fine for low-volume dev use; ~4-8× slower on this host.
|
||||
# ONNX Runtime execution provider. `cuda` uses the GPU (requires the nvidia
|
||||
# container runtime + matching CUDA/cuDNN in the image). `cpu` falls back to
|
||||
# CPU.
|
||||
#
|
||||
# ⚠ ORT's CUDA EP FALLS BACK TO CPU SILENTLY when it cannot initialise — the
|
||||
# server still answers 200 and still returns correct text, just slowly. So
|
||||
# `PROVIDER=cuda` is a REQUEST, not a guarantee, and the only honest check is
|
||||
# to watch `nvidia-smi` during a transcription and confirm a process appears on
|
||||
# the pinned card. See README § Verifying the GPU is actually in use.
|
||||
PARAKEET_PROVIDER=cuda
|
||||
|
||||
# CPU threads per recognizer session. Irrelevant when provider=cuda;
|
||||
|
||||
+90
-63
@@ -4,12 +4,26 @@ NVIDIA Parakeet-TDT 0.6B (int8 ONNX) served by our own thin FastAPI
|
||||
wrapper over [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)
|
||||
(ONNX Runtime + CUDA).
|
||||
|
||||
**Server:** irv-ml1 (Irvine, WireGuard-only)
|
||||
**Port:** 8765 (container 8000)
|
||||
**GPU:** both exposed (`NVIDIA_VISIBLE_DEVICES=all`); sherpa-onnx uses
|
||||
whichever CUDA ExecutionProvider picks
|
||||
**Image:** `local/parakeet:sherpa-onnx-v1` — built from `Dockerfile` +
|
||||
**Server:** fv-ml1 (Fountain Valley, `10.251.50.54`) — moved from irv-ml1 2026-09-15
|
||||
**Port:** 8300 (container 8000)
|
||||
**GPU:** **3**, pinned explicitly via `device_ids` — the utility card
|
||||
**Image:** `local/parakeet:sherpa-onnx-v4` — built from `Dockerfile` +
|
||||
`app.py` + `entrypoint.sh` in this directory; **we own all the code**
|
||||
**Model:** `parakeet-tdt-0.6b-v3` int8, 25 European languages (~464 MiB)
|
||||
|
||||
## Why GPU 3
|
||||
|
||||
fv-ml1 has four RTX PRO 6000 Blackwell Max-Q (96 GB each). Three carry the
|
||||
vLLM serving seats and run 85–98 % full; GPU 3 is the utility card and was
|
||||
empty (2 MiB) at placement time. A 0.6 B int8 ASR model is a rounding error
|
||||
next to those seats, but it still has to go somewhere that is not fighting
|
||||
them for VRAM.
|
||||
|
||||
⚠ The pin is `deploy.resources.reservations.devices[].device_ids: ["3"]`,
|
||||
the fleet convention — **not** `count: all`, which is what the dead on-host
|
||||
stub used and which would have handed this seat all four cards. Inside the
|
||||
container the pinned card presents as `cuda:0`, which is what sherpa-onnx's
|
||||
CUDA execution provider takes by default.
|
||||
|
||||
## Why not the FastAPI community wrappers
|
||||
|
||||
@@ -35,80 +49,93 @@ recognizer API is a three-line call.
|
||||
|
||||
| Host path | Container path | Purpose | Restic? |
|
||||
|---|---|---|---|
|
||||
| `/worktank/parakeet/models/` | `/models` | ONNX encoder+decoder+joiner+tokens (~400 MB int8) | excluded (regenerable — re-downloads from the URL on first run if absent) |
|
||||
| `/tank/parakeet/models/` | `/models` | ONNX encoder+decoder+joiner+tokens (~464 MiB int8) | excluded (regenerable — re-downloads from the URL on first run if absent) |
|
||||
|
||||
## First-time deploy on irv-ml1
|
||||
## ⚠ Verifying the GPU is actually in use
|
||||
|
||||
**ONNX Runtime's CUDA execution provider falls back to CPU silently.** It logs
|
||||
a warning if you are looking, keeps the process alive, answers `200`, and
|
||||
returns *correct transcriptions* — just far slower. So `PROVIDER=cuda` in
|
||||
`.env` is a request, not a guarantee, and "the service is up and the text is
|
||||
right" does **not** establish that the GPU is doing the work.
|
||||
|
||||
Blackwell is the reason this matters here rather than being pedantry: these
|
||||
cards are `sm_120`, newer than the compute capabilities ORT's prebuilt CUDA
|
||||
binaries have historically shipped kernels for, and the irv-ml1 host this
|
||||
stack came from was Ampere `sm_86`. The move is exactly the kind that turns a
|
||||
green service into a CPU service without a single error.
|
||||
|
||||
The honest check is to watch the card while a transcription runs:
|
||||
|
||||
```bash
|
||||
# 1. Push compose + Dockerfile + app + entrypoint
|
||||
scripts/deploy-stack.sh irv-ml1 parakeet
|
||||
# on fv-ml1 — terminal 1
|
||||
watch -n0.2 'nvidia-smi --query-compute-apps=pid,process_name,used_memory \
|
||||
--format=csv -i 3'
|
||||
|
||||
# 2. Make sure the models dir exists (one-time, already done from the
|
||||
# earlier Shadowfita deploy; this is idempotent)
|
||||
ssh -t irv-ml1 'sudo mkdir -p /worktank/parakeet/models && \
|
||||
sudo chown -R lkraven:lkraven /worktank/parakeet'
|
||||
|
||||
# 3. Build the image and bring up. First boot does a ~400 MB model
|
||||
# download via the entrypoint; allow 1–2 minutes before /healthz
|
||||
# flips healthy.
|
||||
ssh irv-ml1 '
|
||||
cd /opt/docker/compose/parakeet && \
|
||||
cp -n .env.example .env && \
|
||||
docker compose config >/dev/null && \
|
||||
docker compose build && \
|
||||
docker compose up -d && \
|
||||
docker compose logs -f --tail=30
|
||||
'
|
||||
# terminal 2 — send real audio, not silence
|
||||
curl -s -F file=@sample.wav http://127.0.0.1:8300/v1/audio/transcriptions
|
||||
```
|
||||
|
||||
## Smoke test
|
||||
A process must appear **on GPU 3** for the duration. If GPU 3 stays empty, the
|
||||
CUDA EP did not initialise and you are on CPU regardless of what `.env` says.
|
||||
Confirm with the container's own startup log, which names the providers ORT
|
||||
actually registered:
|
||||
|
||||
```bash
|
||||
# Over WG from the workstation
|
||||
curl -F "file=@sample.wav" http://10.100.79.3:8765/transcribe
|
||||
# → {"text": "hello world"}
|
||||
|
||||
# OpenAI-shape alias (for clients that only know /v1/audio/transcriptions)
|
||||
curl -F "file=@sample.wav" http://10.100.79.3:8765/v1/audio/transcriptions
|
||||
docker logs parakeet 2>&1 | grep -i 'provider\|cuda\|onnxruntime'
|
||||
```
|
||||
|
||||
## Switching to the v3 (multilingual) model
|
||||
Timing alone is **not** sufficient evidence either way: the int8 model is fast
|
||||
enough on a 96-thread EPYC that a CPU fallback still looks brisk on short
|
||||
clips. Use the process check as the discriminator and treat throughput as a
|
||||
secondary signal.
|
||||
|
||||
The env var `PARAKEET_MODEL_URL` picks the release tarball. To swap
|
||||
from the English-only v2 to the 25-language v3:
|
||||
## LiteLLM alias
|
||||
|
||||
Reached fleet-wide through the gateway rather than by name, engine-neutral so
|
||||
the backend can be swapped without touching consumers — the same pattern as
|
||||
`ext-tts`:
|
||||
|
||||
| alias | mode | backend |
|
||||
|---|---|---|
|
||||
| `ext-stt` | `audio_transcription` | `http://10.251.50.54:8300/v1` |
|
||||
| `whisper-1` | `audio_transcription` | same — OpenAI-compatible name so stock SDK clients work unchanged |
|
||||
|
||||
⚠ **The alias uses a raw IP on purpose.** `ana-docker` (where LiteLLM runs)
|
||||
resolves no `.internal` names at all — its `/etc/resolv.conf` points at
|
||||
`1.1.1.1`/`1.0.0.1`, and the only reason the `ext-tts` backend resolves is a
|
||||
hand-pinned `extra_hosts: irv-ml1.nh3.internal:10.6.110.50` in the LiteLLM
|
||||
compose. Adding a second hosts entry would mean recreating the container and
|
||||
bouncing the gateway for every consumer; an IP costs nothing and cannot go
|
||||
stale silently. See the DNS follow-up in `persistent-memory.md`.
|
||||
|
||||
Aliases live in LiteLLM's **Postgres store** (`store_model_in_db: true`), not
|
||||
in `config.yaml` — that is where the `ext-tts` family lives too, and it means
|
||||
adding one needs no gateway restart. It also means `config.yaml` is not a
|
||||
complete picture of what the gateway serves: check `/v1/models` or
|
||||
`/model/info`, never just the file.
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
ssh irv-ml1 '
|
||||
cd /opt/docker/compose/parakeet && \
|
||||
sed -i "s|v2-int8|v3-int8|" .env && \
|
||||
# Wipe the v2 weights so the entrypoint re-downloads v3 on next up:
|
||||
rm -f /worktank/parakeet/models/*.onnx /worktank/parakeet/models/tokens.txt && \
|
||||
docker compose up -d && \
|
||||
docker compose logs -f --tail=30
|
||||
'
|
||||
scripts/deploy-stack.sh fv-ml1 parakeet --compose
|
||||
# on fv-ml1, first time only:
|
||||
# cp .env.example .env # then edit
|
||||
docker compose -f /opt/docker/compose/parakeet/compose.yaml build
|
||||
docker compose -f /opt/docker/compose/parakeet/compose.yaml up -d
|
||||
```
|
||||
|
||||
## Upgrade sherpa-onnx or change the base image
|
||||
First boot downloads ~464 MiB of ONNX weights into `/tank/parakeet/models/`;
|
||||
the healthcheck's `start_period` is 300 s to cover it. Subsequent starts skip
|
||||
the download.
|
||||
|
||||
Bump `PARAKEET_TAG` in `.env` to force a rebuild of the local image
|
||||
after editing the `Dockerfile`, then:
|
||||
## Switching model variants
|
||||
|
||||
```bash
|
||||
scripts/deploy-stack.sh irv-ml1 parakeet
|
||||
ssh irv-ml1 'cd /opt/docker/compose/parakeet && docker compose build && docker compose up -d'
|
||||
```
|
||||
One line in `.env`, then `docker compose up -d` (not `restart` — the model URL
|
||||
is read by the entrypoint at container creation) and delete the old files from
|
||||
`/tank/parakeet/models/` so the entrypoint re-downloads:
|
||||
|
||||
Model files under `/worktank/parakeet/models/` are preserved across
|
||||
image rebuilds.
|
||||
- **v3** (default) — 25 European languages
|
||||
- **v2** — English only; slightly better on English-only material
|
||||
|
||||
## File layout
|
||||
|
||||
```
|
||||
stacks/parakeet/
|
||||
├── Dockerfile # CUDA 12.8 + cuDNN 9 base, sherpa-onnx-cu12 wheel
|
||||
├── app.py # FastAPI — ~60 lines
|
||||
├── entrypoint.sh # downloads model on first run, then uvicorn
|
||||
├── compose.yaml # one service, bind-mounts the models dir
|
||||
├── .env.example # template; real .env lives on the server
|
||||
└── README.md # this file
|
||||
```
|
||||
Both are ~460 MiB int8 tarballs from the same k2-fsa release page.
|
||||
|
||||
+30
-1
@@ -6,7 +6,7 @@ Load the encoder/decoder/joiner/tokens once at startup; serve:
|
||||
GET /healthz — used by the docker healthcheck
|
||||
|
||||
No VAD chunking, no Silero preprocessing — parakeet-tdt handles long-form natively
|
||||
and the int8 ONNX model on a 24 GB GPU eats everything we're likely to throw at it.
|
||||
and the int8 ONNX model is a rounding error against this host's 96 GB cards.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -14,6 +14,7 @@ from __future__ import annotations
|
||||
import io
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
@@ -59,8 +60,36 @@ def _load_recognizer() -> sherpa_onnx.OfflineRecognizer:
|
||||
)
|
||||
|
||||
|
||||
def _warm(rec: "sherpa_onnx.OfflineRecognizer") -> None:
|
||||
"""Decode one throwaway buffer before the server accepts traffic.
|
||||
|
||||
⚠ NOT an optimisation — it moves a 45 s stall out of the first real request.
|
||||
ONNX Runtime's CUDA EP compiles and autotunes its kernels lazily, on the first
|
||||
decode, and on this host (RTX PRO 6000 Blackwell, sm_120) that measured **45.7 s**
|
||||
while every subsequent call was ~0.48 s. Without this, the first caller after any
|
||||
container restart sees a 45 s hang and most clients — LiteLLM's default request
|
||||
timeout included — give up long before it returns, which reads as "the service is
|
||||
broken" rather than "the service is warming".
|
||||
|
||||
The healthcheck's `start_period` (300 s) is what makes paying it here safe.
|
||||
"""
|
||||
try:
|
||||
t0 = time.monotonic()
|
||||
stream = rec.create_stream()
|
||||
# 1 s of silence at 16 kHz: enough to force the full encoder/decoder/joiner
|
||||
# path to compile, cheap enough not to matter.
|
||||
stream.accept_waveform(16000, np.zeros(16000, dtype=np.float32))
|
||||
rec.decode_stream(stream)
|
||||
logger.info("warmup decode complete in %.1fs — CUDA kernels compiled", time.monotonic() - t0)
|
||||
except Exception:
|
||||
# A failed warmup must not stop the server: the model is loaded and real
|
||||
# requests would still work, just with the stall back on the first caller.
|
||||
logger.exception("warmup decode failed; first real request will absorb the stall")
|
||||
|
||||
|
||||
app = FastAPI(title="Parakeet ASR (sherpa-onnx)")
|
||||
recognizer = _load_recognizer()
|
||||
_warm(recognizer)
|
||||
|
||||
|
||||
def _decode(raw: bytes) -> str:
|
||||
|
||||
@@ -6,7 +6,18 @@
|
||||
# prebuilt int8 quantized Parakeet-TDT from k2-fsa — and wrote our own ~50-line
|
||||
# wrapper we own end-to-end.
|
||||
#
|
||||
# Model weights (~400 MB int8) download on first run via the entrypoint to
|
||||
# HOST: fv-ml1, GPU 3 (relocated from irv-ml1 2026-09-15). GPU 3 is the utility
|
||||
# card — the other three carry the vLLM serving seats and run 85-98% full, so a
|
||||
# seat placed anywhere else would fight them for VRAM.
|
||||
#
|
||||
# ⚠ GPU pin is `deploy.resources.reservations.devices[].device_ids`, the fleet
|
||||
# convention — NOT `runtime: nvidia` + NVIDIA_VISIBLE_DEVICES, and NOT
|
||||
# `count: all` (which is what the dead on-host stub did, and would have let this
|
||||
# tiny ASR seat see all four cards including the three that are full).
|
||||
# device_ids ["3"] presents that card as cuda:0 INSIDE the container, which is
|
||||
# what sherpa-onnx's CUDAExecutionProvider takes by default.
|
||||
#
|
||||
# Model weights (~460 MB int8) download on first run via the entrypoint to
|
||||
# ${PARAKEET_MODELS_DIR}/ (persistent host bind mount). Subsequent starts skip
|
||||
# the download.
|
||||
#
|
||||
@@ -25,11 +36,9 @@ services:
|
||||
dockerfile: Dockerfile
|
||||
container_name: parakeet
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
ports:
|
||||
- "${PARAKEET_BIND:-0.0.0.0}:${PARAKEET_PORT}:8000"
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=0
|
||||
- MODEL_DIR=/models
|
||||
- MODEL_URL=${PARAKEET_MODEL_URL}
|
||||
- PROVIDER=${PARAKEET_PROVIDER:-cuda}
|
||||
@@ -37,17 +46,31 @@ services:
|
||||
- LOG_LEVEL=${PARAKEET_LOG_LEVEL:-INFO}
|
||||
volumes:
|
||||
- ${PARAKEET_MODELS_DIR}:/models
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${PARAKEET_GPU:-3}"]
|
||||
capabilities: [gpu]
|
||||
networks:
|
||||
- tnet
|
||||
healthcheck:
|
||||
# Image ships wget (apt) but not curl — use wget so the check actually runs.
|
||||
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8000/healthz || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# First boot may include a ~400 MB model download.
|
||||
# First boot may include a ~460 MB model download.
|
||||
start_period: 300s
|
||||
labels:
|
||||
- homepage.group=AI - Audio Tools
|
||||
- homepage.name=Parakeet ASR
|
||||
- homepage.icon=mdi-microphone
|
||||
- homepage.description=Parakeet-TDT speech-to-text via sherpa-onnx (irv-ml1)
|
||||
- homepage.href=http://irv-ml1.nh3.internal:${PARAKEET_PORT}
|
||||
- homepage.description=Parakeet-TDT speech-to-text via sherpa-onnx (fv-ml1 GPU 3)
|
||||
- homepage.href=http://10.251.50.54:${PARAKEET_PORT}
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
|
||||
Reference in New Issue
Block a user