feat(parakeet): stand up Parakeet STT on fv-ml1 GPU 3 + LiteLLM ext-stt/whisper-1

Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card
and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry
the vLLM seats at 84-95.5 GB of 96.

Changes:

- compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used
  `count: all`, which would have handed a 0.6B ASR seat all four cards);
  join traefik-net; port 8300; homepage href to the live FV address.
- .env.example: default to the v3 int8 model (25 European languages, 464 MiB)
  rather than English-only v2; models to /tank/parakeet/models.
- app.py: warm the recognizer at startup before uvicorn accepts traffic.

The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes
lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at
45.1s on a second container) against ~0.50s warm. A 45s first request is
indistinguishable from a hang and LiteLLM's default timeout abandons it long
before it returns. Decoding 1s of silence at load moves the cost inside the
healthcheck's 300s start_period; first real request after restart is now 0.65s.

Verification, because "provider=cuda" in the log is only an echo of the env var:
ORT falls back to CPU silently and still returns correct text, so the service
being up and the transcript being right establishes nothing. The discriminator is
a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS
sentence transcribes near-exactly (positive), 3s of digital silence returns
empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread
0.47-0.65s, single-stream, one clip: a smoke measurement with its harness
stated, not a benchmark.

Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1`
(OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's
Postgres store where the ext-tts family already lives — no gateway restart, and
config.yaml is consequently not a complete picture of what the gateway serves.
Both verified end to end.

The aliases use a raw IP deliberately: ana-docker resolves no .internal names at
all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a
hand-pinned extra_hosts entry. A second hosts entry would mean recreating the
container and bouncing the gateway for every consumer.

Also records the svos_miranda plugin validation pass and its structural findings,
and notes that the irv-ml1 parakeet is still running — there are two now, and
retiring the old one is the operator's call.
This commit is contained in:
vh
2026-09-15 01:41:41 -07:00
parent c6b6435c52
commit b9b14b5baf
8 changed files with 366 additions and 93 deletions
+30 -1
View File
@@ -6,7 +6,7 @@ Load the encoder/decoder/joiner/tokens once at startup; serve:
GET /healthz — used by the docker healthcheck
No VAD chunking, no Silero preprocessing — parakeet-tdt handles long-form natively
and the int8 ONNX model on a 24 GB GPU eats everything we're likely to throw at it.
and the int8 ONNX model is a rounding error against this host's 96 GB cards.
"""
from __future__ import annotations
@@ -14,6 +14,7 @@ from __future__ import annotations
import io
import logging
import os
import time
from pathlib import Path
import numpy as np
@@ -59,8 +60,36 @@ def _load_recognizer() -> sherpa_onnx.OfflineRecognizer:
)
def _warm(rec: "sherpa_onnx.OfflineRecognizer") -> None:
"""Decode one throwaway buffer before the server accepts traffic.
⚠ NOT an optimisation — it moves a 45 s stall out of the first real request.
ONNX Runtime's CUDA EP compiles and autotunes its kernels lazily, on the first
decode, and on this host (RTX PRO 6000 Blackwell, sm_120) that measured **45.7 s**
while every subsequent call was ~0.48 s. Without this, the first caller after any
container restart sees a 45 s hang and most clients — LiteLLM's default request
timeout included — give up long before it returns, which reads as "the service is
broken" rather than "the service is warming".
The healthcheck's `start_period` (300 s) is what makes paying it here safe.
"""
try:
t0 = time.monotonic()
stream = rec.create_stream()
# 1 s of silence at 16 kHz: enough to force the full encoder/decoder/joiner
# path to compile, cheap enough not to matter.
stream.accept_waveform(16000, np.zeros(16000, dtype=np.float32))
rec.decode_stream(stream)
logger.info("warmup decode complete in %.1fs — CUDA kernels compiled", time.monotonic() - t0)
except Exception:
# A failed warmup must not stop the server: the model is loaded and real
# requests would still work, just with the stall back on the first caller.
logger.exception("warmup decode failed; first real request will absorb the stall")
app = FastAPI(title="Parakeet ASR (sherpa-onnx)")
recognizer = _load_recognizer()
_warm(recognizer)
def _decode(raw: bytes) -> str: