Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry the vLLM seats at 84-95.5 GB of 96. Changes: - compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used `count: all`, which would have handed a 0.6B ASR seat all four cards); join traefik-net; port 8300; homepage href to the live FV address. - .env.example: default to the v3 int8 model (25 European languages, 464 MiB) rather than English-only v2; models to /tank/parakeet/models. - app.py: warm the recognizer at startup before uvicorn accepts traffic. The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at 45.1s on a second container) against ~0.50s warm. A 45s first request is indistinguishable from a hang and LiteLLM's default timeout abandons it long before it returns. Decoding 1s of silence at load moves the cost inside the healthcheck's 300s start_period; first real request after restart is now 0.65s. Verification, because "provider=cuda" in the log is only an echo of the env var: ORT falls back to CPU silently and still returns correct text, so the service being up and the transcript being right establishes nothing. The discriminator is a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS sentence transcribes near-exactly (positive), 3s of digital silence returns empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread 0.47-0.65s, single-stream, one clip: a smoke measurement with its harness stated, not a benchmark. Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1` (OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's Postgres store where the ext-tts family already lives — no gateway restart, and config.yaml is consequently not a complete picture of what the gateway serves. Both verified end to end. The aliases use a raw IP deliberately: ana-docker resolves no .internal names at all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a hand-pinned extra_hosts entry. A second hosts entry would mean recreating the container and bouncing the gateway for every consumer. Also records the svos_miranda plugin validation pass and its structural findings, and notes that the irv-ml1 parakeet is still running — there are two now, and retiring the old one is the operator's call.
3.9 KiB
Parakeet STT on fv-ml1 GPU 3 (2026-09-15)
Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias.
What it is
stacks/parakeet/ — Parakeet-TDT 0.6B v3 int8 ONNX (25 European languages,
464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container
parakeet, port 8300, GPU 3 pinned by device_ids. Image
local/parakeet:sherpa-onnx-v4 (5.09 GB).
Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather than rewritten — the Ampere→Blackwell move was the only real question.
Why GPU 3
GPU 0 = 84/96 GB, GPU 1 = 92.9/96, GPU 2 = 95.5/96 (the vLLM seats). GPU 3 was
at 2 MiB. The dead on-host stub used count: all, which would have handed this
seat all four cards; replaced with an explicit device_ids: ["3"] per the fleet
convention. Inside the container the pinned card presents as cuda:0, which is
what ORT's CUDA EP takes by default.
⚠ The finding worth keeping: a 45-second first decode
ONNX Runtime's CUDA EP compiles and autotunes lazily, on the first decode, not at session creation. On sm_120:
| measured | |
|---|---|
| first decode, cold container | 45.7 s (n=1), reproduced at 45.1 s on a second container |
| warm, 8.52 s clip | 0.50 s median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50) |
≈17× realtime warm, single-stream, one 8.52 s clip, int8, GPU 3 idle otherwise. That is a smoke measurement with its harness stated, not a benchmark — no concurrency sweep, no length sweep, one clip.
A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's
default timeout would abandon it. _warm() in app.py now decodes 1 s of silence
before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s
start_period. First real request after restart: 0.65 s.
⚠⚠ "provider=cuda" is not evidence the GPU is being used
ORT's CUDA EP falls back to CPU silently — the process lives, answers 200, and
returns correct text, just slowly. Our own log line loading OfflineRecognizer (provider=cuda...) merely echoes the env var and proves nothing.
The discriminator that actually settles it:
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 3
-> 1588301, /opt/venv/bin/python3, 922 MiB
Timing is not a sufficient check either — the int8 model is fast enough on a 96-thread EPYC that a CPU fallback still looks brisk on short clips.
Controls run, both directions:
- positive — known TTS sentence in, near-exact transcript out (two word errors, both attributable to the source audio: an inserted "um", "Foun Valley").
- null — 3 s of digital silence →
{"text": ""}. The instrument does not manufacture signal.
LiteLLM
Two aliases, both mode: audio_transcription → http://10.251.50.54:8300/v1:
ext-stt (engine-neutral fleet name, mirrors ext-tts) and whisper-1
(OpenAI-compatible drop-in). Both verified end-to-end through the gateway.
Registered via POST /model/new, i.e. the Postgres store, not config.yaml —
that is where the ext-tts family lives, and it needs no gateway restart.
⚠ Corollary: config.yaml is NOT a complete picture of what the gateway serves
(it lists 35 models; the gateway serves 40, and carries stale entries like
granite-4.1-8b). Read /v1/models or /model/info, never just the file.
⚠ Raw IP on purpose — see the ana-docker DNS row in the index.
Loose ends
- The irv-ml1 parakeet is still running (healthz 200 on
100.64.0.6:8765). Two Parakeets now. Retiring the old one is the operator's call — not touched. /opt/docker/compose/parakeetand/tank/parakeetnormalised toroot:docker 2775; the rest of fv-ml1's deploy tree is stilllkraven:lkraven(it was not part of the 5-host normalisation).servers/fv-ml1/README.mdis still broadly stale — it claims 2 GPUs and a 2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected.