# Parakeet STT on fv-ml1 GPU 3 (2026-09-15) Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias. ## What it is `stacks/parakeet/` — Parakeet-TDT 0.6B **v3** int8 ONNX (25 European languages, 464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container `parakeet`, port **8300**, **GPU 3** pinned by `device_ids`. Image `local/parakeet:sherpa-onnx-v4` (5.09 GB). Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather than rewritten — the Ampere→Blackwell move was the only real question. ## Why GPU 3 GPU 0 = 84/96 GB, GPU 1 = 92.9/96, GPU 2 = 95.5/96 (the vLLM seats). **GPU 3 was at 2 MiB.** The dead on-host stub used `count: all`, which would have handed this seat all four cards; replaced with an explicit `device_ids: ["3"]` per the fleet convention. Inside the container the pinned card presents as `cuda:0`, which is what ORT's CUDA EP takes by default. ## ⚠ The finding worth keeping: a 45-second first decode ONNX Runtime's CUDA EP compiles and autotunes lazily, on the **first decode**, not at session creation. On sm_120: | | measured | |---|---| | first decode, cold container | **45.7 s** (n=1), reproduced at **45.1 s** on a second container | | warm, 8.52 s clip | **0.50 s** median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50) | ≈17× realtime warm, single-stream, one 8.52 s clip, int8, GPU 3 idle otherwise. That is a smoke measurement with its harness stated, **not** a benchmark — no concurrency sweep, no length sweep, one clip. A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's default timeout would abandon it. `_warm()` in `app.py` now decodes 1 s of silence before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s `start_period`. First real request after restart: **0.65 s**. ## ⚠⚠ "provider=cuda" is not evidence the GPU is being used ORT's CUDA EP **falls back to CPU silently** — the process lives, answers 200, and returns *correct text*, just slowly. Our own log line `loading OfflineRecognizer (provider=cuda...)` merely echoes the env var and proves nothing. The discriminator that actually settles it: ``` nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 3 -> 1588301, /opt/venv/bin/python3, 922 MiB ``` Timing is **not** a sufficient check either — the int8 model is fast enough on a 96-thread EPYC that a CPU fallback still looks brisk on short clips. Controls run, both directions: - **positive** — known TTS sentence in, near-exact transcript out (two word errors, both attributable to the source audio: an inserted "um", "Foun Valley"). - **null** — 3 s of digital silence → `{"text": ""}`. The instrument does not manufacture signal. ## LiteLLM Two aliases, both `mode: audio_transcription` → `http://10.251.50.54:8300/v1`: `ext-stt` (engine-neutral fleet name, mirrors `ext-tts`) and `whisper-1` (OpenAI-compatible drop-in). Both verified end-to-end through the gateway. Registered via `POST /model/new`, i.e. the **Postgres store**, not `config.yaml` — that is where the `ext-tts` family lives, and it needs no gateway restart. ⚠ Corollary: `config.yaml` is NOT a complete picture of what the gateway serves (it lists 35 models; the gateway serves 40, and carries stale entries like `granite-4.1-8b`). Read `/v1/models` or `/model/info`, never just the file. ⚠ **Raw IP on purpose** — see the ana-docker DNS row in the index. ## Loose ends - The **irv-ml1 parakeet is still running** (healthz 200 on `100.64.0.6:8765`). Two Parakeets now. Retiring the old one is the operator's call — not touched. - `/opt/docker/compose/parakeet` and `/tank/parakeet` normalised to `root:docker 2775`; the rest of fv-ml1's deploy tree is still `lkraven:lkraven` (it was not part of the 5-host normalisation). - `servers/fv-ml1/README.md` is still broadly stale — it claims 2 GPUs and a 2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected.