# Parakeet STT on fv-ml1 GPU 3 (2026-09-15) Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias. ## What it is `stacks/parakeet/` — Parakeet-TDT 0.6B **v3** int8 ONNX (25 European languages, 464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container `parakeet`, port **8300**, **GPU 0** pinned by `device_ids`. Image `local/parakeet:sherpa-onnx-v4` (5.09 GB). Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather than rewritten — the Ampere→Blackwell move was the only real question. ## ⚠ Placement — got this wrong first, operator caught it Placed on the empty **GPU 3** initially, reading "the utility gpu" as "the spare card". Operator's correction: *"1gb total vram pressure — and you didn't load it on gpu 0?"* He is right, and the reason is sharper than "it fits anywhere". **vLLM sizes its KV cache as a fraction of TOTAL VRAM, not free VRAM.** So a resident tenant on an otherwise-clean card does not cost its own megabytes — it costs a future full-size seat's profiling margin. `flash-next` needs **93 GiB of 96**. A 96 GB card at 2 MiB is a card that can still take that; the same card at 922 MiB is a card where the next big seat's `--gpu-memory-utilization` has to be hand-trimmed, and the flash-next history in this repo shows exactly how thin and how silent that failure gets. The right question is not "where does 800 MiB fit" but "whose headroom is cheapest to spend": | GPU | committed util | spare | |---|---|---| | **0** | 0.40 + 0.48 = **0.88** | ~13 GB ← moved here | | 1 | **0.975** (six small seats) | ~4.3 GB | | 2 | **0.96** (flash-next) | ~1.8 GB | | 3 | — | **kept empty as reserve** | Moved the same night: one env var (`PARAKEET_GPU`) plus `compose up -d`. GPU 3 back to 2 MiB / 97,247 MiB free. Post-move n=5 on the same clip: 0.68 / 0.54 / 0.54 / 0.52 / 0.53 s, median 0.54 s — **indistinguishable from the GPU 3 median of 0.50 s at this sample size**; the spreads overlap and no difference is claimed. The dead on-host stub used `count: all`, which would have handed this seat all four cards; replaced with an explicit `device_ids` pin per the fleet convention. Inside the container the pinned card presents as `cuda:0`, which is what ORT's CUDA EP takes by default. ## ⚠ The finding worth keeping: a 45-second first decode ONNX Runtime's CUDA EP compiles and autotunes lazily, on the **first decode**, not at session creation. On sm_120: | | measured | |---|---| | first decode, cold container | **45.7 s** (n=1), reproduced at **45.1 s** on a second container | | warm, 8.52 s clip | **0.50 s** median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50) | ≈17x realtime warm, single-stream, one 8.52 s clip, int8. ⚠ Measured on GPU 3 while it was idle; the seat now lives on GPU 0 beside the hot serving path, so treat that number as a best case. That is a smoke measurement with its harness stated, **not** a benchmark — no concurrency sweep, no length sweep, one clip. A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's default timeout would abandon it. `_warm()` in `app.py` now decodes 1 s of silence before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s `start_period`. First real request after restart: **0.65 s**. ## ⚠⚠ "provider=cuda" is not evidence the GPU is being used ORT's CUDA EP **falls back to CPU silently** — the process lives, answers 200, and returns *correct text*, just slowly. Our own log line `loading OfflineRecognizer (provider=cuda...)` merely echoes the env var and proves nothing. The discriminator that actually settles it: ``` nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 0 -> 1594431, /opt/venv/bin/python3, 794 MiB (beside two VLLM::EngineCore entries) ``` Timing is **not** a sufficient check either — the int8 model is fast enough on a 96-thread EPYC that a CPU fallback still looks brisk on short clips. Controls run, both directions: - **positive** — known TTS sentence in, near-exact transcript out (two word errors, both attributable to the source audio: an inserted "um", "Foun Valley"). - **null** — 3 s of digital silence → `{"text": ""}`. The instrument does not manufacture signal. ## LiteLLM Two aliases, both `mode: audio_transcription` → `http://10.251.50.54:8300/v1`: `ext-stt` (engine-neutral fleet name, mirrors `ext-tts`) and `whisper-1` (OpenAI-compatible drop-in). Both verified end-to-end through the gateway. Registered via `POST /model/new`, i.e. the **Postgres store**, not `config.yaml` — that is where the `ext-tts` family lives, and it needs no gateway restart. ⚠ Corollary: `config.yaml` is NOT a complete picture of what the gateway serves (it lists 35 models; the gateway serves 40, and carries stale entries like `granite-4.1-8b`). Read `/v1/models` or `/model/info`, never just the file. ⚠ **Raw IP on purpose** — see the ana-docker DNS row in the index. ## Loose ends - ✅ **irv-ml1 parakeet RETIRED 2026-09-15** (operator ruling, on tts-dev's bench evidence). `docker compose down`; retirement banner prepended to its on-host README naming the replacement. Checked for consumers first: **no gateway alias pointed at it**, and every other `8765`/`parakeet` reference on that host was a comment in a port-allocation register, not a dependency. Model files left on disk at `/worktank/parakeet/models/` (regenerable). One Parakeet now. - `/opt/docker/compose/parakeet` and `/tank/parakeet` normalised to `root:docker 2775`; the rest of fv-ml1's deploy tree is still `lkraven:lkraven` (it was not part of the 5-host normalisation). - `servers/fv-ml1/README.md` is still broadly stale — it claims 2 GPUs and a 2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected. ## ✅ The bench, and why the IRV seat was retired Endpoints sent to **tts-dev** 2026-09-15; **IRV retired the same night on the result.** **Result** (same clips, same client, same night, vs the Whisper incumbent): | clip | whisper-large-v3 | IRV v2 / 3090 | FV v3 / Blackwell | |---|---|---|---| | 1.84 s | 457 ms | 354 ms | **155 ms** | | 6.24 s | 690 ms | **1010 ms** | **391 ms** | IRV lost at both lengths and was *slower than the incumbent* at 6.24 s. Their length sweep (n=9/cell, first 3 discarded) fits **~58 ms fixed + 56 ms per audio-second**, asymptote **~17.8x realtime** — independently reproducing our 17x on a different clip and harness. Gateway hop measured **below their harness resolution** (±30 ms), so `ext-stt` is the right consumer path rather than a direct port. ⚠ **Their between-run variance is ±20%**, because GPU 0 carries the live chat path. Our 0.50 s median was taken on an idle GPU 3 — a best case, not a comparable. ⭐ **tts-dev retracted their own plan's 60-120 ms projection**: published RTFx is **batched throughput on datacenter hardware, not single-stream latency** — the two differ by **~200x**. Consequence that outlived the win: STT was never the bottleneck (~217 ms STT / 464 ms LLM / 478 ms TTS at a 3 s utterance). **Consumer:** `talk`'s push-to-talk ("Grima") went live the same night through `/api/listen` -> `ext-stt`, 16 kHz mono decimated 3:1 in an AudioWorklet. ⭐ **Their acceptance gate is worth copying.** They drove a real Chromium handed our known clip as its microphone, through the page's real handlers. It caught a bug every cheaper check passed: a JS `'didn\'t'` inside a Python string arrives as `'didn't'`, closing the string and killing the whole inline script — while the page still renders, `import app` passes and `node --check` passes, because the file still holds the backslash. **Same shape as the silent-CPU-fallback trap: a check that reads the artifact AS STORED cannot see a transformation between storage and execution.** `node --check` reads the pre-Python file; `provider=cuda` in a log echoes configured intent. Both check the INPUT to a transformation and are reported as if they checked its output. ## The two seats, for the record | | FV (new) | IRV (existing, up 2 months) | |---|---|---| | endpoint | `http://10.251.50.54:8300/v1/audio/transcriptions` | `http://100.64.0.6:8765/...` or `http://10.6.110.50:8765/...` | | model | parakeet-tdt-0.6b-**v3** int8, 25 languages | parakeet-tdt-0.6b-**v2** int8, English only | | GPU | RTX PRO 6000 Blackwell **sm_120**, GPU 0, shares with 2 vLLM seats | RTX 3090 **sm_86**, shares with 4 processes, 4.0 GB free | | image | `local/parakeet:sherpa-onnx-v4` (has startup warmup) | `local/parakeet:sherpa-onnx-v2` (no warmup) | ⚠ **`10.100.79.3:8765` is DEAD** — the retired wg0 lifeline, still the href on IRV's Homepage card. Same for `Speaches ASR` at `10.100.79.3:8204`. ⚠ **These were never an A/B pair — four things differ at once** (model version, GPU architecture, card contention, image). A WER delta is a **v2-vs-v3** result, not an FV-vs-IRV one. Offered tts-dev a v2 container on FV as a second compose project so accuracy can be varied one factor at a time; not built unless they take it up.