Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry the vLLM seats at 84-95.5 GB of 96. Changes: - compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used `count: all`, which would have handed a 0.6B ASR seat all four cards); join traefik-net; port 8300; homepage href to the live FV address. - .env.example: default to the v3 int8 model (25 European languages, 464 MiB) rather than English-only v2; models to /tank/parakeet/models. - app.py: warm the recognizer at startup before uvicorn accepts traffic. The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at 45.1s on a second container) against ~0.50s warm. A 45s first request is indistinguishable from a hang and LiteLLM's default timeout abandons it long before it returns. Decoding 1s of silence at load moves the cost inside the healthcheck's 300s start_period; first real request after restart is now 0.65s. Verification, because "provider=cuda" in the log is only an echo of the env var: ORT falls back to CPU silently and still returns correct text, so the service being up and the transcript being right establishes nothing. The discriminator is a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS sentence transcribes near-exactly (positive), 3s of digital silence returns empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread 0.47-0.65s, single-stream, one clip: a smoke measurement with its harness stated, not a benchmark. Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1` (OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's Postgres store where the ext-tts family already lives — no gateway restart, and config.yaml is consequently not a complete picture of what the gateway serves. Both verified end to end. The aliases use a raw IP deliberately: ana-docker resolves no .internal names at all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a hand-pinned extra_hosts entry. A second hosts entry would mean recreating the container and bouncing the gateway for every consumer. Also records the svos_miranda plugin validation pass and its structural findings, and notes that the irv-ml1 parakeet is still running — there are two now, and retiring the old one is the operator's call.
87 lines
3.9 KiB
Markdown
87 lines
3.9 KiB
Markdown
# Parakeet STT on fv-ml1 GPU 3 (2026-09-15)
|
||
|
||
Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias.
|
||
|
||
## What it is
|
||
|
||
`stacks/parakeet/` — Parakeet-TDT 0.6B **v3** int8 ONNX (25 European languages,
|
||
464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container
|
||
`parakeet`, port **8300**, **GPU 3** pinned by `device_ids`. Image
|
||
`local/parakeet:sherpa-onnx-v4` (5.09 GB).
|
||
|
||
Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather
|
||
than rewritten — the Ampere→Blackwell move was the only real question.
|
||
|
||
## Why GPU 3
|
||
|
||
GPU 0 = 84/96 GB, GPU 1 = 92.9/96, GPU 2 = 95.5/96 (the vLLM seats). **GPU 3 was
|
||
at 2 MiB.** The dead on-host stub used `count: all`, which would have handed this
|
||
seat all four cards; replaced with an explicit `device_ids: ["3"]` per the fleet
|
||
convention. Inside the container the pinned card presents as `cuda:0`, which is
|
||
what ORT's CUDA EP takes by default.
|
||
|
||
## ⚠ The finding worth keeping: a 45-second first decode
|
||
|
||
ONNX Runtime's CUDA EP compiles and autotunes lazily, on the **first decode**, not
|
||
at session creation. On sm_120:
|
||
|
||
| | measured |
|
||
|---|---|
|
||
| first decode, cold container | **45.7 s** (n=1), reproduced at **45.1 s** on a second container |
|
||
| warm, 8.52 s clip | **0.50 s** median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50) |
|
||
|
||
≈17× realtime warm, single-stream, one 8.52 s clip, int8, GPU 3 idle otherwise.
|
||
That is a smoke measurement with its harness stated, **not** a benchmark — no
|
||
concurrency sweep, no length sweep, one clip.
|
||
|
||
A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's
|
||
default timeout would abandon it. `_warm()` in `app.py` now decodes 1 s of silence
|
||
before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s
|
||
`start_period`. First real request after restart: **0.65 s**.
|
||
|
||
## ⚠⚠ "provider=cuda" is not evidence the GPU is being used
|
||
|
||
ORT's CUDA EP **falls back to CPU silently** — the process lives, answers 200, and
|
||
returns *correct text*, just slowly. Our own log line `loading OfflineRecognizer
|
||
(provider=cuda...)` merely echoes the env var and proves nothing.
|
||
|
||
The discriminator that actually settles it:
|
||
|
||
```
|
||
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 3
|
||
-> 1588301, /opt/venv/bin/python3, 922 MiB
|
||
```
|
||
|
||
Timing is **not** a sufficient check either — the int8 model is fast enough on a
|
||
96-thread EPYC that a CPU fallback still looks brisk on short clips.
|
||
|
||
Controls run, both directions:
|
||
- **positive** — known TTS sentence in, near-exact transcript out (two word errors,
|
||
both attributable to the source audio: an inserted "um", "Foun Valley").
|
||
- **null** — 3 s of digital silence → `{"text": ""}`. The instrument does not
|
||
manufacture signal.
|
||
|
||
## LiteLLM
|
||
|
||
Two aliases, both `mode: audio_transcription` → `http://10.251.50.54:8300/v1`:
|
||
`ext-stt` (engine-neutral fleet name, mirrors `ext-tts`) and `whisper-1`
|
||
(OpenAI-compatible drop-in). Both verified end-to-end through the gateway.
|
||
|
||
Registered via `POST /model/new`, i.e. the **Postgres store**, not `config.yaml` —
|
||
that is where the `ext-tts` family lives, and it needs no gateway restart.
|
||
⚠ Corollary: `config.yaml` is NOT a complete picture of what the gateway serves
|
||
(it lists 35 models; the gateway serves 40, and carries stale entries like
|
||
`granite-4.1-8b`). Read `/v1/models` or `/model/info`, never just the file.
|
||
|
||
⚠ **Raw IP on purpose** — see the ana-docker DNS row in the index.
|
||
|
||
## Loose ends
|
||
|
||
- The **irv-ml1 parakeet is still running** (healthz 200 on `100.64.0.6:8765`).
|
||
Two Parakeets now. Retiring the old one is the operator's call — not touched.
|
||
- `/opt/docker/compose/parakeet` and `/tank/parakeet` normalised to `root:docker
|
||
2775`; the rest of fv-ml1's deploy tree is still `lkraven:lkraven` (it was not
|
||
part of the 5-host normalisation).
|
||
- `servers/fv-ml1/README.md` is still broadly stale — it claims 2 GPUs and a
|
||
2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected.
|