The IRV seat was retired on tts-dev's numbers: it lost to the FV seat at both clip lengths and to whisper-large-v3 at 6.24s. Their length sweep fits ~58ms fixed + 56ms per audio-second with an asymptote of ~17.8x realtime, which independently reproduces our 17x on a different clip and harness, and the gateway hop measured below their harness resolution so ext-stt is the right consumer path. Two caveats recorded against our own numbers: their between-run variance is 20% because GPU 0 carries the live chat path, and our 0.50s median came off an idle GPU 3 - a best case, not a comparable. Their RTFx retraction is the durable part: published RTFx is batched throughput on datacenter hardware rather than single-stream latency, and the two differ by ~200x. Also banks the shape their acceptance gate caught, because it generalises past their repo. A JS escape inside a Python string arrives transformed, closing the string and killing the whole inline script, while the page still renders and both import and node --check pass - the file still holds the backslash. That is the same failure as reading provider=cuda out of a log: a check that reads the artifact as stored cannot see a transformation that happens between storage and execution. Both check the input to a transformation and get reported as if they checked its output.
176 lines
8.9 KiB
Markdown
176 lines
8.9 KiB
Markdown
# Parakeet STT on fv-ml1 GPU 3 (2026-09-15)
|
|
|
|
Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias.
|
|
|
|
## What it is
|
|
|
|
`stacks/parakeet/` — Parakeet-TDT 0.6B **v3** int8 ONNX (25 European languages,
|
|
464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container
|
|
`parakeet`, port **8300**, **GPU 0** pinned by `device_ids`. Image
|
|
`local/parakeet:sherpa-onnx-v4` (5.09 GB).
|
|
|
|
Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather
|
|
than rewritten — the Ampere→Blackwell move was the only real question.
|
|
|
|
## ⚠ Placement — got this wrong first, operator caught it
|
|
|
|
Placed on the empty **GPU 3** initially, reading "the utility gpu" as "the spare
|
|
card". Operator's correction: *"1gb total vram pressure — and you didn't load it on
|
|
gpu 0?"* He is right, and the reason is sharper than "it fits anywhere".
|
|
|
|
**vLLM sizes its KV cache as a fraction of TOTAL VRAM, not free VRAM.** So a
|
|
resident tenant on an otherwise-clean card does not cost its own megabytes — it
|
|
costs a future full-size seat's profiling margin. `flash-next` needs **93 GiB of
|
|
96**. A 96 GB card at 2 MiB is a card that can still take that; the same card at
|
|
922 MiB is a card where the next big seat's `--gpu-memory-utilization` has to be
|
|
hand-trimmed, and the flash-next history in this repo shows exactly how thin and
|
|
how silent that failure gets.
|
|
|
|
The right question is not "where does 800 MiB fit" but "whose headroom is cheapest
|
|
to spend":
|
|
|
|
| GPU | committed util | spare |
|
|
|---|---|---|
|
|
| **0** | 0.40 + 0.48 = **0.88** | ~13 GB ← moved here |
|
|
| 1 | **0.975** (six small seats) | ~4.3 GB |
|
|
| 2 | **0.96** (flash-next) | ~1.8 GB |
|
|
| 3 | — | **kept empty as reserve** |
|
|
|
|
Moved the same night: one env var (`PARAKEET_GPU`) plus `compose up -d`. GPU 3 back
|
|
to 2 MiB / 97,247 MiB free. Post-move n=5 on the same clip: 0.68 / 0.54 / 0.54 /
|
|
0.52 / 0.53 s, median 0.54 s — **indistinguishable from the GPU 3 median of 0.50 s
|
|
at this sample size**; the spreads overlap and no difference is claimed.
|
|
|
|
The dead on-host stub used `count: all`, which would have handed this seat all four
|
|
cards; replaced with an explicit `device_ids` pin per the fleet convention. Inside
|
|
the container the pinned card presents as `cuda:0`, which is what ORT's CUDA EP
|
|
takes by default.
|
|
|
|
## ⚠ The finding worth keeping: a 45-second first decode
|
|
|
|
ONNX Runtime's CUDA EP compiles and autotunes lazily, on the **first decode**, not
|
|
at session creation. On sm_120:
|
|
|
|
| | measured |
|
|
|---|---|
|
|
| first decode, cold container | **45.7 s** (n=1), reproduced at **45.1 s** on a second container |
|
|
| warm, 8.52 s clip | **0.50 s** median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50) |
|
|
|
|
≈17x realtime warm, single-stream, one 8.52 s clip, int8. ⚠ Measured on GPU 3 while
|
|
it was idle; the seat now lives on GPU 0 beside the hot serving path, so treat that
|
|
number as a best case.
|
|
That is a smoke measurement with its harness stated, **not** a benchmark — no
|
|
concurrency sweep, no length sweep, one clip.
|
|
|
|
A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's
|
|
default timeout would abandon it. `_warm()` in `app.py` now decodes 1 s of silence
|
|
before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s
|
|
`start_period`. First real request after restart: **0.65 s**.
|
|
|
|
## ⚠⚠ "provider=cuda" is not evidence the GPU is being used
|
|
|
|
ORT's CUDA EP **falls back to CPU silently** — the process lives, answers 200, and
|
|
returns *correct text*, just slowly. Our own log line `loading OfflineRecognizer
|
|
(provider=cuda...)` merely echoes the env var and proves nothing.
|
|
|
|
The discriminator that actually settles it:
|
|
|
|
```
|
|
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 0
|
|
-> 1594431, /opt/venv/bin/python3, 794 MiB (beside two VLLM::EngineCore entries)
|
|
```
|
|
|
|
Timing is **not** a sufficient check either — the int8 model is fast enough on a
|
|
96-thread EPYC that a CPU fallback still looks brisk on short clips.
|
|
|
|
Controls run, both directions:
|
|
- **positive** — known TTS sentence in, near-exact transcript out (two word errors,
|
|
both attributable to the source audio: an inserted "um", "Foun Valley").
|
|
- **null** — 3 s of digital silence → `{"text": ""}`. The instrument does not
|
|
manufacture signal.
|
|
|
|
## LiteLLM
|
|
|
|
Two aliases, both `mode: audio_transcription` → `http://10.251.50.54:8300/v1`:
|
|
`ext-stt` (engine-neutral fleet name, mirrors `ext-tts`) and `whisper-1`
|
|
(OpenAI-compatible drop-in). Both verified end-to-end through the gateway.
|
|
|
|
Registered via `POST /model/new`, i.e. the **Postgres store**, not `config.yaml` —
|
|
that is where the `ext-tts` family lives, and it needs no gateway restart.
|
|
⚠ Corollary: `config.yaml` is NOT a complete picture of what the gateway serves
|
|
(it lists 35 models; the gateway serves 40, and carries stale entries like
|
|
`granite-4.1-8b`). Read `/v1/models` or `/model/info`, never just the file.
|
|
|
|
⚠ **Raw IP on purpose** — see the ana-docker DNS row in the index.
|
|
|
|
## Loose ends
|
|
|
|
- ✅ **irv-ml1 parakeet RETIRED 2026-09-15** (operator ruling, on tts-dev's bench
|
|
evidence). `docker compose down`; retirement banner prepended to its on-host
|
|
README naming the replacement. Checked for consumers first: **no gateway alias
|
|
pointed at it**, and every other `8765`/`parakeet` reference on that host was a
|
|
comment in a port-allocation register, not a dependency. Model files left on
|
|
disk at `/worktank/parakeet/models/` (regenerable). One Parakeet now.
|
|
- `/opt/docker/compose/parakeet` and `/tank/parakeet` normalised to `root:docker
|
|
2775`; the rest of fv-ml1's deploy tree is still `lkraven:lkraven` (it was not
|
|
part of the 5-host normalisation).
|
|
- `servers/fv-ml1/README.md` is still broadly stale — it claims 2 GPUs and a
|
|
2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected.
|
|
|
|
|
|
## ✅ The bench, and why the IRV seat was retired
|
|
|
|
Endpoints sent to **tts-dev** 2026-09-15; **IRV retired the same night on the result.**
|
|
|
|
**Result** (same clips, same client, same night, vs the Whisper incumbent):
|
|
|
|
| clip | whisper-large-v3 | IRV v2 / 3090 | FV v3 / Blackwell |
|
|
|---|---|---|---|
|
|
| 1.84 s | 457 ms | 354 ms | **155 ms** |
|
|
| 6.24 s | 690 ms | **1010 ms** | **391 ms** |
|
|
|
|
IRV lost at both lengths and was *slower than the incumbent* at 6.24 s. Their length
|
|
sweep (n=9/cell, first 3 discarded) fits **~58 ms fixed + 56 ms per audio-second**,
|
|
asymptote **~17.8x realtime** — independently reproducing our 17x on a different clip
|
|
and harness. Gateway hop measured **below their harness resolution** (±30 ms), so
|
|
`ext-stt` is the right consumer path rather than a direct port.
|
|
|
|
⚠ **Their between-run variance is ±20%**, because GPU 0 carries the live chat path.
|
|
Our 0.50 s median was taken on an idle GPU 3 — a best case, not a comparable.
|
|
|
|
⭐ **tts-dev retracted their own plan's 60-120 ms projection**: published RTFx is
|
|
**batched throughput on datacenter hardware, not single-stream latency** — the two
|
|
differ by **~200x**. Consequence that outlived the win: STT was never the bottleneck
|
|
(~217 ms STT / 464 ms LLM / 478 ms TTS at a 3 s utterance).
|
|
|
|
**Consumer:** `talk`'s push-to-talk ("Grima") went live the same night through
|
|
`/api/listen` -> `ext-stt`, 16 kHz mono decimated 3:1 in an AudioWorklet.
|
|
|
|
⭐ **Their acceptance gate is worth copying.** They drove a real Chromium handed our
|
|
known clip as its microphone, through the page's real handlers. It caught a bug every
|
|
cheaper check passed: a JS `'didn\'t'` inside a Python string arrives as `'didn't'`,
|
|
closing the string and killing the whole inline script — while the page still renders,
|
|
`import app` passes and `node --check` passes, because the file still holds the
|
|
backslash. **Same shape as the silent-CPU-fallback trap: a check that reads the
|
|
artifact AS STORED cannot see a transformation between storage and execution.**
|
|
`node --check` reads the pre-Python file; `provider=cuda` in a log echoes configured
|
|
intent. Both check the INPUT to a transformation and are reported as if they checked
|
|
its output.
|
|
|
|
## The two seats, for the record
|
|
|
|
| | FV (new) | IRV (existing, up 2 months) |
|
|
|---|---|---|
|
|
| endpoint | `http://10.251.50.54:8300/v1/audio/transcriptions` | `http://100.64.0.6:8765/...` or `http://10.6.110.50:8765/...` |
|
|
| model | parakeet-tdt-0.6b-**v3** int8, 25 languages | parakeet-tdt-0.6b-**v2** int8, English only |
|
|
| GPU | RTX PRO 6000 Blackwell **sm_120**, GPU 0, shares with 2 vLLM seats | RTX 3090 **sm_86**, shares with 4 processes, 4.0 GB free |
|
|
| image | `local/parakeet:sherpa-onnx-v4` (has startup warmup) | `local/parakeet:sherpa-onnx-v2` (no warmup) |
|
|
|
|
⚠ **`10.100.79.3:8765` is DEAD** — the retired wg0 lifeline, still the href on IRV's
|
|
Homepage card. Same for `Speaches ASR` at `10.100.79.3:8204`.
|
|
|
|
⚠ **These were never an A/B pair — four things differ at once** (model version,
|
|
GPU architecture, card contention, image). A WER delta is a **v2-vs-v3** result, not
|
|
an FV-vs-IRV one. Offered tts-dev a v2 container on FV as a second compose project so
|
|
accuracy can be varied one factor at a time; not built unless they take it up.
|