Placed on GPU 3 first because it was the empty card. That was the wrong read:
the seat is ~800 MiB, under 1% of a 96 GB card, so the question was never "where
does it fit" but "whose headroom is cheapest to spend".
vLLM sizes its KV cache as a fraction of TOTAL VRAM, not free VRAM. A resident
tenant on an otherwise-clean card therefore does not cost its own megabytes — it
costs the profiling margin of whatever full-size seat lands there later, and
flash-next needs 93 GiB of 96. A 96 GB card at 2 MiB can still take that; the
same card at 922 MiB is one where the next big seat needs its utilization
hand-trimmed, which this repo's flash-next history shows is both thin and silent
when it goes wrong.
Committed utilization per card is the number that governs, not free bytes:
GPU 0 0.40 + 0.48 = 0.88 ~13 GB spare <- moved here
GPU 1 0.52+0.24+0.10+0.055+0.03+0.03 = 0.975 ~4.3 GB
GPU 2 0.96 ~1.8 GB
GPU 3 - kept empty as reserve
GPU 3 is back to 2 MiB / 97,247 MiB free and is now documented as a deliberate
reserve rather than a spare.
Post-move n=5 on the same clip: 0.68 / 0.54 / 0.54 / 0.52 / 0.53 s, median 0.54 s
against 0.50 s on GPU 3. The spreads overlap at this sample size and no difference
is claimed; the GPU 3 figure was taken on an idle card and is now noted as a best
case, since the seat shares GPU 0 with the hot serving path. Silence control and
the gateway round-trip both re-verified after the move.
Also records both Parakeet endpoints (FV v3 on Blackwell, IRV v2 on a 3090) and
the four confounds that make them not an A/B pair, sent to tts-dev for the bench.
Parakeet ASR
NVIDIA Parakeet-TDT 0.6B (int8 ONNX) served by our own thin FastAPI wrapper over sherpa-onnx (ONNX Runtime + CUDA).
Server: fv-ml1 (Fountain Valley, 10.251.50.54) — moved from irv-ml1 2026-09-15
Port: 8300 (container 8000)
GPU: 0, pinned explicitly via device_ids — shares the card with two vLLM seats
Image: local/parakeet:sherpa-onnx-v4 — built from Dockerfile +
app.py + entrypoint.sh in this directory; we own all the code
Model: parakeet-tdt-0.6b-v3 int8, 25 European languages (~464 MiB)
Why GPU 0 and not the empty card
fv-ml1 has four RTX PRO 6000 Blackwell Max-Q (96 GB each). This seat holds ~800 MiB — under 1 % of one card, so the question is not "where does it fit" but "whose headroom can it spend most cheaply".
| GPU | committed --gpu-memory-utilization |
spare |
|---|---|---|
| 0 | 0.40 + 0.48 = 0.88 | ~13 GB ← here |
| 1 | 0.52+0.24+0.10+0.055+0.03+0.03 = 0.975 | ~4.3 GB |
| 2 | 0.96 | ~1.8 GB |
| 3 | — | kept empty |
It was first placed on the empty GPU 3, which was wrong: vLLM sizes its KV
cache as a fraction of TOTAL VRAM, not free VRAM, so any resident tenant on an
otherwise-clean card eats directly into the profiling margin of whatever big seat
lands there later — and flash-next needs 93 GiB of 96. A 96 GB card at 2 MiB is
worth far more than 800 MiB of it. Moved to GPU 0 the same night.
⚠ The pin is deploy.resources.reservations.devices[].device_ids, the fleet
convention — not count: all, which is what the dead on-host stub used and
which would have handed this seat all four cards. Inside the container the pinned
card presents as cuda:0, which is what sherpa-onnx's CUDA execution provider
takes by default.
Why not the FastAPI community wrappers
Both Shadowfita/parakeet-tdt-0.6b-v2-fastapi and
pnivek/Parakeet-ASR-FastAPI look appealing on paper but have open,
unfixed bugs in the actual transcribe path (return-shape mismatches
after an unpinned torchaudio upgrade, torchaudio.tensor which
doesn't exist, etc.). We tried Shadowfita and hit #16+#10 on the
first real request. Rather than babysit someone else's half-tested
code, we moved to sherpa-onnx — ONNX Runtime is a stable base, k2-fsa
publishes prebuilt int8 Parakeet weights per release, and the
recognizer API is a three-line call.
API endpoints
| Method + path | Purpose |
|---|---|
POST /transcribe |
Multipart file upload → {"text": "..."} |
POST /v1/audio/transcriptions |
Same body; OpenAI-compatible path |
GET /healthz |
Health probe (used by docker healthcheck) |
Path layout
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
/tank/parakeet/models/ |
/models |
ONNX encoder+decoder+joiner+tokens (~464 MiB int8) | excluded (regenerable — re-downloads from the URL on first run if absent) |
⚠ Verifying the GPU is actually in use
ONNX Runtime's CUDA execution provider falls back to CPU silently. It logs
a warning if you are looking, keeps the process alive, answers 200, and
returns correct transcriptions — just far slower. So PROVIDER=cuda in
.env is a request, not a guarantee, and "the service is up and the text is
right" does not establish that the GPU is doing the work.
Blackwell is the reason this matters here rather than being pedantry: these
cards are sm_120, newer than the compute capabilities ORT's prebuilt CUDA
binaries have historically shipped kernels for, and the irv-ml1 host this
stack came from was Ampere sm_86. The move is exactly the kind that turns a
green service into a CPU service without a single error.
The honest check is to watch the card while a transcription runs:
# on fv-ml1 — terminal 1
watch -n0.2 'nvidia-smi --query-compute-apps=pid,process_name,used_memory \
--format=csv -i 0'
# terminal 2 — send real audio, not silence
curl -s -F file=@sample.wav http://127.0.0.1:8300/v1/audio/transcriptions
A process must appear on GPU 0 for the duration — ~800 MiB alongside the two
much larger VLLM::EngineCore entries. If it never appears, the
CUDA EP did not initialise and you are on CPU regardless of what .env says.
Confirm with the container's own startup log, which names the providers ORT
actually registered:
docker logs parakeet 2>&1 | grep -i 'provider\|cuda\|onnxruntime'
Timing alone is not sufficient evidence either way: the int8 model is fast enough on a 96-thread EPYC that a CPU fallback still looks brisk on short clips. Use the process check as the discriminator and treat throughput as a secondary signal.
LiteLLM alias
Reached fleet-wide through the gateway rather than by name, engine-neutral so
the backend can be swapped without touching consumers — the same pattern as
ext-tts:
| alias | mode | backend |
|---|---|---|
ext-stt |
audio_transcription |
http://10.251.50.54:8300/v1 |
whisper-1 |
audio_transcription |
same — OpenAI-compatible name so stock SDK clients work unchanged |
⚠ The alias uses a raw IP on purpose. ana-docker (where LiteLLM runs)
resolves no .internal names at all — its /etc/resolv.conf points at
1.1.1.1/1.0.0.1, and the only reason the ext-tts backend resolves is a
hand-pinned extra_hosts: irv-ml1.nh3.internal:10.6.110.50 in the LiteLLM
compose. Adding a second hosts entry would mean recreating the container and
bouncing the gateway for every consumer; an IP costs nothing and cannot go
stale silently. See the DNS follow-up in persistent-memory.md.
Aliases live in LiteLLM's Postgres store (store_model_in_db: true), not
in config.yaml — that is where the ext-tts family lives too, and it means
adding one needs no gateway restart. It also means config.yaml is not a
complete picture of what the gateway serves: check /v1/models or
/model/info, never just the file.
Deploy
scripts/deploy-stack.sh fv-ml1 parakeet --compose
# on fv-ml1, first time only:
# cp .env.example .env # then edit
docker compose -f /opt/docker/compose/parakeet/compose.yaml build
docker compose -f /opt/docker/compose/parakeet/compose.yaml up -d
First boot downloads ~464 MiB of ONNX weights into /tank/parakeet/models/;
the healthcheck's start_period is 300 s to cover it. Subsequent starts skip
the download.
Switching model variants
One line in .env, then docker compose up -d (not restart — the model URL
is read by the entrypoint at container creation) and delete the old files from
/tank/parakeet/models/ so the entrypoint re-downloads:
- v3 (default) — 25 European languages
- v2 — English only; slightly better on English-only material
Both are ~460 MiB int8 tarballs from the same k2-fsa release page.