Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-15-parakeet-stt-fv-ml1.md
T
vh b9b14b5baf feat(parakeet): stand up Parakeet STT on fv-ml1 GPU 3 + LiteLLM ext-stt/whisper-1
Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card
and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry
the vLLM seats at 84-95.5 GB of 96.

Changes:

- compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used
  `count: all`, which would have handed a 0.6B ASR seat all four cards);
  join traefik-net; port 8300; homepage href to the live FV address.
- .env.example: default to the v3 int8 model (25 European languages, 464 MiB)
  rather than English-only v2; models to /tank/parakeet/models.
- app.py: warm the recognizer at startup before uvicorn accepts traffic.

The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes
lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at
45.1s on a second container) against ~0.50s warm. A 45s first request is
indistinguishable from a hang and LiteLLM's default timeout abandons it long
before it returns. Decoding 1s of silence at load moves the cost inside the
healthcheck's 300s start_period; first real request after restart is now 0.65s.

Verification, because "provider=cuda" in the log is only an echo of the env var:
ORT falls back to CPU silently and still returns correct text, so the service
being up and the transcript being right establishes nothing. The discriminator is
a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS
sentence transcribes near-exactly (positive), 3s of digital silence returns
empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread
0.47-0.65s, single-stream, one clip: a smoke measurement with its harness
stated, not a benchmark.

Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1`
(OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's
Postgres store where the ext-tts family already lives — no gateway restart, and
config.yaml is consequently not a complete picture of what the gateway serves.
Both verified end to end.

The aliases use a raw IP deliberately: ana-docker resolves no .internal names at
all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a
hand-pinned extra_hosts entry. A second hosts entry would mean recreating the
container and bouncing the gateway for every consumer.

Also records the svos_miranda plugin validation pass and its structural findings,
and notes that the irv-ml1 parakeet is still running — there are two now, and
retiring the old one is the operator's call.
2026-09-15 01:41:41 -07:00

3.9 KiB
Raw Blame History

Parakeet STT on fv-ml1 GPU 3 (2026-09-15)

Operator asked for an STT service on fv-ml1's utility GPU plus a LiteLLM alias.

What it is

stacks/parakeet/ — Parakeet-TDT 0.6B v3 int8 ONNX (25 European languages, 464 MiB) under sherpa-onnx, behind ~90 lines of FastAPI we own. Container parakeet, port 8300, GPU 3 pinned by device_ids. Image local/parakeet:sherpa-onnx-v4 (5.09 GB).

Not greenfield: the stack already existed, targeting irv-ml1. Retargeted rather than rewritten — the Ampere→Blackwell move was the only real question.

Why GPU 3

GPU 0 = 84/96 GB, GPU 1 = 92.9/96, GPU 2 = 95.5/96 (the vLLM seats). GPU 3 was at 2 MiB. The dead on-host stub used count: all, which would have handed this seat all four cards; replaced with an explicit device_ids: ["3"] per the fleet convention. Inside the container the pinned card presents as cuda:0, which is what ORT's CUDA EP takes by default.

⚠ The finding worth keeping: a 45-second first decode

ONNX Runtime's CUDA EP compiles and autotunes lazily, on the first decode, not at session creation. On sm_120:

measured
first decode, cold container 45.7 s (n=1), reproduced at 45.1 s on a second container
warm, 8.52 s clip 0.50 s median (n=5: 0.65 / 0.53 / 0.48 / 0.47 / 0.50)

≈17× realtime warm, single-stream, one 8.52 s clip, int8, GPU 3 idle otherwise. That is a smoke measurement with its harness stated, not a benchmark — no concurrency sweep, no length sweep, one clip.

A 45 s first request is indistinguishable from a hang to any caller, and LiteLLM's default timeout would abandon it. _warm() in app.py now decodes 1 s of silence before uvicorn accepts traffic, so the cost lands inside the healthcheck's 300 s start_period. First real request after restart: 0.65 s.

⚠⚠ "provider=cuda" is not evidence the GPU is being used

ORT's CUDA EP falls back to CPU silently — the process lives, answers 200, and returns correct text, just slowly. Our own log line loading OfflineRecognizer (provider=cuda...) merely echoes the env var and proves nothing.

The discriminator that actually settles it:

nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv -i 3
-> 1588301, /opt/venv/bin/python3, 922 MiB

Timing is not a sufficient check either — the int8 model is fast enough on a 96-thread EPYC that a CPU fallback still looks brisk on short clips.

Controls run, both directions:

  • positive — known TTS sentence in, near-exact transcript out (two word errors, both attributable to the source audio: an inserted "um", "Foun Valley").
  • null — 3 s of digital silence → {"text": ""}. The instrument does not manufacture signal.

LiteLLM

Two aliases, both mode: audio_transcription → http://10.251.50.54:8300/v1: ext-stt (engine-neutral fleet name, mirrors ext-tts) and whisper-1 (OpenAI-compatible drop-in). Both verified end-to-end through the gateway.

Registered via POST /model/new, i.e. the Postgres store, not config.yaml — that is where the ext-tts family lives, and it needs no gateway restart. ⚠ Corollary: config.yaml is NOT a complete picture of what the gateway serves (it lists 35 models; the gateway serves 40, and carries stale entries like granite-4.1-8b). Read /v1/models or /model/info, never just the file.

⚠ Raw IP on purpose — see the ana-docker DNS row in the index.

Loose ends

  • The irv-ml1 parakeet is still running (healthz 200 on 100.64.0.6:8765). Two Parakeets now. Retiring the old one is the operator's call — not touched.
  • /opt/docker/compose/parakeet and /tank/parakeet normalised to root:docker 2775; the rest of fv-ml1's deploy tree is still lkraven:lkraven (it was not part of the 5-host normalisation).
  • servers/fv-ml1/README.md is still broadly stale — it claims 2 GPUs and a 2026-07-22 stack list. Only the parakeet/GPU-3 rows were corrected.