voxtral + kyutai-tts: fix wrong image tag / wrong endpoint paths; fish-s2: env-selectable model variant
Three fixes from the second-wave deploy attempts:
* voxtral: vllm/vllm-omni doesn't publish a `latest` tag — pull
failed with "manifest unknown". Pinned VOXTRAL_VLLM_TAG to v0.18.0
(released 2026-03-29, the day after the Voxtral 4B TTS release —
first cut with Voxtral support).
* kyutai-tts: NillPointer wrapper exposes ONLY /health (root) and
POST /v1/audio/speech. No /v1/models, no /v1/audio/voices —
those return 404. Verified by /openapi.json against the live
container. Compose healthcheck + playbook wait + verify steps
all repointed at the actual paths. POST /v1/audio/speech is now
smoke-tested with a RIFF WAV assertion (same pattern as fish-s2).
* fish-s2: added FISH_S2_MODEL env var so the model variant is
swappable via .env without rebuilding. Both s2-pro (default) and
s1-mini are pre-pulled into the bind-mount; LLAMA_CHECKPOINT_PATH
+ DECODER_CHECKPOINT_PATH now use ${FISH_S2_MODEL:-s2-pro}.
s1-mini was originally gated on fishaudio's HF org (401), but
niobures/OpenAudio-S1 mirrors the same files openly — pulled
from there via a one-shot snapshot_download.
This commit is contained in:
@@ -73,26 +73,32 @@ steps:
|
||||
- name: docker compose up -d
|
||||
shell: cd {{ compose_dir }} && docker compose up -d
|
||||
|
||||
- name: Wait for /v1/models to respond (allow ~10 min for first download + warmup)
|
||||
- name: Wait for /health to respond (allow ~10 min for first download + warmup)
|
||||
shell: |
|
||||
for i in $(seq 1 120); do
|
||||
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/v1/models && exit 0
|
||||
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/health && exit 0
|
||||
sleep 5
|
||||
done
|
||||
exit 1
|
||||
changed_when: "false"
|
||||
|
||||
verify:
|
||||
- name: /v1/models returns valid JSON
|
||||
shell: |
|
||||
curl -sf http://localhost:{{ host_port }}/v1/models \
|
||||
| python3 -c "import json,sys; json.load(sys.stdin)"
|
||||
- name: /health returns 200
|
||||
shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/health
|
||||
changed_when: "false"
|
||||
|
||||
- name: /v1/audio/voices returns valid JSON
|
||||
- name: /v1/audio/speech returns a real WAV (POST with text body)
|
||||
# NillPointer wrapper exposes ONLY /health and POST /v1/audio/speech
|
||||
# (no /v1/models, no /v1/audio/voices). Smoke by POSTing and
|
||||
# asserting a RIFF WAV comes back.
|
||||
shell: |
|
||||
curl -sf http://localhost:{{ host_port }}/v1/audio/voices \
|
||||
| python3 -c "import json,sys; json.load(sys.stdin)"
|
||||
out=$(mktemp --suffix=.wav)
|
||||
curl -sf -X POST http://localhost:{{ host_port }}/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"tts-1.6b-en_fr","input":"Verify."}' \
|
||||
-o "$out" --max-time 30
|
||||
file -b "$out" | grep -q '^RIFF.*WAVE'
|
||||
rm -f "$out"
|
||||
changed_when: "false"
|
||||
|
||||
- name: Container is running
|
||||
|
||||
Reference in New Issue
Block a user