The 2026-09-06 headscale cutover retired irv-ml1's wg0 tunnel IP 10.100.79.3 (now 10.6.110.50). Repointed all LIVE canonical refs to the DNS NAME so the next move can't re-break them: homepage.href/siteMonitor labels across 25 stack composes, load-bearing env defaults (asset-engine INFERENCE_HOST, open-webui AUDIO_TTS_OPENAI_API_BASE_URL, skaldsong SKALDSONG_TTS_BASE_URL, zonos-gateway ZONOS_URL, dia), homepage services.yaml manual cards (Voice Design Studio, IRV-ML1), and servers/irv-ml1/ssh-target. Updated the stale 'WG tunnel' comment to the mesh reality. Left as-is: README curl-examples and .env.example comments (docs), and historical mentions in CLAUDE.md/persistent-memory. NOTE: applying the label repoints to the RUNNING irv-ml1 containers needs a recreate per service (labels read at creation); deployed .env values are separate from these canonical defaults.
Voxtral TTS
mistralai/Voxtral-4B-TTS-2603 — Mistral AI's 4B open-weight streaming TTS, served via the vLLM-Omni production serving stack (Mistral co-developed). Released March 28, 2026.
⚠️ License
CC BY-NC. Personal use, research, and internal tooling are fine. Don't ship Voxtral output in any commercial product without re-licensing from Mistral. The other TTS in this fleet (Kokoro, Chatterbox, Fish S2-Pro, IndexTTS-2, Qwen3-TTS, CosyVoice) are all open-licensed and clean for commercial work.
Why this stack exists
Multilingual streaming with serious speed:
| use case | |
|---|---|
| Voxtral | multilingual EN/FR/DE/ES/IT/PT/NL/HI streaming, 70 ms model latency |
| Kokoro | low-latency English, fixed voice library |
| Chatterbox Turbo | low-latency English w/ cloning + 9 paralinguistic tags |
| Fish Audio S2-Pro | richest paralinguistic English (15k+ tags) |
| IndexTTS-2 | English voice cloning + emotion vector / text control |
| Qwen3-TTS-1.7B | English voice cloning (slow on official backend) |
| CosyVoice 3 | multilingual (Chinese-leaning) |
| VibeVoice 1.5B | long-form / multi-speaker dialogue |
Voxtral fills the multilingual + low-latency + cloning slot that's been weak in the fleet (CosyVoice is multilingual but slow on English; nothing else is multilingual at all).
Headline numbers
- 70 ms model latency for a typical 10 s sample (500-char input)
- 9.7× realtime factor
- 68.4% blind A/B win rate vs ElevenLabs Flash v2.5 in voice cloning evaluations
- 8 languages: EN, FR, DE, ES, IT, PT, NL, HI
API
vLLM-Omni serves an OpenAI-compatible API at
http://10.100.79.3:8201/v1:
# Single-shot synthesis.
curl -fsS -X POST http://10.100.79.3:8201/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"mistralai/Voxtral-4B-TTS-2603","input":"Hello there.","voice":"alloy","response_format":"wav"}' \
> out.wav
# Streaming.
curl -fsS -X POST http://10.100.79.3:8201/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"mistralai/Voxtral-4B-TTS-2603","input":"long passage…","voice":"alloy","stream":true}' \
| mpv --no-cache -
# vLLM-Omni standard endpoints.
curl http://10.100.79.3:8201/v1/models # confirms model loaded
curl http://10.100.79.3:8201/v1/audio/voices # built-in + cloned voices
Deploy
scripts/elway irv-ml1 --playbook playbooks/deploy-voxtral.yaml
First boot pulls Voxtral-4B (~8 GB BF16) into the HF cache + warms vLLM. Both are cached afterwards.
Hardware footprint
- VRAM: ~16 GB practical (8 GB weights + KV + activation). Pinned to GPU 1 (RTX A6000) by default — comfortable headroom. The 3090's 24 GB CAN fit but it's tight for long streaming sessions.
- Disk: ~8 GB for the Voxtral checkpoint + HF cache.
Voice library
Drop reference WAV / FLAC into /worktank/voxtral/voices/ on the
host. The wrapper scans on request — no restart needed.