Files
esh-pfi-infrastructure/stacks/voxtral/README.md
T
vh 16d018ff96 stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B
Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:

* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
  4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
  Headline: 15,000+ paralinguistic / emotion tags via natural language
  ([laugh] [whispers] [super happy] etc.) — a step-function over
  Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
  EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
  open. ~17 GB VRAM.

* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
  4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
  win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
  (EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
  serving stack) — published Docker image, no local build. ~16 GB VRAM.
  CC BY-NC license — personal/research use only; flagged in README.

* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
  Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
  solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
  official deploy is Rust + websockets only; using NillPointer's
  community OpenAI-compat wrapper to bridge to /v1/audio/speech so
  it slots into the same bench harness. ~4-6 GB VRAM.

Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.

Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
2026-04-27 22:40:10 -07:00

86 lines
3.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Voxtral TTS
[mistralai/Voxtral-4B-TTS-2603](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603)
— Mistral AI's 4B open-weight streaming TTS, served via the
[vLLM-Omni](https://github.com/vllm-project/vllm-omni) production
serving stack (Mistral co-developed). Released March 28, 2026.
## ⚠️ License
**CC BY-NC.** Personal use, research, and internal tooling are fine.
**Don't ship Voxtral output in any commercial product** without
re-licensing from Mistral. The other TTS in this fleet (Kokoro,
Chatterbox, Fish S2-Pro, IndexTTS-2, Qwen3-TTS, CosyVoice) are all
open-licensed and clean for commercial work.
## Why this stack exists
Multilingual streaming with serious speed:
| | use case |
|---|---|
| **Voxtral** | multilingual EN/FR/DE/ES/IT/PT/NL/HI streaming, 70 ms model latency |
| Kokoro | low-latency English, fixed voice library |
| Chatterbox Turbo | low-latency English w/ cloning + 9 paralinguistic tags |
| Fish Audio S2-Pro | richest paralinguistic English (15k+ tags) |
| IndexTTS-2 | English voice cloning + emotion vector / text control |
| Qwen3-TTS-1.7B | English voice cloning (slow on official backend) |
| CosyVoice 3 | multilingual (Chinese-leaning) |
| VibeVoice 1.5B | long-form / multi-speaker dialogue |
Voxtral fills the **multilingual + low-latency + cloning** slot
that's been weak in the fleet (CosyVoice is multilingual but slow on
English; nothing else is multilingual at all).
## Headline numbers
- **70 ms** model latency for a typical 10 s sample (500-char input)
- **9.7×** realtime factor
- **68.4%** blind A/B win rate vs ElevenLabs Flash v2.5 in voice
cloning evaluations
- **8 languages**: EN, FR, DE, ES, IT, PT, NL, HI
## API
vLLM-Omni serves an OpenAI-compatible API at
`http://10.100.79.3:8197/v1`:
```bash
# Single-shot synthesis.
curl -fsS -X POST http://10.100.79.3:8197/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"mistralai/Voxtral-4B-TTS-2603","input":"Hello there.","voice":"alloy","response_format":"wav"}' \
> out.wav
# Streaming.
curl -fsS -X POST http://10.100.79.3:8197/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"mistralai/Voxtral-4B-TTS-2603","input":"long passage…","voice":"alloy","stream":true}' \
| mpv --no-cache -
# vLLM-Omni standard endpoints.
curl http://10.100.79.3:8197/v1/models # confirms model loaded
curl http://10.100.79.3:8197/v1/audio/voices # built-in + cloned voices
```
## Deploy
```bash
scripts/elway irv-ml1 --playbook playbooks/deploy-voxtral.yaml
```
First boot pulls Voxtral-4B (~8 GB BF16) into the HF cache + warms
vLLM. Both are cached afterwards.
## Hardware footprint
- **VRAM**: ~16 GB practical (8 GB weights + KV + activation). Pinned
to GPU 1 (RTX A6000) by default — comfortable headroom. The 3090's
24 GB CAN fit but it's tight for long streaming sessions.
- **Disk**: ~8 GB for the Voxtral checkpoint + HF cache.
## Voice library
Drop reference WAV / FLAC into `/worktank/voxtral/voices/` on the
host. The wrapper scans on request — no restart needed.