stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B

Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:

* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
  4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
  Headline: 15,000+ paralinguistic / emotion tags via natural language
  ([laugh] [whispers] [super happy] etc.) — a step-function over
  Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
  EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
  open. ~17 GB VRAM.

* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
  4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
  win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
  (EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
  serving stack) — published Docker image, no local build. ~16 GB VRAM.
  CC BY-NC license — personal/research use only; flagged in README.

* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
  Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
  solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
  official deploy is Rust + websockets only; using NillPointer's
  community OpenAI-compat wrapper to bridge to /v1/audio/speech so
  it slots into the same bench harness. ~4-6 GB VRAM.

Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.

Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
This commit is contained in:
vh
2026-04-27 22:40:10 -07:00
parent db42a7cc17
commit 16d018ff96
12 changed files with 872 additions and 0 deletions
+105
View File
@@ -0,0 +1,105 @@
# Fish Audio S2-Pro
[fishaudio/s2-pro](https://huggingface.co/fishaudio/s2-pro) — the most
expressive open-source TTS model as of 2026-04, served via the
official [fishaudio/fish-speech](https://github.com/fishaudio/fish-speech)
inference engine.
## Why this stack exists
Three of the existing TTS already cover the basics — Kokoro for raw
speed, Chatterbox for speed-with-cloning, IndexTTS-2 for precision
emotion control. Fish Audio S2-Pro fills a different slot:
**dramatically richer paralinguistic control via natural-language
tags** (15,000+ vs Chatterbox Turbo's 9 fixed tags), with comparable
latency (~150 ms streaming) and voice cloning.
Released March 9, 2026; we missed it during the original irv-ml1
build-out in early April.
| | use case |
|---|---|
| **Fish Audio S2-Pro** | richest emotive / paralinguistic English TTS — 15k+ tags |
| Kokoro | low-latency English, fixed voice library |
| Chatterbox Turbo | low-latency English w/ cloning + 9 paralinguistic tags |
| IndexTTS-2 | English voice cloning + emotion vector / text control |
| Qwen3-TTS-1.7B | English voice cloning (slow on official backend) |
| CosyVoice 3 | multilingual (Chinese-leaning) |
| VibeVoice 1.5B | long-form / multi-speaker dialogue |
## Architecture
Dual-AR (Slow + Fast):
- **Slow AR** operates along the time axis, predicts the primary
semantic codebook.
- **Fast AR** generates the remaining 9 residual codebooks per time
step, reconstructing fine-grained acoustic detail.
Trained on 10M+ hours of audio across 80+ languages with
reinforcement-learning alignment. Win rates per upstream:
| benchmark | S2-Pro |
|---|---|
| EmergentTTS-Eval paralinguistics | 91.61% |
| Blind A/B vs ElevenLabs Flash v2.5 (multilingual) | strong |
## Headline features
- **15,000+ paralinguistic / emotion tags** via natural language:
```
[laugh] [whispers] [super happy] [sigh] [excited] [heavy breathing]
[angry] [sleepy] [crying] [surprise] ...
```
Drop them inline in the input text. Different shape from
IndexTTS-2's 8-vector emotion control — this is "say it like this"
markup directly in the prompt, with a far larger vocabulary.
- **Voice cloning** from ~5-15 s reference WAV.
- **Multi-speaker / multi-turn** generation natively supported.
- **80+ languages** (English-strong, not Chinese-leaning like
CosyVoice).
## API
OpenAI-compat at `http://10.100.79.3:8195`:
```bash
# Single-shot synthesis with paralinguistic tags inline.
curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"fish-s2","input":"Oh wow [super happy] I cannot believe it. [laugh] What a day.","voice":"glados","response_format":"wav"}' \
> out.wav
# Built-in voices.
curl http://10.100.79.3:8195/v1/audio/voices
# Streaming.
curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"fish-s2","input":"long passage…","voice":"glados","stream":true}' \
| mpv --no-cache -
```
WebUI at `/`. OpenAPI / docs at `/docs`. Healthcheck at `/v1/health`.
## Voice library
Drop reference WAV / MP3 / FLAC into
`/worktank/fish-s2/references/` on the host. The wrapper scans on
request — no restart needed. Use clean ~5-15 s clips, single
speaker, ideally with diverse intonation samples.
## Deploy
```bash
scripts/elway irv-ml1 --playbook playbooks/deploy-fish-s2.yaml
```
First boot pulls the s2-pro checkpoint (~9 GB BF16) into the HF
cache + warms torch.compile (adds ~60 s). Both are cached afterwards.
## Hardware footprint
- **VRAM**: ~17 GB practical, 24 GB recommended. Pinned to GPU 1
(RTX A6000) by default — plenty of headroom for long contexts and
large mmproj if a future checkpoint adds vision.
- **Disk**: ~9 GB for the model checkpoint + HF cache.