stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B
Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:
* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
Headline: 15,000+ paralinguistic / emotion tags via natural language
([laugh] [whispers] [super happy] etc.) — a step-function over
Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
open. ~17 GB VRAM.
* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
(EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
serving stack) — published Docker image, no local build. ~16 GB VRAM.
CC BY-NC license — personal/research use only; flagged in README.
* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
official deploy is Rust + websockets only; using NillPointer's
community OpenAI-compat wrapper to bridge to /v1/audio/speech so
it slots into the same bench harness. ~4-6 GB VRAM.
Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.
Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
This commit is contained in:
@@ -0,0 +1,105 @@
|
||||
# Fish Audio S2-Pro
|
||||
|
||||
[fishaudio/s2-pro](https://huggingface.co/fishaudio/s2-pro) — the most
|
||||
expressive open-source TTS model as of 2026-04, served via the
|
||||
official [fishaudio/fish-speech](https://github.com/fishaudio/fish-speech)
|
||||
inference engine.
|
||||
|
||||
## Why this stack exists
|
||||
|
||||
Three of the existing TTS already cover the basics — Kokoro for raw
|
||||
speed, Chatterbox for speed-with-cloning, IndexTTS-2 for precision
|
||||
emotion control. Fish Audio S2-Pro fills a different slot:
|
||||
**dramatically richer paralinguistic control via natural-language
|
||||
tags** (15,000+ vs Chatterbox Turbo's 9 fixed tags), with comparable
|
||||
latency (~150 ms streaming) and voice cloning.
|
||||
|
||||
Released March 9, 2026; we missed it during the original irv-ml1
|
||||
build-out in early April.
|
||||
|
||||
| | use case |
|
||||
|---|---|
|
||||
| **Fish Audio S2-Pro** | richest emotive / paralinguistic English TTS — 15k+ tags |
|
||||
| Kokoro | low-latency English, fixed voice library |
|
||||
| Chatterbox Turbo | low-latency English w/ cloning + 9 paralinguistic tags |
|
||||
| IndexTTS-2 | English voice cloning + emotion vector / text control |
|
||||
| Qwen3-TTS-1.7B | English voice cloning (slow on official backend) |
|
||||
| CosyVoice 3 | multilingual (Chinese-leaning) |
|
||||
| VibeVoice 1.5B | long-form / multi-speaker dialogue |
|
||||
|
||||
## Architecture
|
||||
|
||||
Dual-AR (Slow + Fast):
|
||||
- **Slow AR** operates along the time axis, predicts the primary
|
||||
semantic codebook.
|
||||
- **Fast AR** generates the remaining 9 residual codebooks per time
|
||||
step, reconstructing fine-grained acoustic detail.
|
||||
|
||||
Trained on 10M+ hours of audio across 80+ languages with
|
||||
reinforcement-learning alignment. Win rates per upstream:
|
||||
|
||||
| benchmark | S2-Pro |
|
||||
|---|---|
|
||||
| EmergentTTS-Eval paralinguistics | 91.61% |
|
||||
| Blind A/B vs ElevenLabs Flash v2.5 (multilingual) | strong |
|
||||
|
||||
## Headline features
|
||||
|
||||
- **15,000+ paralinguistic / emotion tags** via natural language:
|
||||
```
|
||||
[laugh] [whispers] [super happy] [sigh] [excited] [heavy breathing]
|
||||
[angry] [sleepy] [crying] [surprise] ...
|
||||
```
|
||||
Drop them inline in the input text. Different shape from
|
||||
IndexTTS-2's 8-vector emotion control — this is "say it like this"
|
||||
markup directly in the prompt, with a far larger vocabulary.
|
||||
- **Voice cloning** from ~5-15 s reference WAV.
|
||||
- **Multi-speaker / multi-turn** generation natively supported.
|
||||
- **80+ languages** (English-strong, not Chinese-leaning like
|
||||
CosyVoice).
|
||||
|
||||
## API
|
||||
|
||||
OpenAI-compat at `http://10.100.79.3:8195`:
|
||||
|
||||
```bash
|
||||
# Single-shot synthesis with paralinguistic tags inline.
|
||||
curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"fish-s2","input":"Oh wow [super happy] I cannot believe it. [laugh] What a day.","voice":"glados","response_format":"wav"}' \
|
||||
> out.wav
|
||||
|
||||
# Built-in voices.
|
||||
curl http://10.100.79.3:8195/v1/audio/voices
|
||||
|
||||
# Streaming.
|
||||
curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"fish-s2","input":"long passage…","voice":"glados","stream":true}' \
|
||||
| mpv --no-cache -
|
||||
```
|
||||
|
||||
WebUI at `/`. OpenAPI / docs at `/docs`. Healthcheck at `/v1/health`.
|
||||
|
||||
## Voice library
|
||||
|
||||
Drop reference WAV / MP3 / FLAC into
|
||||
`/worktank/fish-s2/references/` on the host. The wrapper scans on
|
||||
request — no restart needed. Use clean ~5-15 s clips, single
|
||||
speaker, ideally with diverse intonation samples.
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
scripts/elway irv-ml1 --playbook playbooks/deploy-fish-s2.yaml
|
||||
```
|
||||
|
||||
First boot pulls the s2-pro checkpoint (~9 GB BF16) into the HF
|
||||
cache + warms torch.compile (adds ~60 s). Both are cached afterwards.
|
||||
|
||||
## Hardware footprint
|
||||
|
||||
- **VRAM**: ~17 GB practical, 24 GB recommended. Pinned to GPU 1
|
||||
(RTX A6000) by default — plenty of headroom for long contexts and
|
||||
large mmproj if a future checkpoint adds vision.
|
||||
- **Disk**: ~9 GB for the model checkpoint + HF cache.
|
||||
Reference in New Issue
Block a user