stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B
Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:
* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
Headline: 15,000+ paralinguistic / emotion tags via natural language
([laugh] [whispers] [super happy] etc.) — a step-function over
Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
open. ~17 GB VRAM.
* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
(EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
serving stack) — published Docker image, no local build. ~16 GB VRAM.
CC BY-NC license — personal/research use only; flagged in README.
* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
official deploy is Rust + websockets only; using NillPointer's
community OpenAI-compat wrapper to bridge to /v1/audio/speech so
it slots into the same bench harness. ~4-6 GB VRAM.
Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.
Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
This commit is contained in:
@@ -0,0 +1,32 @@
|
||||
# Kyutai TTS stack tunables. Copy to `.env` on irv-ml1 before
|
||||
# deploying.
|
||||
|
||||
# ── build pin ────────────────────────────────────────────────────────
|
||||
# SHA of NillPointer/Kyutai-TTS-Server to build from. Use the FULL
|
||||
# 40-char SHA — docker buildx's git source resolver doesn't accept
|
||||
# short hashes.
|
||||
KYUTAI_TTS_SHA=main
|
||||
|
||||
# Local image tag — bump when you change build context to force a
|
||||
# fresh layer build.
|
||||
KYUTAI_TTS_TAG=v1
|
||||
|
||||
# Kyutai model on HF. Available variants:
|
||||
# kyutai/tts-1.6b-en_fr — bilingual EN/FR, 1.6B params (default)
|
||||
# kyutai/pocket-tts — 100M, CPU-realtime, EN-only (lighter alt)
|
||||
KYUTAI_TTS_MODEL=kyutai/tts-1.6b-en_fr
|
||||
|
||||
# ── network ──────────────────────────────────────────────────────────
|
||||
KYUTAI_TTS_PORT=8198
|
||||
KYUTAI_TTS_BIND=0.0.0.0
|
||||
|
||||
# ── runtime / GPU ────────────────────────────────────────────────────
|
||||
# GPU pinning. "0" = RTX 3090 (24 GB), "1" = RTX A6000 (48 GB).
|
||||
# 1.6B Kyutai needs ~4-6 GB practical, fits comfortably on either.
|
||||
# Pinned to GPU 0 (3090) by default since the A6000 is hosting the
|
||||
# heavier Fish S2-Pro / Voxtral.
|
||||
KYUTAI_TTS_GPU_DEVICES=0
|
||||
|
||||
# ── persistent storage on the host ───────────────────────────────────
|
||||
KYUTAI_TTS_CACHE_DIR=/worktank/kyutai-tts/hf_cache
|
||||
KYUTAI_TTS_VOICES_DIR=/worktank/kyutai-tts/voices
|
||||
@@ -0,0 +1,83 @@
|
||||
# Kyutai TTS
|
||||
|
||||
[kyutai/tts-1.6b-en_fr](https://huggingface.co/kyutai/tts-1.6b-en_fr)
|
||||
— Kyutai's flagship streaming TTS (1.6B params, EN/FR bilingual,
|
||||
trained on 2.5M hours), served via the
|
||||
[NillPointer/Kyutai-TTS-Server](https://github.com/NillPointer/Kyutai-TTS-Server)
|
||||
community OpenAI-compatible wrapper.
|
||||
|
||||
## Why this stack exists
|
||||
|
||||
Kyutai's claim is the **lowest streaming latency in this size class**:
|
||||
220 ms in solo setup; up to 32 simultaneous streams under 350 ms on
|
||||
a single L40-class GPU. Worth bench-comparing against:
|
||||
|
||||
| | claimed latency | use case |
|
||||
|---|---|---|
|
||||
| **Kyutai TTS** | **~220 ms** | streaming EN/FR, conversational dialogue heritage |
|
||||
| Kokoro | ~26 ms TTFB measured | low-latency English, fixed voice library |
|
||||
| Chatterbox Turbo | ~1.2 s TTFB measured | English w/ cloning + 9 paralinguistic tags |
|
||||
| Fish Audio S2-Pro | ~150 ms claimed | richest paralinguistic English |
|
||||
| Voxtral | ~70 ms model latency | multilingual EN/FR/DE/ES/IT/PT/NL/HI |
|
||||
|
||||
## Deployment notes
|
||||
|
||||
Kyutai's official deployment path is **Rust + websockets only** (no
|
||||
HTTP, no OpenAI-compat). That doesn't fit the OpenAI-`/v1/audio/speech`
|
||||
contract the rest of our TTS fleet uses. The
|
||||
NillPointer/Kyutai-TTS-Server community wrapper bridges Kyutai's
|
||||
native streaming to the OpenAI HTTP shape, which lets us slot it
|
||||
into the same bench harness as the others.
|
||||
|
||||
**Tradeoff**: the wrapper adds Python overhead on the request path,
|
||||
so measured latency on this stack will be *higher* than Kyutai's
|
||||
220 ms claim (which is for the bare Rust server). If we measure
|
||||
~400-500 ms TTFB end-to-end, the wrapper is the floor — Kyutai itself
|
||||
is hitting its target.
|
||||
|
||||
## Architecture heritage
|
||||
|
||||
Kyutai's TTS shares the **Mimi** neural codec + **Moshi** dialogue
|
||||
modeling framework. Both target full-duplex conversational AI (Moshi
|
||||
is their flagship speech-text foundation model). The TTS-only model
|
||||
is the "synthesis half" of the stack, distilled for low-latency
|
||||
streaming.
|
||||
|
||||
Trained on **2.5M hours** — a different scaling regime from the
|
||||
others (CosyVoice 5k, Fish 10M, Voxtral undisclosed).
|
||||
|
||||
## API
|
||||
|
||||
OpenAI-compat at `http://10.100.79.3:8198`:
|
||||
|
||||
```bash
|
||||
# Single-shot synthesis.
|
||||
curl -fsS -X POST http://10.100.79.3:8198/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"tts-1.6b-en_fr","input":"Hello there.","voice":"default","response_format":"wav"}' \
|
||||
> out.wav
|
||||
|
||||
# Streaming.
|
||||
curl -fsS -X POST http://10.100.79.3:8198/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"tts-1.6b-en_fr","input":"long passage…","voice":"default","stream":true}' \
|
||||
| mpv --no-cache -
|
||||
|
||||
# Built-in voices.
|
||||
curl http://10.100.79.3:8198/v1/audio/voices
|
||||
```
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
scripts/elway irv-ml1 --playbook playbooks/deploy-kyutai-tts.yaml
|
||||
```
|
||||
|
||||
First boot pulls the kyutai/tts-1.6b-en_fr checkpoint (~3-6 GB) into
|
||||
the HF cache.
|
||||
|
||||
## Hardware footprint
|
||||
|
||||
- **VRAM**: ~4-6 GB practical. Pinned to GPU 0 (RTX 3090) by default
|
||||
— the A6000 is hosting the heavier Fish S2-Pro / Voxtral.
|
||||
- **Disk**: ~6 GB for the checkpoint + HF cache.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Kyutai TTS — 1.6B / 2B-class streaming TTS from Kyutai (the Moshi /
|
||||
# Mimi team), trained on 2.5M hours. 220 ms latency in solo setup; up
|
||||
# to 32 simultaneous streams under 350 ms on an L40-class GPU.
|
||||
#
|
||||
# Served via the NillPointer/Kyutai-TTS-Server community wrapper —
|
||||
# Kyutai's official deployment is Rust + websockets only, which doesn't
|
||||
# fit our OpenAI-compat fleet. The community wrapper bridges Kyutai's
|
||||
# native streaming to the OpenAI /v1/audio/speech contract.
|
||||
#
|
||||
# Why this stack alongside the existing TTS:
|
||||
# * Kyutai's claim is the lowest streaming latency in this size
|
||||
# class (220 ms on a single GPU). Worth bench-comparing against
|
||||
# Chatterbox (~1.2 s) and Fish S2-Pro (~150 ms claimed).
|
||||
# * Trained on 2.5M hours — a different scaling regime from the
|
||||
# others (CosyVoice 5k hrs, Fish 10M hrs).
|
||||
# * Designed for full-duplex dialogue (Moshi heritage) — may surface
|
||||
# conversational quality the others lack.
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
services:
|
||||
kyutai-tts:
|
||||
image: local/kyutai-tts:${KYUTAI_TTS_TAG}
|
||||
build:
|
||||
context: https://github.com/NillPointer/Kyutai-TTS-Server.git#${KYUTAI_TTS_SHA}
|
||||
dockerfile: Dockerfile
|
||||
container_name: kyutai-tts
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
ports:
|
||||
- "${KYUTAI_TTS_BIND:-0.0.0.0}:${KYUTAI_TTS_PORT}:8000"
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${KYUTAI_TTS_GPU_DEVICES:-0}
|
||||
# Kyutai's en/fr bilingual model on HF. Switch to a different
|
||||
# checkpoint via .env without rebuilding.
|
||||
- KYUTAI_MODEL=${KYUTAI_TTS_MODEL:-kyutai/tts-1.6b-en_fr}
|
||||
- HF_HOME=/app/hf_cache
|
||||
volumes:
|
||||
- ${KYUTAI_TTS_CACHE_DIR}:/app/hf_cache
|
||||
- ${KYUTAI_TTS_VOICES_DIR}:/app/voices:ro
|
||||
healthcheck:
|
||||
# The wrapper exposes /v1/models for OpenAI-compat — same shape
|
||||
# as Voxtral / Qwen3-TTS. Use that as the readiness signal.
|
||||
# 127.0.0.1 explicit to dodge IPv4/IPv6 localhost race.
|
||||
test: ["CMD-SHELL", "python3 -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/v1/models', timeout=5).status==200 else 1)\""]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# First boot pulls the Kyutai checkpoint (~3-6 GB) + warms.
|
||||
start_period: 600s
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=Kyutai TTS
|
||||
- homepage.icon=mdi-radio-tower
|
||||
- homepage.description=Ultra-low-latency streaming TTS — 220 ms on solo GPU, EN/FR (irv-ml1)
|
||||
- homepage.href=http://10.100.79.3:${KYUTAI_TTS_PORT}
|
||||
Reference in New Issue
Block a user