Files
esh-pfi-infrastructure/stacks/kyutai-tts/compose.yaml
T
vh 16d018ff96 stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B
Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:

* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
  4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
  Headline: 15,000+ paralinguistic / emotion tags via natural language
  ([laugh] [whispers] [super happy] etc.) — a step-function over
  Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
  EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
  open. ~17 GB VRAM.

* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
  4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
  win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
  (EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
  serving stack) — published Docker image, no local build. ~16 GB VRAM.
  CC BY-NC license — personal/research use only; flagged in README.

* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
  Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
  solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
  official deploy is Rust + websockets only; using NillPointer's
  community OpenAI-compat wrapper to bridge to /v1/audio/speech so
  it slots into the same bench harness. ~4-6 GB VRAM.

Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.

Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
2026-04-27 22:40:10 -07:00

57 lines
2.5 KiB
YAML

# Kyutai TTS — 1.6B / 2B-class streaming TTS from Kyutai (the Moshi /
# Mimi team), trained on 2.5M hours. 220 ms latency in solo setup; up
# to 32 simultaneous streams under 350 ms on an L40-class GPU.
#
# Served via the NillPointer/Kyutai-TTS-Server community wrapper —
# Kyutai's official deployment is Rust + websockets only, which doesn't
# fit our OpenAI-compat fleet. The community wrapper bridges Kyutai's
# native streaming to the OpenAI /v1/audio/speech contract.
#
# Why this stack alongside the existing TTS:
# * Kyutai's claim is the lowest streaming latency in this size
# class (220 ms on a single GPU). Worth bench-comparing against
# Chatterbox (~1.2 s) and Fish S2-Pro (~150 ms claimed).
# * Trained on 2.5M hours — a different scaling regime from the
# others (CosyVoice 5k hrs, Fish 10M hrs).
# * Designed for full-duplex dialogue (Moshi heritage) — may surface
# conversational quality the others lack.
#
# All tunables live in .env — edit that, not this file.
services:
kyutai-tts:
image: local/kyutai-tts:${KYUTAI_TTS_TAG}
build:
context: https://github.com/NillPointer/Kyutai-TTS-Server.git#${KYUTAI_TTS_SHA}
dockerfile: Dockerfile
container_name: kyutai-tts
restart: unless-stopped
runtime: nvidia
ports:
- "${KYUTAI_TTS_BIND:-0.0.0.0}:${KYUTAI_TTS_PORT}:8000"
environment:
- NVIDIA_VISIBLE_DEVICES=${KYUTAI_TTS_GPU_DEVICES:-0}
# Kyutai's en/fr bilingual model on HF. Switch to a different
# checkpoint via .env without rebuilding.
- KYUTAI_MODEL=${KYUTAI_TTS_MODEL:-kyutai/tts-1.6b-en_fr}
- HF_HOME=/app/hf_cache
volumes:
- ${KYUTAI_TTS_CACHE_DIR}:/app/hf_cache
- ${KYUTAI_TTS_VOICES_DIR}:/app/voices:ro
healthcheck:
# The wrapper exposes /v1/models for OpenAI-compat — same shape
# as Voxtral / Qwen3-TTS. Use that as the readiness signal.
# 127.0.0.1 explicit to dodge IPv4/IPv6 localhost race.
test: ["CMD-SHELL", "python3 -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/v1/models', timeout=5).status==200 else 1)\""]
interval: 30s
timeout: 10s
retries: 3
# First boot pulls the Kyutai checkpoint (~3-6 GB) + warms.
start_period: 600s
labels:
- homepage.group=AI Systems
- homepage.name=Kyutai TTS
- homepage.icon=mdi-radio-tower
- homepage.description=Ultra-low-latency streaming TTS — 220 ms on solo GPU, EN/FR (irv-ml1)
- homepage.href=http://10.100.79.3:${KYUTAI_TTS_PORT}