stacks/{fish-s2,voxtral,kyutai-tts}: three new TTS deploys for irv-ml1 quality A/B

Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:

* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
  4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
  Headline: 15,000+ paralinguistic / emotion tags via natural language
  ([laugh] [whispers] [super happy] etc.) — a step-function over
  Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
  EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
  open. ~17 GB VRAM.

* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
  4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
  win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
  (EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
  serving stack) — published Docker image, no local build. ~16 GB VRAM.
  CC BY-NC license — personal/research use only; flagged in README.

* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
  Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
  solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
  official deploy is Rust + websockets only; using NillPointer's
  community OpenAI-compat wrapper to bridge to /v1/audio/speech so
  it slots into the same bench harness. ~4-6 GB VRAM.

Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.

Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
This commit is contained in:
vh
2026-04-27 22:40:10 -07:00
parent db42a7cc17
commit 16d018ff96
12 changed files with 872 additions and 0 deletions
+38
View File
@@ -0,0 +1,38 @@
# Voxtral TTS stack tunables. Copy to `.env` on irv-ml1 before
# deploying.
# ── image pin ────────────────────────────────────────────────────────
# vLLM-Omni image tag (Mistral's partner serving stack for Voxtral).
# Use a specific version rather than `latest` — vLLM moves fast and
# Voxtral has version-specific compatibility.
VOXTRAL_VLLM_TAG=latest
# Voxtral model on Hugging Face. The 4B variant is the only released
# checkpoint as of 2026-04. Default BF16 weights are ~8 GB.
VOXTRAL_MODEL=mistralai/Voxtral-4B-TTS-2603
# ── network ──────────────────────────────────────────────────────────
# Host port (container listens on 8000 internally).
VOXTRAL_PORT=8197
VOXTRAL_BIND=0.0.0.0
# ── runtime / GPU ────────────────────────────────────────────────────
# GPU pinning. "0" = RTX 3090 (24 GB), "1" = RTX A6000 (48 GB).
# Voxtral 4B BF16 needs ~16 GB practical (model + KV + activation).
# Pinned to A6000 by default for headroom. The 3090 fits but is tight
# for long streaming sessions.
VOXTRAL_GPU_DEVICES=1
# vLLM GPU memory utilization fraction (0.0-1.0). 0.85 = leave 15%
# headroom for other processes / KV cache spikes. Lower if running
# alongside other GPU workloads on the same device.
VOXTRAL_GPU_UTIL=0.85
# ── persistent storage on the host ───────────────────────────────────
# HF cache — first start pulls the Voxtral checkpoint (~8 GB) into
# this dir. Persistent across container recreates.
VOXTRAL_CACHE_DIR=/worktank/voxtral/hf_cache
# Reference voices for cloning. Read-only mount inside the container.
# Drop ~5-15 s WAV / FLAC clips here.
VOXTRAL_VOICES_DIR=/worktank/voxtral/voices