stacks/index-tts: own FastAPI wrapper for IndexTTS-2 + deploy playbook

Adds a third TTS to the irv-ml1 fleet. IndexTTS-2 is Bilibili's
emotion-controllable zero-shot TTS (paper 2506.21619). Distinguishing
capability vs the existing two: timbre and emotion are disentangled —
clone a voice's timbre from one reference and the emotion from a
different reference, OR set emotion via 8-vector, OR derive it from a
text description. Neither CosyVoice 3 nor Qwen3-TTS-1.7B-Base does
this cleanly in English.

Wrapper is owned end-to-end (~150 lines in app.py) — the only existing
FastAPI fork (csllpr/index-tts-fastapi) targets v1 and is a dormant
single-commit repo. Upstream IndexTTS-2 ships only a Gradio webui.

Layout follows the qwen3-tts pattern:
  stacks/index-tts/
    Dockerfile           — CUDA 12.8 base, IndexTTS pinned to a SHA
    app.py               — FastAPI: POST /v1/audio/speech + /v1/voices
    entrypoint.sh        — one-time HF snapshot_download of the weights
    compose.yaml         — env-driven, GPU pinning support, bind mounts
    .env.example         — port 8192, fp16, paths
    README.md            — API examples + comparison vs the other TTS
  playbooks/deploy-index-tts.yaml  — elway playbook for irv-ml1

Voice and emotion libraries are flat host dirs of WAVs, bind-mounted.
Drop a new <name>.wav and /v1/voices picks it up immediately.

License caveat: IndexTTS-2 weights ship under a custom Bilibili
license (free at our scale, not OSI-open). README documents it.
This commit is contained in:
vh
2026-04-25 10:38:14 -07:00
parent 1f14c6d959
commit b75f020cc9
7 changed files with 604 additions and 0 deletions
+53
View File
@@ -0,0 +1,53 @@
# IndexTTS-2 stack tunables. Copy to `.env` on irv-ml1 before deploying.
# ── build pin ────────────────────────────────────────────────────────
# SHA of index-tts/index-tts to build from. Bump + rebuild when you want
# upstream wrapper updates.
INDEX_TTS_SHA=830f6f8f94a51fea23ab1d639027a86200075a4e
# Local image tag — bump when you change build context (Dockerfile,
# app.py, entrypoint.sh) to force a fresh layer build.
INDEX_TTS_TAG=v1
# ── network ──────────────────────────────────────────────────────────
# Host port. Container listens on 8000 internally.
# Reserved on irv-ml1: 8188 ComfyUI, 8190 CosyVoice, 8191 Qwen3-TTS,
# 8765 Parakeet. 8192 is open.
INDEX_TTS_PORT=8192
# Bind address. 0.0.0.0 exposes on all interfaces (incl. WG tunnel
# interface 10.100.79.3); 127.0.0.1 restricts to local-only.
INDEX_TTS_BIND=0.0.0.0
# ── runtime / GPU ────────────────────────────────────────────────────
# bf16/fp16 vs fp32. Empty = fp32. "1" = fp16. IndexTTS-2 README
# recommends fp16; ~6-10 GB VRAM at fp16, double at fp32.
INDEX_TTS_FP16=1
# GPU pinning. Empty = let IndexTTS auto-pick (cuda:0). Set to "cuda:1"
# to pin to the second GPU (irv-ml1's RTX 3090 vs RTX A6000).
INDEX_TTS_DEVICE=
# Devices visible inside the container. Either "all" (both GPUs) or a
# comma-separated list of indices (e.g. "1" to expose only the second
# card). The DEVICE setting above further restricts within those.
INDEX_TTS_GPU_DEVICES=all
# Logging level for the wrapper itself (IndexTTS internals are noisier
# regardless).
INDEX_TTS_LOG_LEVEL=INFO
# ── persistent storage on the host ───────────────────────────────────
# Model weights (~5-7 GB after first run). Bind-mounted so model state
# survives container recreate. Excluded from restic (regenerable from HF).
INDEX_TTS_CACHE_DIR=/worktank/index-tts/cache
# Voice library — flat dir of <name>.wav files (timbre references).
# Cloned voices need the original reference audio to recreate; included
# in restic.
INDEX_TTS_VOICES_DIR=/worktank/index-tts/voices
# Emotion library — flat dir of <name>.wav files (emotion references,
# typically short clips with strong affect). Optional — without any
# entries here you can still use emotion_vector or emotion_text.
INDEX_TTS_EMOTIONS_DIR=/worktank/index-tts/emotions