Files
esh-pfi-infrastructure/stacks/index-tts
vh 569e1af9ca feat(homepage): split AI fleet into role-based groups on a dedicated AI tab
Move the ~22-service flat "AI Systems" group off the Main tab into a new
four-tab layout (Main / AI / Infrastructure / Toolchain). The AI tab sorts
the inference fleet by function into seven groups:

  AI - Inference        gen, char-rp, char-rp-reasoning, Granite summarizer
  AI - Eval & Retrieval Selene, Skywork Reward, Qwen3 rerank/embed, image-bench
  AI - Gateways & Chat  LiteLLM, Asset Engine, Gateway Chat, Open WebUI, ...
  AI - Speech (TTS)     Chatterbox Fast, Kokoro, mOrpheus
  AI - Audio Tools      Parakeet ASR, YT Voice Clipper
  AI - Image & Media    ComfyUI, Arbo
  AI - Dormant          stopped rollback seats + retired auditions

Relabel each stack's homepage.group so canonical stacks/ matches the live
containers on ana-ml2, ana-docker, and irv-ml1. Dormant stacks were refreshed
with `docker compose up --no-start` so they carry the new label while staying
stopped (compose-start rollback preserved). settings.yaml drives tab/order/
columns; services.yaml and README updated to the new scheme.
2026-07-14 20:05:50 -07:00
..

IndexTTS-2

Bilibili's emotion-controllable zero-shot TTS (paper, code, weights) served behind our own thin FastAPI wrapper.

Why this stack exists alongside the other two TTS

CosyVoice 3 Qwen3-TTS-1.7B-Base IndexTTS-2
Voice cloning (-Base variant only)
English quality medium (Chinese-leaning) high (English-first) medium (better than CosyVoice)
Emotion control instruct mode is Chinese-only inline tags explicit: audio / 8-vector / text
Duration control implicit implicit explicit token-count mode
License Apache 2.0 Apache 2.0 custom (Bilibili — free at our scale)
Wrapper bare CosyVoice CLI groxaxo upstream FastAPI ours, in this dir

The differentiator is disentangled emotion. IndexTTS-2 lets you clone a voice's timbre from one reference and the emotion from a different reference — or skip emotion-audio entirely and supply an 8-vector or a text description. Neither of the other two does this cleanly in English.

API

OpenAI-compat-ish:

# List available voices + emotions
curl http://10.100.79.3:8192/v1/voices

# Basic synthesis (uses speaker WAV's natural emotion)
curl -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "input": "I have all the time in the world.",
    "voice": "glados"
  }' > glados.wav

# Same speaker, emotion taken from a separate reference WAV
curl -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "input": "I have all the time in the world.",
    "voice": "glados",
    "emotion_voice": "menacing",
    "emotion_alpha": 0.9
  }' > glados-menacing.wav

# Same speaker, emotion as 8-vector
# Order: happy, angry, sad, afraid, disgusted, melancholic, surprised, calm
curl -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "input": "I have all the time in the world.",
    "voice": "glados",
    "emotion_vector": [0, 0.7, 0, 0, 0.2, 0, 0, 0]
  }' > glados-angry.wav

# Same speaker, emotion derived from text by bundled QwenEmotion model
curl -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "input": "I have all the time in the world.",
    "voice": "glados",
    "emotion_text": "she said with quiet menace"
  }' > glados-menacing.wav

# Health
curl http://10.100.79.3:8192/healthz

Output is always WAV (PCM_16, 22050 Hz mono — IndexTTS-2's native rate). response_format other than wav is rejected.

Streaming (since 0.2.0)

Add "stream": true to any request to stream the WAV as it generates. IndexTTS-2 emits one chunk per text segment (~120 tokens) as soon as it finishes synthesizing it; long inputs start playing while the rest is still being generated.

# Pipe straight into a player. Time-to-first-audio drops from
# whole-file latency to ~one-segment latency.
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"input":"long passage of text...","voice":"glados","stream":true}' \
  | mpv --no-cache -

# Or save while playing (tee).
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"input":"...","voice":"glados","stream":true}' \
  | tee out.wav | mpv -

The streaming WAV uses a placeholder data-length in the header (0xFFFFFFFF) so chunks can be written before the total is known. Players that read until EOF (mpv, ffplay, aplay, sox, browsers via <audio>) handle this fine. Strict parsers (some metadata extractors, foobar2000 default settings) may complain about the size. If that matters, request without stream and you get a normal closed-length WAV.

Voice library

Flat dirs on the host (bind-mounted; survives container recreates):

/worktank/index-tts/voices/<name>.wav        # timbre references
/worktank/index-tts/emotions/<name>.wav      # emotion references

Drop a new WAV in either dir and /v1/voices picks it up immediately — no restart. Use clean reference clips, 5-30 s each, single speaker. Cloned voices live under restic; cache (model weights) is excluded.

Deploy

scripts/elway irv-ml1 --playbook playbooks/deploy-index-tts.yaml

First build: ~5-10 min for the docker image (CUDA torch + IndexTTS deps), plus ~5-7 GB model download on first container start. Subsequent starts: ~30 s warmup.

Switching GPUs

irv-ml1 has an RTX 3090 (cuda:0) + RTX A6000 (cuda:1). Default is auto-pick (cuda:0). To pin to the A6000 alongside Qwen3-TTS-on-3090:

ssh irv-ml1 '
  cd /opt/docker/compose/index-tts
  sed -i "s|^INDEX_TTS_DEVICE=.*|INDEX_TTS_DEVICE=cuda:1|" .env
  docker compose up -d
'

Gotchas

  • License — IndexTeam/IndexTTS-2 ships under a custom Bilibili license, not Apache/MIT. Free at our scale (the commercial tier kicks in at 100M MAU / RMB 1B revenue). Restricts using outputs to train other AI models. Read INDEX_MODEL_LICENSE in the HF repo before using outputs anywhere external.

  • Sample rate — 22050 Hz is hardcoded upstream. If you need 24 kHz or 48 kHz, resample in the caller.

  • Emotion-source precedence — if multiple emotion controls are specified in one request, the first non-empty one wins in this order: emotion_voice > emotion_vector > emotion_text. The others are silently ignored.

  • Model download — happens in the entrypoint on first start; the config.yaml file in the cache dir is the gate. To force a re-download, delete that file and recreate the container.

  • Bundled example WAVs are LFS pointers, not audio. Upstream stores examples/emo_*.wav and examples/voice_*.wav as Git LFS objects. The image clones the repo without git lfs pull (the index-tts org has exhausted GitHub's LFS bandwidth budget repeatedly, so doing it in the Dockerfile aborts the build). If you want the IndexTTS-2 example clips as starter material, fetch them once via the media CDN — that's a separate code path that doesn't count against the LFS API budget:

    ssh irv-ml1 '
      cd /worktank/index-tts/emotions
      curl -fsSL -o hate.wav https://media.githubusercontent.com/media/index-tts/index-tts/main/examples/emo_hate.wav
      curl -fsSL -o sad.wav  https://media.githubusercontent.com/media/index-tts/index-tts/main/examples/emo_sad.wav
    '
    
  • HF cache pinninginfer_v2.py pins HF_HUB_CACHE at import time to ./checkpoints/hf_cache. The wrapper sets this env var before importing, so auxiliary HF assets (MaskGCT, campplus, BigVGAN, w2v-bert) land alongside the IndexTTS-2 weights and are excluded from restic together.