4549d241a7
Three TTS additions to round out coverage on irv-ml1, each filling a
distinct niche the existing slate doesn't own.
Final coverage matrix (all on irv-ml1):
Kokoro — low-latency English, fixed voice library, ~300ms TTFA
Chatterbox Turbo — low-latency English w/ voice cloning + paralinguistic tags
IndexTTS-2 — English voice cloning + emotion vector / text control
Qwen3-TTS-1.7B-Base — high-quality English voice cloning
CosyVoice 3 — multilingual (Chinese-leaning)
VibeVoice 1.5B — long-form / multi-speaker dialogue
stacks/kokoro:
- port 8193, GPU device 0 (3090)
- pulls ghcr.io/remsky/kokoro-fastapi-gpu:v0.2.4-master (no Dockerfile,
no first-run model download — models baked in)
- 60+ built-in voices, OpenAI-compat with stream=true over chunked HTTP
- Apache-2.0 weights + code, ~1 GB VRAM
stacks/vibevoice:
- port 8194, GPU device 1 (A6000 — for 7B headroom)
- builds groxaxo/VibeVoice-FastAPI1 (more current fork of ncoder-ai)
pinned to 7614c469a145
- default model microsoft/VibeVoice-1.5B (~7 GB bf16 VRAM); env var
swap to rsxdalv/VibeVoice-Large (7B) or FabioSarracino/VibeVoice-Large-Q8
- multi-speaker dialogue via /v1/vibevoice/generate with Speaker N: format
- long-form niche only — not low-latency
stacks/chatterbox:
- port 8196, GPU device 0 (3090)
- builds devnen/Chatterbox-TTS-Server (most active Turbo-supporting wrapper)
- default model ResembleAI/chatterbox-turbo (~2.5 GB fp16, ~75ms latency)
- paralinguistic tags inline ([laugh] [whisper] etc) — different shape
from IndexTTS-2's emotion vector; fills the speed+cloning niche
Kokoro/IndexTTS don't cover together
- mandatory PerTh watermark on outputs (Resemble policy)
Three matching playbooks under playbooks/deploy-{kokoro,vibevoice,
chatterbox}.yaml. All idempotent, creates-/when-gated.
Cold-deploy disk on /worktank/: ~7 GB Kokoro + ~19 GB VibeVoice 1.5B
+ ~12 GB Chatterbox = ~38 GB total. VRAM concurrent: ~10-11 GB across
both GPUs.
Skipped from the original four-stack proposal: VibeVoice Realtime
(overlaps Kokoro's niche; Kokoro wins on latency, license, and not
needing a build).
69 lines
2.9 KiB
YAML
69 lines
2.9 KiB
YAML
# VibeVoice 1.5B long-form TTS via groxaxo/VibeVoice-FastAPI1
|
|
# (fork of ncoder-ai/VibeVoice-FastAPI). Multi-speaker dialogue support
|
|
# via the extended /v1/vibevoice/generate endpoint with a Speaker N:
|
|
# script format. OpenAI-compat /v1/audio/speech also exposed.
|
|
#
|
|
# Why this stack exists alongside the other TTS:
|
|
# * Long-form / podcast-quality slot — VibeVoice is Microsoft's
|
|
# diffusion-based long-form TTS designed for multi-speaker output.
|
|
# * Dialogue mode: feed `Speaker 0: ... \n Speaker 1: ...` and the
|
|
# model handles voice switching natively.
|
|
# * Trade-off: NOT streaming-friendly — generation is single-shot
|
|
# latent denoising over the whole sequence, then vocode. For
|
|
# low-latency English, use Kokoro or Chatterbox Turbo instead.
|
|
#
|
|
# Image is built locally from the upstream Dockerfile via docker
|
|
# buildx git-context (no source vendored on the host). Pinned to a
|
|
# SHA in .env so rebuilds are reproducible.
|
|
#
|
|
# Default model is VibeVoice-1.5B (~7 GB bf16 VRAM). Switch to
|
|
# rsxdalv/VibeVoice-Large for the 7B variant (~18 GB) — pin to A6000
|
|
# in that case.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
vibevoice:
|
|
image: local/vibevoice:${VIBEVOICE_TAG}
|
|
build:
|
|
context: https://github.com/groxaxo/VibeVoice-FastAPI1.git#${VIBEVOICE_SHA}
|
|
dockerfile: Dockerfile
|
|
container_name: vibevoice
|
|
restart: unless-stopped
|
|
runtime: nvidia
|
|
ports:
|
|
- "${VIBEVOICE_BIND:-0.0.0.0}:${VIBEVOICE_PORT}:8001"
|
|
environment:
|
|
- NVIDIA_VISIBLE_DEVICES=${VIBEVOICE_GPU_DEVICES:-1}
|
|
- VIBEVOICE_MODEL_PATH=${VIBEVOICE_MODEL:-microsoft/VibeVoice-1.5B}
|
|
- VIBEVOICE_DEVICE=cuda
|
|
- VIBEVOICE_INFERENCE_STEPS=${VIBEVOICE_INFERENCE_STEPS:-10}
|
|
- VIBEVOICE_DTYPE=${VIBEVOICE_DTYPE:-bfloat16}
|
|
- VIBEVOICE_ATTN_IMPLEMENTATION=${VIBEVOICE_ATTN:-flash_attention_2}
|
|
- VIBEVOICE_QUANTIZATION=${VIBEVOICE_QUANT:-}
|
|
- TORCH_COMPILE=${VIBEVOICE_TORCH_COMPILE:-false}
|
|
- TORCH_COMPILE_MODE=${VIBEVOICE_TORCH_COMPILE_MODE:-default}
|
|
- VOICES_DIR=/app/voices
|
|
- DEFAULT_CFG_SCALE=${VIBEVOICE_CFG_SCALE:-1.8}
|
|
- MAX_GENERATION_LENGTH=${VIBEVOICE_MAX_GEN_LEN:-5400}
|
|
- HF_HOME=/root/.cache/huggingface
|
|
volumes:
|
|
- ${VIBEVOICE_VOICES_DIR}:/app/voices:ro
|
|
- ${VIBEVOICE_CACHE_DIR}:/root/.cache/huggingface
|
|
healthcheck:
|
|
# Upstream Dockerfile exposes /health.
|
|
test: ["CMD-SHELL", "curl -fsS -o /dev/null http://localhost:8001/health || exit 1"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
# First boot pulls VibeVoice-1.5B (~7 GB) on a cold cache and
|
|
# may also build flash-attn / torch.compile JIT cache on the
|
|
# first inference. Generous deadline.
|
|
start_period: 900s
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=VibeVoice
|
|
- homepage.icon=mdi-podcast
|
|
- homepage.description=Long-form / multi-speaker dialogue TTS (irv-ml1)
|
|
- homepage.href=http://10.100.79.3:${VIBEVOICE_PORT}
|