qwen3-tts: add stack + deploy playbook for irv-ml1
Alibaba's open-weight TTS (Apache 2.0, Jan 2026), deployed via groxaxo/Qwen3-TTS-Openai-Fastapi wrapper. Built locally from a pinned git SHA via docker buildx's git context — no source vendored. 1.7B flagship model by default; 0.6B available via QWEN3_TTS_MODEL env override. Why we need a second TTS stack: cosyvoice 3 emits Chinese phonemes for English content per upstream FunAudioLLM/CosyVoice#1790 (unfixed). Qwen3-TTS is from the same Alibaba team but with English first-class in the checkpoint — 10 languages, 97 ms streaming TTFB, instruction-driven emotion. Coexists with cosyvoice on irv-ml1 (port 8191; cosyvoice keeps 8190). Voice cloning shape DIFFERS from cosyvoice: profile-based, not voice-id. Profiles live under voice_library/profiles/<name>/ and are referenced as voice="clone:<name>". Path layout: /worktank/qwen3-tts/{cache,voices}/, with cache excluded from restic (regenerable from HF Hub) and voices included (cloned profiles need original reference audio to recreate). playbooks/deploy-qwen3-tts.yaml: 10 steps + 5 verify, idempotent; the wait step polls /health for up to ~10 min to absorb first-run model download. Stack only — restic profile update for /worktank/qwen3-tts/voices/ to follow when this is empirically validated against the GLaDOS voice (the "did Qwen inherit the Chinese-bias bug?" question).
This commit is contained in:
@@ -0,0 +1,63 @@
|
||||
# Qwen3-TTS — Alibaba's open-weight TTS, deployed via the
|
||||
# groxaxo/Qwen3-TTS-Openai-Fastapi wrapper.
|
||||
#
|
||||
# Why this stack exists alongside cosyvoice: CosyVoice 3 emits
|
||||
# Chinese phonemes for non-Chinese inputs (upstream issue
|
||||
# FunAudioLLM/CosyVoice#1790, no fix). Qwen3-TTS is from the same
|
||||
# Alibaba team but with English first-class — 10 languages, 97 ms
|
||||
# streaming TTFB, voice cloning, instruction-driven emotional
|
||||
# expression. Released Jan 2026, Apache 2.0.
|
||||
#
|
||||
# Build: no prebuilt image; pinned to a SHA via docker buildx's git
|
||||
# context URL so subsequent rebuilds are reproducible. ~5–10 min on
|
||||
# first build (CUDA torch + transformers).
|
||||
#
|
||||
# Model: 1.7B flagship (~6–8 GB VRAM with bfloat16) by default; the
|
||||
# host has plenty of VRAM. Switch to the 0.6B in .env if you ever
|
||||
# need more headroom.
|
||||
#
|
||||
# Voice cloning shape DIFFERS from cosyvoice: profile-based, not
|
||||
# voice-id. Profiles live under voice_library/profiles/<name>/ with
|
||||
# meta.json + reference.wav, and are referenced as
|
||||
# `voice="clone:<name>"` in /v1/audio/speech requests.
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
services:
|
||||
qwen3-tts:
|
||||
image: local/qwen3-tts:${QWEN3_TTS_TAG}
|
||||
build:
|
||||
context: https://github.com/groxaxo/Qwen3-TTS-Openai-Fastapi.git#${QWEN3_TTS_SHA}
|
||||
dockerfile: Dockerfile
|
||||
container_name: qwen3-tts
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
ports:
|
||||
- "${QWEN3_TTS_BIND:-0.0.0.0}:${QWEN3_TTS_PORT}:8880"
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=all
|
||||
- PORT=8880
|
||||
- TTS_BACKEND=${QWEN3_TTS_BACKEND:-official}
|
||||
- TTS_MODEL_NAME=${QWEN3_TTS_MODEL:-Qwen/Qwen3-TTS-12Hz-1.7B}
|
||||
- TTS_WARMUP_ON_START=${QWEN3_TTS_WARMUP:-true}
|
||||
- TTS_MAX_CONCURRENT=${QWEN3_TTS_MAX_CONCURRENT:-1}
|
||||
- ENABLE_VOICE_STUDIO=${QWEN3_TTS_VOICE_STUDIO:-true}
|
||||
- VOICE_LIBRARY_DIR=/root/qwen3-tts/voice_library
|
||||
- HF_HOME=/root/.cache/huggingface
|
||||
volumes:
|
||||
- ${QWEN3_TTS_CACHE_DIR}:/root/.cache/huggingface
|
||||
- ${QWEN3_TTS_VOICES_DIR}:/root/qwen3-tts/voice_library
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "curl -fsS http://localhost:8880/health >/dev/null || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# First boot pulls torch + Qwen3-TTS-12Hz-1.7B (~6 GB) and
|
||||
# optionally warms the model — give it a generous budget.
|
||||
start_period: 600s
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=Qwen3-TTS
|
||||
- homepage.icon=mdi-account-voice
|
||||
- homepage.description=Multilingual TTS with English-first emotion (irv-ml1)
|
||||
- homepage.href=http://10.100.79.3:${QWEN3_TTS_PORT}
|
||||
Reference in New Issue
Block a user