qwen3-tts: add stack + deploy playbook for irv-ml1
Alibaba's open-weight TTS (Apache 2.0, Jan 2026), deployed via groxaxo/Qwen3-TTS-Openai-Fastapi wrapper. Built locally from a pinned git SHA via docker buildx's git context — no source vendored. 1.7B flagship model by default; 0.6B available via QWEN3_TTS_MODEL env override. Why we need a second TTS stack: cosyvoice 3 emits Chinese phonemes for English content per upstream FunAudioLLM/CosyVoice#1790 (unfixed). Qwen3-TTS is from the same Alibaba team but with English first-class in the checkpoint — 10 languages, 97 ms streaming TTFB, instruction-driven emotion. Coexists with cosyvoice on irv-ml1 (port 8191; cosyvoice keeps 8190). Voice cloning shape DIFFERS from cosyvoice: profile-based, not voice-id. Profiles live under voice_library/profiles/<name>/ and are referenced as voice="clone:<name>". Path layout: /worktank/qwen3-tts/{cache,voices}/, with cache excluded from restic (regenerable from HF Hub) and voices included (cloned profiles need original reference audio to recreate). playbooks/deploy-qwen3-tts.yaml: 10 steps + 5 verify, idempotent; the wait step polls /health for up to ~10 min to absorb first-run model download. Stack only — restic profile update for /worktank/qwen3-tts/voices/ to follow when this is empirically validated against the GLaDOS voice (the "did Qwen inherit the Chinese-bias bug?" question).
This commit is contained in:
@@ -0,0 +1,52 @@
|
||||
# Qwen3-TTS stack tunables. Copy to `.env` on irv-ml1 before deploying.
|
||||
|
||||
# ── build pin ────────────────────────────────────────────────────────
|
||||
# SHA of groxaxo/Qwen3-TTS-Openai-Fastapi to build from. Bump + rebuild
|
||||
# when you want upstream wrapper updates.
|
||||
QWEN3_TTS_SHA=10323ce778c48a75dbda93d0a4891983fb371f58
|
||||
|
||||
# Local image tag — bump when you change build context to force a
|
||||
# fresh layer build.
|
||||
QWEN3_TTS_TAG=v1
|
||||
|
||||
# ── network ──────────────────────────────────────────────────────────
|
||||
# Host port (container listens on 8880 internally).
|
||||
QWEN3_TTS_PORT=8191
|
||||
|
||||
# Bind address. 0.0.0.0 exposes on all interfaces (incl. WG tunnel
|
||||
# interface 10.100.79.3); 127.0.0.1 restricts to local-only.
|
||||
QWEN3_TTS_BIND=0.0.0.0
|
||||
|
||||
# ── runtime ──────────────────────────────────────────────────────────
|
||||
# Inference backend. `official` = default upstream; `optimized` =
|
||||
# faster but slightly less robust; `vllm_omni` = vLLM-backed (needs
|
||||
# more VRAM); `pytorch` = bare pytorch path.
|
||||
QWEN3_TTS_BACKEND=official
|
||||
|
||||
# Model variant. 1.7B = flagship, 6–8 GB VRAM with bfloat16, best
|
||||
# quality + control. 0.6B = lightweight, ~2–3 GB VRAM, faster, slightly
|
||||
# less expressive.
|
||||
QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-1.7B
|
||||
|
||||
# Warm the model on container start so the first synthesis request
|
||||
# doesn't pay the load latency. Adds ~30 s to startup. Recommended.
|
||||
QWEN3_TTS_WARMUP=true
|
||||
|
||||
# Concurrency cap on synthesis requests. Single GPU + 1.7B model →
|
||||
# leave at 1 unless you're load-testing.
|
||||
QWEN3_TTS_MAX_CONCURRENT=1
|
||||
|
||||
# Mount the gradio voice-studio UI at /voice-studio for browser-side
|
||||
# voice cloning. Set "false" to disable for headless deployments.
|
||||
QWEN3_TTS_VOICE_STUDIO=true
|
||||
|
||||
# ── persistent storage on the host ───────────────────────────────────
|
||||
# HuggingFace cache (model weights, ~5 GB after first run). Bind-mounted
|
||||
# so model state survives container recreate. Excluded from restic
|
||||
# (regenerable from HF Hub).
|
||||
QWEN3_TTS_CACHE_DIR=/worktank/qwen3-tts/cache
|
||||
|
||||
# Cloned voice profiles (meta.json + reference.wav per voice). Precious
|
||||
# — cloned voices need the original reference audio to recreate.
|
||||
# Included in restic.
|
||||
QWEN3_TTS_VOICES_DIR=/worktank/qwen3-tts/voices
|
||||
Reference in New Issue
Block a user