# Qwen3-TTS stack tunables. Copy to `.env` on irv-ml1 before deploying. # ── build pin ──────────────────────────────────────────────────────── # SHA of groxaxo/Qwen3-TTS-Openai-Fastapi to build from. Bump + rebuild # when you want upstream wrapper updates. QWEN3_TTS_SHA=10323ce778c48a75dbda93d0a4891983fb371f58 # Local image tag — bump when you change build context to force a # fresh layer build. v2 = first GPU build (target=production); v1 # was the accidental CPU-only image (last stage of upstream's # multi-stage Dockerfile). QWEN3_TTS_TAG=v2 # ── network ────────────────────────────────────────────────────────── # Host port (container listens on 8880 internally). QWEN3_TTS_PORT=8191 # Bind address. 0.0.0.0 exposes on all interfaces (incl. WG tunnel # interface 10.100.79.3); 127.0.0.1 restricts to local-only. QWEN3_TTS_BIND=0.0.0.0 # ── runtime ────────────────────────────────────────────────────────── # Inference backend. Five options exposed by /app/api/backends/factory.py: # official — default. Naive HF transformers autoregressive loop. # Reliable, but the GPU sits at ~27% utilization and # TTFB on a sentence is 8-12s (no streaming endpoint # tuning). 2026-04-27 bench confirmed. # optimized — torch.compile + CUDA graphs + real-time streaming. # Should be the path to upstream's claimed 97 ms TTFB. # CRASHES the container during its hardcoded warmup as # of 2026-04-27 (silent ExitCode 0, no Traceback, repeats # every ~22s under restart policy). TTS_WARMUP_ON_START= # false doesn't help — the optimized backend's warmup is # separate from the factory-level warmup. Don't use until # upstream root-cause is identified or worked around. # vllm_omni — vLLM-Omni-backed. Faster but needs more VRAM and a # vLLM dependency we haven't pulled. Untested. # pytorch — CPU-only fallback. # openvino — Intel CPU experimental. QWEN3_TTS_BACKEND=official # Model variant. Upstream publishes four checkpoints on HF: # Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice — flagship, supports voice cloning + preset voices (RECOMMENDED) # Qwen/Qwen3-TTS-12Hz-1.7B-Base — flagship, foundation only — REJECTS all synthesis through this wrapper # Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice — lightweight cloning-capable variant # Qwen/Qwen3-TTS-12Hz-0.6B-Base — lightweight foundation only # 1.7B = ~6–8 GB VRAM bfloat16, best quality. 0.6B = ~2–3 GB. # # IMPORTANT: the wrapper's only synthesis code path goes through the # voice cloning machinery. The -Base variants do NOT expose # generate_custom_voice; they error with "does not support # generate_custom_voice" on every request, no matter what voice you # specify (built-ins like Ryan, Vivian etc. count as custom voices). # Use -CustomVoice. The Base variant is foundation-model territory # that needs a different harness to be useful. # (This file had the labels backwards through 2026-04-27 — fixed # after the 2026-04-27 voice-fleet bench surfaced the failure mode.) QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice # Warm the model on container start so the first synthesis request # doesn't pay the load latency. Adds ~30 s to startup. Recommended. QWEN3_TTS_WARMUP=true # Concurrency cap on synthesis requests. Single GPU + 1.7B model → # leave at 1 unless you're load-testing. QWEN3_TTS_MAX_CONCURRENT=1 # Mount the gradio voice-studio UI at /voice-studio for browser-side # voice cloning. Set "false" to disable for headless deployments. QWEN3_TTS_VOICE_STUDIO=true # ── persistent storage on the host ─────────────────────────────────── # HuggingFace cache (model weights, ~5 GB after first run). Bind-mounted # so model state survives container recreate. Excluded from restic # (regenerable from HF Hub). QWEN3_TTS_CACHE_DIR=/worktank/qwen3-tts/cache # Cloned voice profiles (meta.json + reference.wav per voice). Precious # — cloned voices need the original reference audio to recreate. # Included in restic. QWEN3_TTS_VOICES_DIR=/worktank/qwen3-tts/voices