Files
vh db42a7cc17 qwen3-tts/.env.example: document the optimized-backend crashloop trap
Investigation of the slow (8-12s) qwen3-tts TTFB found the upstream
wrapper has 5 backend options. The advertised path to fast TTFB is
TTS_BACKEND=optimized (torch.compile + CUDA graphs + real-time
streaming). It loads cleanly but crashes the container during its
hardcoded warmup phase — silent exit (ExitCode 0, no traceback,
no OOM kill), repeats every ~22s under restart policy.

TTS_WARMUP_ON_START=false suppresses the factory-level warmup but
the optimized backend has its own internal warmup that fires
regardless and triggers the crash.

Updated the .env.example block to enumerate all 5 backend options
with their actual current behavior so future-Claude doesn't re-walk
this path. official is staying as the default.
2026-04-27 22:12:04 -07:00
..

Qwen3-TTS

Alibaba's open-weight TTS (Apache 2.0, released Jan 2026), deployed via the groxaxo/Qwen3-TTS-Openai-Fastapi OpenAI-compatible wrapper.

Server: irv-ml1 (Irvine, WireGuard-only) Port: 8191 (container 8880) GPUs: both exposed (NVIDIA_VISIBLE_DEVICES=all); 1.7B model fits on either the RTX 3090 (24 GB) or A6000 (48 GB) with headroom Image: local/qwen3-tts:v1 — built locally from a pinned git SHA of the wrapper repo via docker buildx's git URL context

Why this stack alongside cosyvoice

CosyVoice 3 (the other stack on this host) emits Chinese-flavored phonemes when given English content. Confirmed against upstream issue FunAudioLLM/CosyVoice#1790 — closed without a fix as of 2026-04-14.

Qwen3-TTS is from the same Alibaba team but built English-first into the checkpoint: 10 languages, 97 ms streaming TTFB, instruction-driven emotional expression, voice cloning. It's the better choice for any English-narration use; keep CosyVoice 3 around for Chinese / dialect work where it shines.

API endpoints

Method + path Purpose
POST /v1/audio/speech OpenAI-compatible TTS (drop-in for OpenAI clients)
GET /v1/voices List cloned voice profiles
GET /v1/models List available model checkpoints
GET /health Healthcheck
GET / Web UI
GET /voice-studio Gradio voice-cloning UI (when ENABLE_VOICE_STUDIO=true)

Voice cloning — different shape from cosyvoice

Profile-based, not voice-id-based:

  1. Clone a voice via the /voice-studio web UI (uploads reference audio + transcript, names the profile, e.g. glados).
  2. The wrapper writes /root/qwen3-tts/voice_library/profiles/<name>/{meta.json,reference.wav} inside the container (bind-mounted to /worktank/qwen3-tts/voices/profiles/<name>/ on the host).
  3. Reference the profile in synthesis requests as voice="clone:<name>".
# OpenAI-shape, English with emotion via instruction
curl -X POST http://10.100.79.3:8191/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen/Qwen3-TTS-12Hz-1.7B",
    "voice": "clone:glados",
    "input": "You should leave now. Visitors are not welcome.",
    "instructions": "speak with cold contempt",
    "response_format": "wav"
  }' \
  -o glados_en_angry.wav

(Note instructions field name; OpenAI standard. Whether English instructions actually drive emotion correctly here is the open question this stack exists to test — see the deploy playbook's verify section.)

Path layout

Host path Container path Purpose Restic?
/worktank/qwen3-tts/cache/ /root/.cache/huggingface Model cache (~5 GB after first run) excluded
/worktank/qwen3-tts/voices/ /root/qwen3-tts/voice_library Cloned profiles included (precious)

Deploy

Via elway — see playbooks/deploy-qwen3-tts.yaml in the eshpfi-management root. The playbook builds the image, ensures the host dirs, installs compose + .env, brings up. First boot downloads the 1.7B model from HF (~6 GB), which is the slowest step.

scripts/elway irv-ml1 --playbook playbooks/deploy-qwen3-tts.yaml

To pin a different upstream wrapper SHA, set --var sha=<new-sha> or edit .env on the server and rebuild.

Switching to the smaller model

If 1.7B is too heavy or you need to share GPUs more aggressively:

ssh irv-ml1 '
  cd /opt/docker/compose/qwen3-tts && \
  sed -i "s|^QWEN3_TTS_MODEL=.*|QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-0.6B-Base|" .env && \
  docker compose up -d
'

The new model auto-downloads on next start (~23 GB).