stacks/qwen3-tts: target=production + user=root + correct HF model id

Three fixes from the first deploy attempt on irv-ml1:

- build.target=production. Upstream Dockerfile is multistage; the last
  stage `cpu-base` was selected by default, producing a CPU-only image
  with no flash-attn and `torch ... whl/cpu`.
- user: "0:0". Upstream image declares USER appuser but writes runtime
  state under /root (mode 0700). appuser cannot traverse /root, so
  /v1/voices 500s on PermissionError. Run as root to sidestep.
- QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice. The bare
  `1.7B` id we had isn't a real HF identifier; upstream publishes
  -CustomVoice / -Base variants of each size. Use -CustomVoice so
  `voice="clone:<name>"` works.

Tag bumped to v2 to keep the v1 cpu image distinguishable in the local
registry.

After: all 5 verify steps pass, GPU synthesis ~5s for 3-4s of audio,
three contrasting English `instructions` produce three distinct
hashes — emotion steering actually works (unlike CosyVoice's English
path).
This commit is contained in:
vh
2026-04-24 17:30:09 -07:00
parent 4f7bf3b0b6
commit 7c560a67fb
2 changed files with 25 additions and 6 deletions
+12 -6
View File
@@ -6,8 +6,10 @@
QWEN3_TTS_SHA=10323ce778c48a75dbda93d0a4891983fb371f58
# Local image tag — bump when you change build context to force a
# fresh layer build.
QWEN3_TTS_TAG=v1
# fresh layer build. v2 = first GPU build (target=production); v1
# was the accidental CPU-only image (last stage of upstream's
# multi-stage Dockerfile).
QWEN3_TTS_TAG=v2
# ── network ──────────────────────────────────────────────────────────
# Host port (container listens on 8880 internally).
@@ -23,10 +25,14 @@ QWEN3_TTS_BIND=0.0.0.0
# more VRAM); `pytorch` = bare pytorch path.
QWEN3_TTS_BACKEND=official
# Model variant. 1.7B = flagship, 6–8 GB VRAM with bfloat16, best
# quality + control. 0.6B = lightweight, ~2–3 GB VRAM, faster, slightly
# less expressive.
QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-1.7B
# Model variant. Upstream publishes four checkpoints on HF:
# Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice — flagship, voice cloning
# Qwen/Qwen3-TTS-12Hz-1.7B-Base — flagship, no cloning
# Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice — lightweight, voice cloning
# Qwen/Qwen3-TTS-12Hz-0.6B-Base — lightweight, no cloning
# 1.7B = ~6–8 GB VRAM bfloat16, best quality. 0.6B = ~2–3 GB.
# Use -CustomVoice for `voice="clone:<name>"` to work.
QWEN3_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
# Warm the model on container start so the first synthesis request
# doesn't pay the load latency. Adds ~30 s to startup. Recommended.