feat(chatterbox-fast): Phase 3 scaffold — Dockerfile, compose, .env.example

Container artifacts to deploy alongside the live chatterbox (:8196) on irv-ml1.
- Dockerfile: thin overlay FROM local/chatterbox:v1 (sibling's image, has the
  chatterbox lib + torch + fastapi) + COPY scheduler.py app.py; runs uvicorn.
- compose.yaml: mirrors the sibling chatterbox stack (runtime: nvidia +
  NVIDIA_VISIBLE_DEVICES; host IP:port, no traefik-net — these GPU TTS services
  aren't traefik-fronted). Port 8197, /health healthcheck, homepage labels,
  reuses /worktank/chatterbox/{cache,reference_audio}.
- .env.example: GPU default device 1 (A6000) — turbo is fp32, 3090 free VRAM is
  tight; port reservations; perf-lever toggles.

Not yet deployed — awaiting operator go (shared GPU host, runs beside production).
This commit is contained in:
vh
2026-06-01 23:30:36 -07:00
parent a95aa75947
commit 5c8d174f8e
4 changed files with 135 additions and 1 deletions
+47
View File
@@ -0,0 +1,47 @@
# chatterbox-fast stack tunables. Copy to `.env` on irv-ml1 before deploying.
# ── image / build ────────────────────────────────────────────────────
# Local image tag for this stack. Bump to force a fresh layer build.
CBF_TAG=v1
# Base image tag — the sibling `chatterbox` stack's local image, which
# carries the chatterbox lib + torch + fastapi. Must exist on irv-ml1
# (built by the `chatterbox` stack). Bump in lockstep if that rebuilds.
CBF_BASE_TAG=v1
# ── network ──────────────────────────────────────────────────────────
# Host port. Container listens on 8197 internally.
# Reserved on irv-ml1: 8188 ComfyUI, 8190 CosyVoice, 8191 Qwen3-TTS,
# 8192 IndexTTS-2, 8193 Kokoro, 8194 VibeVoice, 8196 Chatterbox,
# 8765 Parakeet. 8197 picked here.
CBF_PORT=8197
# Bind address. 0.0.0.0 exposes on all interfaces (incl. the WG tunnel
# interface 10.100.79.3); 127.0.0.1 restricts to local-only.
CBF_BIND=0.0.0.0
# ── runtime / GPU ────────────────────────────────────────────────────
# Device visible inside the container.
# 1 = RTX A6000 (~12 GB free; the safe default).
# 0 = RTX 3090 — TIGHT: turbo loads FP32 (not fp16), and the 3090 idles
# ~20.5 GB used (shared dev stack), leaving ~3.8 GB free. Measure the
# actual footprint before pinning here; it likely will NOT fit in fp32.
CBF_GPU_DEVICES=1
# Default reference voice (a *.wav stem in CBF_REFERENCE_DIR, or an absolute
# path). The server prepares this at startup so first request is warm.
CBF_DEFAULT_VOICE=glados_25s
# Perf levers (Ampere-safe, free). Off with 0. Measured: they don't move TTFA
# (AR-decode-bound) but don't hurt; bf16 is deferred (fp32 model, no clean cast).
CBF_TF32=1
CBF_SDPA_FLASH=1
# ── persistent storage on the host ───────────────────────────────────
# Reference / predefined voice wavs (mounted at /refs). Shared with the
# `chatterbox` stack. `_`-prefixed files (bench/A-B scratch) are ignored.
CBF_REFERENCE_DIR=/worktank/chatterbox/reference_audio
# HuggingFace cache — Chatterbox-Turbo weights. Reused from the `chatterbox`
# stack (already populated, ~3.8 GB); no re-download. Excluded from restic.
CBF_CACHE_DIR=/worktank/chatterbox/cache