stacks: add Kokoro, VibeVoice 1.5B, Chatterbox Turbo (TTS slate fill-in)

Three TTS additions to round out coverage on irv-ml1, each filling a
distinct niche the existing slate doesn't own.

Final coverage matrix (all on irv-ml1):
  Kokoro              — low-latency English, fixed voice library, ~300ms TTFA
  Chatterbox Turbo    — low-latency English w/ voice cloning + paralinguistic tags
  IndexTTS-2          — English voice cloning + emotion vector / text control
  Qwen3-TTS-1.7B-Base — high-quality English voice cloning
  CosyVoice 3         — multilingual (Chinese-leaning)
  VibeVoice 1.5B      — long-form / multi-speaker dialogue

stacks/kokoro:
  - port 8193, GPU device 0 (3090)
  - pulls ghcr.io/remsky/kokoro-fastapi-gpu:v0.2.4-master (no Dockerfile,
    no first-run model download — models baked in)
  - 60+ built-in voices, OpenAI-compat with stream=true over chunked HTTP
  - Apache-2.0 weights + code, ~1 GB VRAM

stacks/vibevoice:
  - port 8194, GPU device 1 (A6000 — for 7B headroom)
  - builds groxaxo/VibeVoice-FastAPI1 (more current fork of ncoder-ai)
    pinned to 7614c469a145
  - default model microsoft/VibeVoice-1.5B (~7 GB bf16 VRAM); env var
    swap to rsxdalv/VibeVoice-Large (7B) or FabioSarracino/VibeVoice-Large-Q8
  - multi-speaker dialogue via /v1/vibevoice/generate with Speaker N: format
  - long-form niche only — not low-latency

stacks/chatterbox:
  - port 8196, GPU device 0 (3090)
  - builds devnen/Chatterbox-TTS-Server (most active Turbo-supporting wrapper)
  - default model ResembleAI/chatterbox-turbo (~2.5 GB fp16, ~75ms latency)
  - paralinguistic tags inline ([laugh] [whisper] etc) — different shape
    from IndexTTS-2's emotion vector; fills the speed+cloning niche
    Kokoro/IndexTTS don't cover together
  - mandatory PerTh watermark on outputs (Resemble policy)

Three matching playbooks under playbooks/deploy-{kokoro,vibevoice,
chatterbox}.yaml. All idempotent, creates-/when-gated.

Cold-deploy disk on /worktank/: ~7 GB Kokoro + ~19 GB VibeVoice 1.5B
+ ~12 GB Chatterbox = ~38 GB total. VRAM concurrent: ~10-11 GB across
both GPUs.

Skipped from the original four-stack proposal: VibeVoice Realtime
(overlaps Kokoro's niche; Kokoro wins on latency, license, and not
needing a build).
This commit is contained in:
vh
2026-04-25 16:18:37 -07:00
parent 54fef0e9d8
commit 4549d241a7
12 changed files with 930 additions and 0 deletions
+86
View File
@@ -0,0 +1,86 @@
# Deploy Kokoro-FastAPI to irv-ml1.
#
# Image is published on GHCR — no Dockerfile to maintain, no first-run
# model download (Kokoro-82M weights are baked in). Stage compose +
# .env, pull, bring up. ~6.5 GB pull on cold cache, ~2-5 min.
#
# Usage:
# scripts/elway irv-ml1 --playbook playbooks/deploy-kokoro.yaml
#
# Idempotent — every step is creates-/when-gated; rerun is safe.
vars:
compose_dir: /opt/docker/compose/kokoro
user_voices_dir: /worktank/kokoro/user_voices
voices_dir: /worktank/kokoro/voices
host_port: "8193"
steps:
# ── host-side dirs ──────────────────────────────────────────────────
- name: Ensure /worktank/kokoro root exists (one-time, sudo)
shell: mkdir -p /worktank/kokoro
sudo: true
creates: /worktank/kokoro
- name: Chown /worktank/kokoro to lkraven
shell: chown lkraven:lkraven /worktank/kokoro
sudo: true
when: '[ "$(stat -c %U /worktank/kokoro)" != lkraven ]'
- name: Ensure user-voices dir exists
shell: mkdir -p {{ user_voices_dir }}
creates: "{{ user_voices_dir }}"
- name: Ensure voices dir exists (used only if compose mount is enabled)
shell: mkdir -p {{ voices_dir }}
creates: "{{ voices_dir }}"
- name: Ensure compose dir exists
shell: mkdir -p {{ compose_dir }}
creates: "{{ compose_dir }}"
# ── deploy compose files ────────────────────────────────────────────
- name: Upload compose.yaml
upload:
src: stacks/kokoro/compose.yaml
dest: "{{ compose_dir }}/compose.yaml"
mode: "0644"
- name: Seed .env from template (only if absent)
upload:
src: stacks/kokoro/.env.example
dest: "{{ compose_dir }}/.env"
mode: "0644"
when: "[ ! -f {{ compose_dir }}/.env ]"
# ── pull + bring up ─────────────────────────────────────────────────
- name: docker compose pull (first run: ~6.5 GB from GHCR)
shell: cd {{ compose_dir }} && docker compose pull
- name: docker compose up -d
shell: cd {{ compose_dir }} && docker compose up -d
- name: Wait for /v1/audio/voices to respond
shell: |
for i in $(seq 1 60); do
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/v1/audio/voices && exit 0
sleep 5
done
exit 1
changed_when: "false"
verify:
- name: /v1/audio/voices returns 200
shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/v1/audio/voices
changed_when: "false"
- name: /v1/audio/voices includes at least one built-in (af_bella)
shell: curl -sf http://localhost:{{ host_port }}/v1/audio/voices | grep -q af_bella
changed_when: "false"
- name: Container is running
shell: docker inspect kokoro --format '{{.State.Status}}' | grep -q running
changed_when: "false"