refactor(dots-tts): extract TTS stack to tts-stack repo; pointer stub + move voices out

TTS development moves to a dedicated repo (~/development/tts-stack) so a separate
agent can own tuning/dev. Mirrors the chatterbox-fast extraction:

- stacks/dots-tts/ reduced to a pointer README (code/Dockerfile/compose/tests/env
  now canonical in tts-stack).
- voices/ canonical corpus moved out to tts-stack/voices/. Blast-radius checked:
  no eshpfi playbook/script reads the corpus (other voices/ refs are unrelated
  host paths under /worktank/...).
- persistent-memory updated: TTS dev extracted + stood down; reverses the earlier
  "corpus home = eshpfi voices/" call.

The ~15 experimental TTS compose wrappers stay here as reference (catalogued in
tts-stack/KNOWLEDGE.md). Live service on irv-ml1:8198 is unaffected (runs from a
copy on the host).
This commit is contained in:
vh
2026-08-11 07:46:11 -07:00
parent a80f6e958f
commit 62672c9850
20 changed files with 24 additions and 746 deletions
+21 -55
View File
@@ -1,64 +1,30 @@
# dots-tts
# dots-tts — moved to its own repository
OpenAI-compatible zero-shot voice-clone TTS over **dots.tts** (rednote-hilab) —
2B continuous-AR, native **48kHz**, `optimize=True` CUDA graphs → **RTF ~0.22** on
the irv-ml1 3090. Thin FastAPI wrapper around `DotsTtsRuntime` (chosen over SGLang
Omni: Omni's batching is MeanFlow-only and unneeded for a single consumer; the raw
runtime already streams at the same RTF and is ~100 lines we control).
The dots.tts serving stack now lives in the dedicated **tts-stack** repo:
- **Host:** irv-ml1, port **8198** (chatterbox-fast is :8197 — they co-reside on the 3090)
- **Model:** `dots-studio/dots.tts-soar`, bf16, num_steps=10
- **Voices:** every `<name>.wav` (+ `<name>.txt` transcript) in the mounted voices dir,
sourced from the [`voices/`](../../voices/) canonical corpus via `derive.py dots`.
> **`~/development/tts-stack`** → `stacks/dots-tts/`
> (gitea `vh/tts-stack` on gitea.phasefinal.com once pushed)
## API
Extracted from this workspace on 2026-08-11 so a dedicated agent can drive all TTS
tuning / development. Like chatterbox-fast, dots-tts is **authored software with a
test suite**, so it follows the sister-repo pattern rather than staying a thin
compose wrapper here. The new repo owns the code, Dockerfile, tests, the voice
corpus (formerly `voices/` here), the deploy runbook, and the running TTS
knowledge base.
```
GET /health -> {status, model, sample_rate, voices[]}
GET /v1/voices -> {voices[]}
POST /v1/audio/speech -> audio
body: {input, voice, response_format?("wav"|"pcm"), stream?}
```
## Deployed service
`stream:true` returns a WAV stream (placeholder-header + PCM frames, 48kHz mono
s16le) — the same shape the Zonos/chatterbox consumers already handle. Non-stream
returns a complete WAV (or raw PCM with `response_format:"pcm"`).
Live on **irv-ml1:8198** (`local/dots-tts:v3`, OpenAI `/v1/audio/speech`). The
host stack dir is `/opt/docker/compose/dots-tts/`. Deploy + rollback runbook and
all engine knowledge live in tts-stack (`docs/infrastructure.md`, `KNOWLEDGE.md`).
```bash
curl -X POST http://10.100.79.3:8198/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Well, look who finally showed up.","voice":"glados"}' \
--output out.wav
```
## Voice corpus
## Deploy
The canonical voice corpus (was `eshpfi-management/voices/`) moved to
`tts-stack/voices/`.
Reference sets come from the canonical corpus, not this stack — derive then point
the mount at them:
## Other TTS engines
```bash
# 1. produce dots refs from the corpus (on a box with the whisper venv):
python voices/derive.py dots # -> voices/derived/dots/*.{wav,txt}
# 2. build + run on irv-ml1 (cp .env.example .env first; adjust mounts):
scripts/deploy-stack.sh irv-ml1 dots-tts # or, on the host:
docker compose build && docker compose up -d
```
The model (~5GB) is **not** baked — it's read from the mounted `HF_HOME`
(`DOTS_HFCACHE_DIR`). First boot downloads it there if absent.
## Voice cloning gotcha
dots.tts clones from `(reference wav + its transcript)` and **leaks reference
audio into the output** if the transcript is inaccurate or ends mid-clause. The
`voices/` corpus + `derive.py` handle this (sentence-bounded trim + accurate
transcript); don't hand this server a raw reference wav without a matching `.txt`.
## Notes
- **GPU:** `NVIDIA_VISIBLE_DEVICES=0` = the 3090 in Docker (PCI order). `optimize=True`
is **incompatible with `PYTORCH_CUDA_ALLOC_CONF=expandable_segments`** (CUDA-graph
capture error) — don't set it.
- **Variants:** `dots.tts-mf` (MeanFlow, faster) is a drop-in via `DOTS_MODEL`; soar
is the quality pick and single-consumer doesn't need mf's batching.
The experimental TTS compose wrappers evaluated along the way (cosyvoice, dia,
kokoro, vibevoice, zonos, …) remain under `stacks/` here as reference; their
verdicts are catalogued in `tts-stack/KNOWLEDGE.md`.