refactor(dots-tts): extract TTS stack to tts-stack repo; pointer stub + move voices out
TTS development moves to a dedicated repo (~/development/tts-stack) so a separate agent can own tuning/dev. Mirrors the chatterbox-fast extraction: - stacks/dots-tts/ reduced to a pointer README (code/Dockerfile/compose/tests/env now canonical in tts-stack). - voices/ canonical corpus moved out to tts-stack/voices/. Blast-radius checked: no eshpfi playbook/script reads the corpus (other voices/ refs are unrelated host paths under /worktank/...). - persistent-memory updated: TTS dev extracted + stood down; reverses the earlier "corpus home = eshpfi voices/" call. The ~15 experimental TTS compose wrappers stay here as reference (catalogued in tts-stack/KNOWLEDGE.md). Live service on irv-ml1:8198 is unaffected (runs from a copy on the host).
This commit is contained in:
+21
-55
@@ -1,64 +1,30 @@
|
||||
# dots-tts
|
||||
# dots-tts — moved to its own repository
|
||||
|
||||
OpenAI-compatible zero-shot voice-clone TTS over **dots.tts** (rednote-hilab) —
|
||||
2B continuous-AR, native **48kHz**, `optimize=True` CUDA graphs → **RTF ~0.22** on
|
||||
the irv-ml1 3090. Thin FastAPI wrapper around `DotsTtsRuntime` (chosen over SGLang
|
||||
Omni: Omni's batching is MeanFlow-only and unneeded for a single consumer; the raw
|
||||
runtime already streams at the same RTF and is ~100 lines we control).
|
||||
The dots.tts serving stack now lives in the dedicated **tts-stack** repo:
|
||||
|
||||
- **Host:** irv-ml1, port **8198** (chatterbox-fast is :8197 — they co-reside on the 3090)
|
||||
- **Model:** `dots-studio/dots.tts-soar`, bf16, num_steps=10
|
||||
- **Voices:** every `<name>.wav` (+ `<name>.txt` transcript) in the mounted voices dir,
|
||||
sourced from the [`voices/`](../../voices/) canonical corpus via `derive.py dots`.
|
||||
> **`~/development/tts-stack`** → `stacks/dots-tts/`
|
||||
> (gitea `vh/tts-stack` on gitea.phasefinal.com once pushed)
|
||||
|
||||
## API
|
||||
Extracted from this workspace on 2026-08-11 so a dedicated agent can drive all TTS
|
||||
tuning / development. Like chatterbox-fast, dots-tts is **authored software with a
|
||||
test suite**, so it follows the sister-repo pattern rather than staying a thin
|
||||
compose wrapper here. The new repo owns the code, Dockerfile, tests, the voice
|
||||
corpus (formerly `voices/` here), the deploy runbook, and the running TTS
|
||||
knowledge base.
|
||||
|
||||
```
|
||||
GET /health -> {status, model, sample_rate, voices[]}
|
||||
GET /v1/voices -> {voices[]}
|
||||
POST /v1/audio/speech -> audio
|
||||
body: {input, voice, response_format?("wav"|"pcm"), stream?}
|
||||
```
|
||||
## Deployed service
|
||||
|
||||
`stream:true` returns a WAV stream (placeholder-header + PCM frames, 48kHz mono
|
||||
s16le) — the same shape the Zonos/chatterbox consumers already handle. Non-stream
|
||||
returns a complete WAV (or raw PCM with `response_format:"pcm"`).
|
||||
Live on **irv-ml1:8198** (`local/dots-tts:v3`, OpenAI `/v1/audio/speech`). The
|
||||
host stack dir is `/opt/docker/compose/dots-tts/`. Deploy + rollback runbook and
|
||||
all engine knowledge live in tts-stack (`docs/infrastructure.md`, `KNOWLEDGE.md`).
|
||||
|
||||
```bash
|
||||
curl -X POST http://10.100.79.3:8198/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Well, look who finally showed up.","voice":"glados"}' \
|
||||
--output out.wav
|
||||
```
|
||||
## Voice corpus
|
||||
|
||||
## Deploy
|
||||
The canonical voice corpus (was `eshpfi-management/voices/`) moved to
|
||||
`tts-stack/voices/`.
|
||||
|
||||
Reference sets come from the canonical corpus, not this stack — derive then point
|
||||
the mount at them:
|
||||
## Other TTS engines
|
||||
|
||||
```bash
|
||||
# 1. produce dots refs from the corpus (on a box with the whisper venv):
|
||||
python voices/derive.py dots # -> voices/derived/dots/*.{wav,txt}
|
||||
|
||||
# 2. build + run on irv-ml1 (cp .env.example .env first; adjust mounts):
|
||||
scripts/deploy-stack.sh irv-ml1 dots-tts # or, on the host:
|
||||
docker compose build && docker compose up -d
|
||||
```
|
||||
|
||||
The model (~5GB) is **not** baked — it's read from the mounted `HF_HOME`
|
||||
(`DOTS_HFCACHE_DIR`). First boot downloads it there if absent.
|
||||
|
||||
## Voice cloning gotcha
|
||||
|
||||
dots.tts clones from `(reference wav + its transcript)` and **leaks reference
|
||||
audio into the output** if the transcript is inaccurate or ends mid-clause. The
|
||||
`voices/` corpus + `derive.py` handle this (sentence-bounded trim + accurate
|
||||
transcript); don't hand this server a raw reference wav without a matching `.txt`.
|
||||
|
||||
## Notes
|
||||
|
||||
- **GPU:** `NVIDIA_VISIBLE_DEVICES=0` = the 3090 in Docker (PCI order). `optimize=True`
|
||||
is **incompatible with `PYTORCH_CUDA_ALLOC_CONF=expandable_segments`** (CUDA-graph
|
||||
capture error) — don't set it.
|
||||
- **Variants:** `dots.tts-mf` (MeanFlow, faster) is a drop-in via `DOTS_MODEL`; soar
|
||||
is the quality pick and single-consumer doesn't need mf's batching.
|
||||
The experimental TTS compose wrappers evaluated along the way (cosyvoice, dia,
|
||||
kokoro, vibevoice, zonos, …) remain under `stacks/` here as reference; their
|
||||
verdicts are catalogued in `tts-stack/KNOWLEDGE.md`.
|
||||
|
||||
Reference in New Issue
Block a user