feat(dots-tts): ship OpenAI-compatible dots.tts TTS stack on irv-ml1:8198

Thin FastAPI wrapper over DotsTtsRuntime (soar, optimize=True, RTF ~0.22),
serialized single-consumer; OpenAI /v1/audio/speech (stream + non-stream),
voices from the voices/ corpus derived set. Live + healthy alongside
chatterbox-fast on the 3090; nothing repointed. Dockerfile needs
build-essential (torch.compile/inductor JITs via gcc at runtime) + persisted
inductor cache. Remaining Phase-2: ratatoskr client cutover.
This commit is contained in:
vh
2026-08-10 01:07:37 -07:00
parent fca1a545f1
commit c8acf60449
6 changed files with 325 additions and 1 deletions
+64
View File
@@ -0,0 +1,64 @@
# dots-tts
OpenAI-compatible zero-shot voice-clone TTS over **dots.tts** (rednote-hilab) —
2B continuous-AR, native **48kHz**, `optimize=True` CUDA graphs → **RTF ~0.22** on
the irv-ml1 3090. Thin FastAPI wrapper around `DotsTtsRuntime` (chosen over SGLang
Omni: Omni's batching is MeanFlow-only and unneeded for a single consumer; the raw
runtime already streams at the same RTF and is ~100 lines we control).
- **Host:** irv-ml1, port **8198** (chatterbox-fast is :8197 — they co-reside on the 3090)
- **Model:** `dots-studio/dots.tts-soar`, bf16, num_steps=10
- **Voices:** every `<name>.wav` (+ `<name>.txt` transcript) in the mounted voices dir,
sourced from the [`voices/`](../../voices/) canonical corpus via `derive.py dots`.
## API
```
GET /health -> {status, model, sample_rate, voices[]}
GET /v1/voices -> {voices[]}
POST /v1/audio/speech -> audio
body: {input, voice, response_format?("wav"|"pcm"), stream?}
```
`stream:true` returns a WAV stream (placeholder-header + PCM frames, 48kHz mono
s16le) — the same shape the Zonos/chatterbox consumers already handle. Non-stream
returns a complete WAV (or raw PCM with `response_format:"pcm"`).
```bash
curl -X POST http://10.100.79.3:8198/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Well, look who finally showed up.","voice":"glados"}' \
--output out.wav
```
## Deploy
Reference sets come from the canonical corpus, not this stack — derive then point
the mount at them:
```bash
# 1. produce dots refs from the corpus (on a box with the whisper venv):
python voices/derive.py dots # -> voices/derived/dots/*.{wav,txt}
# 2. build + run on irv-ml1 (cp .env.example .env first; adjust mounts):
scripts/deploy-stack.sh irv-ml1 dots-tts # or, on the host:
docker compose build && docker compose up -d
```
The model (~5GB) is **not** baked — it's read from the mounted `HF_HOME`
(`DOTS_HFCACHE_DIR`). First boot downloads it there if absent.
## Voice cloning gotcha
dots.tts clones from `(reference wav + its transcript)` and **leaks reference
audio into the output** if the transcript is inaccurate or ends mid-clause. The
`voices/` corpus + `derive.py` handle this (sentence-bounded trim + accurate
transcript); don't hand this server a raw reference wav without a matching `.txt`.
## Notes
- **GPU:** `NVIDIA_VISIBLE_DEVICES=0` = the 3090 in Docker (PCI order). `optimize=True`
is **incompatible with `PYTORCH_CUDA_ALLOC_CONF=expandable_segments`** (CUDA-graph
capture error) — don't set it.
- **Variants:** `dots.tts-mf` (MeanFlow, faster) is a drop-in via `DOTS_MODEL`; soar
is the quality pick and single-consumer doesn't need mf's batching.