feat(dots-tts): ship OpenAI-compatible dots.tts TTS stack on irv-ml1:8198
Thin FastAPI wrapper over DotsTtsRuntime (soar, optimize=True, RTF ~0.22), serialized single-consumer; OpenAI /v1/audio/speech (stream + non-stream), voices from the voices/ corpus derived set. Live + healthy alongside chatterbox-fast on the 3090; nothing repointed. Dockerfile needs build-essential (torch.compile/inductor JITs via gcc at runtime) + persisted inductor cache. Remaining Phase-2: ratatoskr client cutover.
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
# dots-tts
|
||||
|
||||
OpenAI-compatible zero-shot voice-clone TTS over **dots.tts** (rednote-hilab) —
|
||||
2B continuous-AR, native **48kHz**, `optimize=True` CUDA graphs → **RTF ~0.22** on
|
||||
the irv-ml1 3090. Thin FastAPI wrapper around `DotsTtsRuntime` (chosen over SGLang
|
||||
Omni: Omni's batching is MeanFlow-only and unneeded for a single consumer; the raw
|
||||
runtime already streams at the same RTF and is ~100 lines we control).
|
||||
|
||||
- **Host:** irv-ml1, port **8198** (chatterbox-fast is :8197 — they co-reside on the 3090)
|
||||
- **Model:** `dots-studio/dots.tts-soar`, bf16, num_steps=10
|
||||
- **Voices:** every `<name>.wav` (+ `<name>.txt` transcript) in the mounted voices dir,
|
||||
sourced from the [`voices/`](../../voices/) canonical corpus via `derive.py dots`.
|
||||
|
||||
## API
|
||||
|
||||
```
|
||||
GET /health -> {status, model, sample_rate, voices[]}
|
||||
GET /v1/voices -> {voices[]}
|
||||
POST /v1/audio/speech -> audio
|
||||
body: {input, voice, response_format?("wav"|"pcm"), stream?}
|
||||
```
|
||||
|
||||
`stream:true` returns a WAV stream (placeholder-header + PCM frames, 48kHz mono
|
||||
s16le) — the same shape the Zonos/chatterbox consumers already handle. Non-stream
|
||||
returns a complete WAV (or raw PCM with `response_format:"pcm"`).
|
||||
|
||||
```bash
|
||||
curl -X POST http://10.100.79.3:8198/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Well, look who finally showed up.","voice":"glados"}' \
|
||||
--output out.wav
|
||||
```
|
||||
|
||||
## Deploy
|
||||
|
||||
Reference sets come from the canonical corpus, not this stack — derive then point
|
||||
the mount at them:
|
||||
|
||||
```bash
|
||||
# 1. produce dots refs from the corpus (on a box with the whisper venv):
|
||||
python voices/derive.py dots # -> voices/derived/dots/*.{wav,txt}
|
||||
|
||||
# 2. build + run on irv-ml1 (cp .env.example .env first; adjust mounts):
|
||||
scripts/deploy-stack.sh irv-ml1 dots-tts # or, on the host:
|
||||
docker compose build && docker compose up -d
|
||||
```
|
||||
|
||||
The model (~5GB) is **not** baked — it's read from the mounted `HF_HOME`
|
||||
(`DOTS_HFCACHE_DIR`). First boot downloads it there if absent.
|
||||
|
||||
## Voice cloning gotcha
|
||||
|
||||
dots.tts clones from `(reference wav + its transcript)` and **leaks reference
|
||||
audio into the output** if the transcript is inaccurate or ends mid-clause. The
|
||||
`voices/` corpus + `derive.py` handle this (sentence-bounded trim + accurate
|
||||
transcript); don't hand this server a raw reference wav without a matching `.txt`.
|
||||
|
||||
## Notes
|
||||
|
||||
- **GPU:** `NVIDIA_VISIBLE_DEVICES=0` = the 3090 in Docker (PCI order). `optimize=True`
|
||||
is **incompatible with `PYTORCH_CUDA_ALLOC_CONF=expandable_segments`** (CUDA-graph
|
||||
capture error) — don't set it.
|
||||
- **Variants:** `dots.tts-mf` (MeanFlow, faster) is a drop-in via `DOTS_MODEL`; soar
|
||||
is the quality pick and single-consumer doesn't need mf's batching.
|
||||
Reference in New Issue
Block a user