Deployed on irv-ml1 beside live chatterbox (:8196): healthy on :8197, TTFB ~0.5s, no starvation. Measured VRAM 5.34 GB (fp32) settles the placement: the 3090's ~3.8 GB free does NOT fit, A6000 (device 1) is the only viable card.
130 lines
6.3 KiB
Markdown
130 lines
6.3 KiB
Markdown
# chatterbox-fast — streaming TTS engine
|
||
|
||
Custom streaming server on top of `ChatterboxTurboTTS` that delivers
|
||
**sub-second time-to-first-audio** while keeping turbo's full quality. Workload:
|
||
**single-stream interactive**. Deployed (Phase 3) **alongside** the live
|
||
`chatterbox` (:8196) on irv-ml1, burned in, then catalog-flipped.
|
||
|
||
Design: [`docs/design/chatterbox-fast-plan.md`](../../docs/design/chatterbox-fast-plan.md)
|
||
(canonical plan). The abandoned native-frame-streaming arc is recorded in
|
||
`persistent-memory.md` → *Tried and abandoned*.
|
||
|
||
## How it works — adaptive buffer-ratchet chunking
|
||
|
||
The engine never splits mid-sentence (keeps each chunk prosodically coherent).
|
||
Instead it rides Chatterbox's faster-than-realtime generation (RTF ~3.4–3.8×):
|
||
|
||
1. **Chunk 1 = first sentence**, generated and emitted immediately (~0.66s
|
||
first-audio). Latency-critical.
|
||
2. **While chunk N plays, generate chunk N+1** by greedily accumulating whole
|
||
sentences until the next would exceed `margin × audio_buffered_remaining`.
|
||
3. Each chunk's playback buys wall-clock for a ~3× bigger next chunk, so after
|
||
2-3 joins the rest of the paragraph is one big near-full-context chunk.
|
||
Context loss is confined to those few sentence-boundary joins.
|
||
4. Driven off **measured** RTF + sec-per-char (EMA), not constants.
|
||
5. **Starvation relief:** if a mid-stream sentence is too long to generate
|
||
within the current buffer, its *clause* boundaries are exposed so chunks pack
|
||
to commas (natural pauses) — never a mid-clause split. A long *comma-less*
|
||
sentence after a short opener is the one unavoidable case: the rule is
|
||
honored and the brief gap is **flagged** (`drained > 0`), never hidden.
|
||
|
||
This only works because RTF > 1 — a sub-realtime model (e.g. Fish) would starve
|
||
regardless of chunking. That is why this is the chatterbox-specific answer.
|
||
|
||
## Files
|
||
|
||
| file | role |
|
||
|---|---|
|
||
| `scheduler.py` | The adaptive-chunk scheduler. **GPU-free, pure logic** — the meat. |
|
||
| `test_scheduler.py` | GPU-free simulation: asserts no-starvation + ratchet. `python test_scheduler.py` or `pytest`. |
|
||
| `app.py` | FastAPI server: model holder + `POST /tts` (StreamingResponse) + `GET /health`. |
|
||
| `bench.py` | Client: ground-truth TTFB + real 1×-consumer starvation check; saves `.wav` for A/B. |
|
||
| `Dockerfile` | Thin overlay: `FROM local/chatterbox:v1` + our two modules. |
|
||
| `compose.yaml` · `.env.example` | Deploy on irv-ml1 alongside the live `chatterbox`. |
|
||
|
||
## Deploy (Phase 3)
|
||
|
||
```bash
|
||
scripts/deploy-stack.sh irv-ml1 chatterbox-fast # push compose+code to the host
|
||
# then on irv-ml1, in /opt/docker/compose/chatterbox-fast/ (after copying .env):
|
||
docker compose build && docker compose up -d
|
||
```
|
||
|
||
`Dockerfile` is `FROM local/chatterbox:v1` (the sibling stack's image — must exist
|
||
on irv-ml1) + `COPY scheduler.py app.py`. GPU pin and voices/cache paths come from
|
||
`.env` (see `.env.example`). GPU is **device 1 (A6000)** — measured footprint is
|
||
**5.34 GB** (turbo loads fp32), so the 3090's ~3.8 GB free does **not** fit it.
|
||
|
||
**Deployed 2026-06-02** alongside the live `chatterbox` (:8196): healthy on
|
||
:8197, TTFB ~0.5s, no starvation, ~7 GB free left on the A6000.
|
||
|
||
## API
|
||
|
||
`POST /tts` → streamed audio. Body:
|
||
|
||
```json
|
||
{ "text": "...", "voice": "glados_25s", "format": "pcm",
|
||
"stream": true, "exaggeration": 0.5, "temperature": 0.8,
|
||
"top_p": 0.95, "top_k": 1000, "repetition_penalty": 1.2 }
|
||
```
|
||
|
||
- `format`: `pcm` (raw s16le @ 24 kHz, lowest latency, default) or `wav`.
|
||
- `stream: false` → whole-text one-shot (the A/B quality baseline).
|
||
- `voice`: predefined name (a `*.wav` in `CBF_VOICES_DIR`) or an absolute path
|
||
to a clone reference. Omit → server default.
|
||
- `margin` / `margin_first` / `rtf_prior`: optional scheduler overrides.
|
||
|
||
`GET /health` → `{status, sr, device, default_voice, voices_dir}`.
|
||
|
||
`GET /voices` → `{voices: [stem…], default}` — predefined `*.wav` stems in
|
||
`CBF_VOICES_DIR` (`_`-prefixed scratch/A-B files excluded). Clone refs are passed
|
||
per-request as an absolute path and aren't listed.
|
||
|
||
## Config (env)
|
||
|
||
| var | default | meaning |
|
||
|---|---|---|
|
||
| `CBF_MODEL_DEVICE` | `cuda` | `cuda` / `cuda:0` / `cpu` |
|
||
| `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs |
|
||
| `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) |
|
||
| `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind |
|
||
| `CBF_TF32` | `1` | TF32 matmul/cudnn (free; off with `0`) |
|
||
| `CBF_SDPA_FLASH` | `1` | flash + mem-efficient SDPA backend |
|
||
|
||
### Perf notes (measured 2026-06-02, turbo on A6000)
|
||
|
||
- Model loads in **float32** (not the fp16 older notes assumed).
|
||
- **TF32 + SDPA do not move TTFA** (~0.5s): the first-sentence latency is bound by
|
||
the sequential AR token decode (T3 Llama, batch-1), not matmul throughput. They
|
||
stay on (free, help the larger chunks marginally).
|
||
- **bf16 deferred:** the lever that *would* help batch-1 decode, but `from_pretrained()`
|
||
has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make
|
||
a clean cast nontrivial. Not worth the quality risk while ~0.5s TTFA is fine.
|
||
- **torch.compile: deferred** (research flags a batch-1 regression).
|
||
|
||
## Dev / test on irv-ml1
|
||
|
||
```bash
|
||
# (from this dir) copy the server into the chatterbox image and run it on GPU 1:
|
||
scp app.py scheduler.py bench.py lkraven@10.100.79.3:/tmp/cbf/
|
||
IMG=$(ssh lkraven@10.100.79.3 "docker images --format '{{.Repository}}:{{.Tag}}' | grep -i chatterbox | grep -v '<none>' | head -1")
|
||
ssh lkraven@10.100.79.3 "docker run --rm --gpus '\"device=1\"' -e NVIDIA_VISIBLE_DEVICES=1 \
|
||
-e HF_HOME=/app/hf_cache -e CBF_VOICES_DIR=/refs -e CBF_DEFAULT_VOICE=glados_25s \
|
||
-p 8197:8197 -v /worktank/chatterbox/cache:/app/hf_cache \
|
||
-v /worktank/chatterbox/reference_audio:/refs -v /tmp/cbf:/cbf \
|
||
$IMG python /cbf/app.py"
|
||
|
||
# then, from the host (or anywhere on the WG net):
|
||
python bench.py --host http://10.100.79.3:8197 --out /refs/_fast.wav
|
||
python bench.py --host http://10.100.79.3:8197 --oneshot --out /refs/_oneshot.wav
|
||
```
|
||
|
||
Pull the samples to listen: `scp lkraven@10.100.79.3:/worktank/chatterbox/reference_audio/_*.wav ~/chatterbox-ab/`.
|
||
|
||
## Acceptance (plan §6)
|
||
|
||
- **Latency:** first-audio < ~0.8s on the deployment GPU.
|
||
- **No starvation:** `bench.py` reports "stayed ahead"; `test_scheduler.py` green.
|
||
- **Quality:** operator ear-A/B the streamed output vs the one-shot — join-context
|
||
loss should be ~imperceptible for multi-sentence text.
|