feat(chatterbox-fast): Phase 1 streaming server — adaptive-chunk scheduler
Build the streaming TTS server MVP per docs/design/chatterbox-fast-plan.md §4. - scheduler.py: adaptive buffer-ratchet chunker (the meat) — GPU-free pure logic. First sentence emitted alone for low TTFA, then chunks ratchet ~3x by packing whole sentences to margin x buffered-audio; drives off measured RTF + sec/char (EMA). relieve_leader() clause-splits a too-big mid-stream sentence to avoid starvation (joins land on commas); a long comma-less sentence is the one honored-but-flagged limitation. - test_scheduler.py: GPU-free simulation, 13 tests — asserts no-starvation (incl. overestimated RTF) and the ratchet. - app.py: FastAPI model holder + POST /tts StreamingResponse (raw PCM s16le default, wav optional, stream/oneshot) + GET /health. - bench.py: client — ground-truth TTFB + real 1x-consumer starvation check. Live test on irv-ml1 (turbo, A6000, GLaDOS voice): streaming TTFB 499ms vs oneshot 5230ms (~10x), stayed ahead of a 1x player (no starvation), ratchet 1.64->4.08->8.60->8.60s audio, measured RTF self-corrected 3.38->4.01. Kill the superseded docs/design/chatterbox-fast.md — its §5 windowed-token streaming was the abandoned native-frame-streaming arc; the adaptive-chunk plan supersedes it. Repoint persistent-memory + README at the canonical plan.
This commit is contained in:
@@ -0,0 +1,96 @@
|
||||
# chatterbox-fast — streaming TTS engine
|
||||
|
||||
Custom streaming server on top of `ChatterboxTurboTTS` that delivers
|
||||
**sub-second time-to-first-audio** while keeping turbo's full quality. Workload:
|
||||
**single-stream interactive**. Deployed (Phase 3) **alongside** the live
|
||||
`chatterbox` (:8196) on irv-ml1, burned in, then catalog-flipped.
|
||||
|
||||
Design: [`docs/design/chatterbox-fast-plan.md`](../../docs/design/chatterbox-fast-plan.md)
|
||||
(canonical plan). The abandoned native-frame-streaming arc is recorded in
|
||||
`persistent-memory.md` → *Tried and abandoned*.
|
||||
|
||||
## How it works — adaptive buffer-ratchet chunking
|
||||
|
||||
The engine never splits mid-sentence (keeps each chunk prosodically coherent).
|
||||
Instead it rides Chatterbox's faster-than-realtime generation (RTF ~3.4–3.8×):
|
||||
|
||||
1. **Chunk 1 = first sentence**, generated and emitted immediately (~0.66s
|
||||
first-audio). Latency-critical.
|
||||
2. **While chunk N plays, generate chunk N+1** by greedily accumulating whole
|
||||
sentences until the next would exceed `margin × audio_buffered_remaining`.
|
||||
3. Each chunk's playback buys wall-clock for a ~3× bigger next chunk, so after
|
||||
2-3 joins the rest of the paragraph is one big near-full-context chunk.
|
||||
Context loss is confined to those few sentence-boundary joins.
|
||||
4. Driven off **measured** RTF + sec-per-char (EMA), not constants.
|
||||
5. **Starvation relief:** if a mid-stream sentence is too long to generate
|
||||
within the current buffer, its *clause* boundaries are exposed so chunks pack
|
||||
to commas (natural pauses) — never a mid-clause split. A long *comma-less*
|
||||
sentence after a short opener is the one unavoidable case: the rule is
|
||||
honored and the brief gap is **flagged** (`drained > 0`), never hidden.
|
||||
|
||||
This only works because RTF > 1 — a sub-realtime model (e.g. Fish) would starve
|
||||
regardless of chunking. That is why this is the chatterbox-specific answer.
|
||||
|
||||
## Files
|
||||
|
||||
| file | role |
|
||||
|---|---|
|
||||
| `scheduler.py` | The adaptive-chunk scheduler. **GPU-free, pure logic** — the meat. |
|
||||
| `test_scheduler.py` | GPU-free simulation: asserts no-starvation + ratchet. `python test_scheduler.py` or `pytest`. |
|
||||
| `app.py` | FastAPI server: model holder + `POST /tts` (StreamingResponse) + `GET /health`. |
|
||||
| `bench.py` | Client: ground-truth TTFB + real 1×-consumer starvation check; saves `.wav` for A/B. |
|
||||
|
||||
Phase 3 will add `compose.yaml`, `Dockerfile`, `.env.example`.
|
||||
|
||||
## API
|
||||
|
||||
`POST /tts` → streamed audio. Body:
|
||||
|
||||
```json
|
||||
{ "text": "...", "voice": "glados_25s", "format": "pcm",
|
||||
"stream": true, "exaggeration": 0.5, "temperature": 0.8,
|
||||
"top_p": 0.95, "top_k": 1000, "repetition_penalty": 1.2 }
|
||||
```
|
||||
|
||||
- `format`: `pcm` (raw s16le @ 24 kHz, lowest latency, default) or `wav`.
|
||||
- `stream: false` → whole-text one-shot (the A/B quality baseline).
|
||||
- `voice`: predefined name (a `*.wav` in `CBF_VOICES_DIR`) or an absolute path
|
||||
to a clone reference. Omit → server default.
|
||||
- `margin` / `margin_first` / `rtf_prior`: optional scheduler overrides.
|
||||
|
||||
`GET /health` → `{status, sr, device, default_voice, voices_dir}`.
|
||||
|
||||
## Config (env)
|
||||
|
||||
| var | default | meaning |
|
||||
|---|---|---|
|
||||
| `CBF_MODEL_DEVICE` | `cuda` | `cuda` / `cuda:0` / `cpu` |
|
||||
| `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs |
|
||||
| `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) |
|
||||
| `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind |
|
||||
|
||||
## Dev / test on irv-ml1
|
||||
|
||||
```bash
|
||||
# (from this dir) copy the server into the chatterbox image and run it on GPU 1:
|
||||
scp app.py scheduler.py bench.py lkraven@10.100.79.3:/tmp/cbf/
|
||||
IMG=$(ssh lkraven@10.100.79.3 "docker images --format '{{.Repository}}:{{.Tag}}' | grep -i chatterbox | grep -v '<none>' | head -1")
|
||||
ssh lkraven@10.100.79.3 "docker run --rm --gpus '\"device=1\"' -e NVIDIA_VISIBLE_DEVICES=1 \
|
||||
-e HF_HOME=/app/hf_cache -e CBF_VOICES_DIR=/refs -e CBF_DEFAULT_VOICE=glados_25s \
|
||||
-p 8197:8197 -v /worktank/chatterbox/cache:/app/hf_cache \
|
||||
-v /worktank/chatterbox/reference_audio:/refs -v /tmp/cbf:/cbf \
|
||||
$IMG python /cbf/app.py"
|
||||
|
||||
# then, from the host (or anywhere on the WG net):
|
||||
python bench.py --host http://10.100.79.3:8197 --out /refs/_fast.wav
|
||||
python bench.py --host http://10.100.79.3:8197 --oneshot --out /refs/_oneshot.wav
|
||||
```
|
||||
|
||||
Pull the samples to listen: `scp lkraven@10.100.79.3:/worktank/chatterbox/reference_audio/_*.wav ~/chatterbox-ab/`.
|
||||
|
||||
## Acceptance (plan §6)
|
||||
|
||||
- **Latency:** first-audio < ~0.8s on the deployment GPU.
|
||||
- **No starvation:** `bench.py` reports "stayed ahead"; `test_scheduler.py` green.
|
||||
- **Quality:** operator ear-A/B the streamed output vs the one-shot — join-context
|
||||
loss should be ~imperceptible for multi-sentence text.
|
||||
Reference in New Issue
Block a user