# chatterbox-fast — streaming TTS engine Custom streaming server on top of `ChatterboxTurboTTS` that delivers **sub-second time-to-first-audio** while keeping turbo's full quality. Workload: **single-stream interactive**. Deployed (Phase 3) **alongside** the live `chatterbox` (:8196) on irv-ml1, burned in, then catalog-flipped. Design: [`docs/design/chatterbox-fast-plan.md`](../../docs/design/chatterbox-fast-plan.md) (canonical plan). The abandoned native-frame-streaming arc is recorded in `persistent-memory.md` → *Tried and abandoned*. ## How it works — adaptive buffer-ratchet chunking The engine never splits mid-sentence (keeps each chunk prosodically coherent). Instead it rides Chatterbox's faster-than-realtime generation (RTF ~3.4–3.8×): 1. **Chunk 1 = first sentence**, generated and emitted immediately (~0.66s first-audio). Latency-critical. 2. **While chunk N plays, generate chunk N+1** by greedily accumulating whole sentences until the next would exceed `margin × audio_buffered_remaining`. 3. Each chunk's playback buys wall-clock for a ~3× bigger next chunk, so after 2-3 joins the rest of the paragraph is one big near-full-context chunk. Context loss is confined to those few sentence-boundary joins. 4. Driven off **measured** RTF + sec-per-char (EMA), not constants. 5. **Starvation relief:** if a mid-stream sentence is too long to generate within the current buffer, its *clause* boundaries are exposed so chunks pack to commas (natural pauses) — never a mid-clause split. A long *comma-less* sentence after a short opener is the one unavoidable case: the rule is honored and the brief gap is **flagged** (`drained > 0`), never hidden. This only works because RTF > 1 — a sub-realtime model (e.g. Fish) would starve regardless of chunking. That is why this is the chatterbox-specific answer. ## Files | file | role | |---|---| | `scheduler.py` | The adaptive-chunk scheduler. **GPU-free, pure logic** — the meat. | | `test_scheduler.py` | GPU-free simulation: asserts no-starvation + ratchet. `python test_scheduler.py` or `pytest`. | | `app.py` | FastAPI server: model holder + `POST /tts` (StreamingResponse) + `GET /health`. | | `bench.py` | Client: ground-truth TTFB + real 1×-consumer starvation check; saves `.wav` for A/B. | | `Dockerfile` | Thin overlay: `FROM local/chatterbox:v1` + our two modules. | | `compose.yaml` · `.env.example` | Deploy on irv-ml1 alongside the live `chatterbox`. | ## Deploy (Phase 3) ```bash scripts/deploy-stack.sh irv-ml1 chatterbox-fast # push compose+code to the host # then on irv-ml1, in /opt/docker/compose/chatterbox-fast/ (after copying .env): docker compose build && docker compose up -d ``` `Dockerfile` is `FROM local/chatterbox:v1` (the sibling stack's image — must exist on irv-ml1) + `COPY scheduler.py app.py`. GPU pin and voices/cache paths come from `.env` (see `.env.example`). GPU is **device 1 (A6000)** — measured footprint is **5.34 GB** (turbo loads fp32), so the 3090's ~3.8 GB free does **not** fit it. **Deployed 2026-06-02** alongside the live `chatterbox` (:8196): healthy on :8197, TTFB ~0.5s, no starvation, ~7 GB free left on the A6000. ## API `POST /tts` → streamed audio. Body: ```json { "text": "...", "voice": "glados_25s", "format": "pcm", "stream": true, "exaggeration": 0.5, "temperature": 0.8, "top_p": 0.95, "top_k": 1000, "repetition_penalty": 1.2 } ``` - `format`: `pcm` (raw s16le @ 24 kHz, lowest latency, default) or `wav`. - `stream: false` → whole-text one-shot (the A/B quality baseline). - `voice`: predefined name (a `*.wav` in `CBF_VOICES_DIR`) or an absolute path to a clone reference. Omit → server default. - `margin` / `margin_first` / `rtf_prior`: optional scheduler overrides. `GET /health` → `{status, sr, device, default_voice, voices_dir}`. `GET /voices` → `{voices: [stem…], default}` — predefined `*.wav` stems in `CBF_VOICES_DIR` (`_`-prefixed scratch/A-B files excluded). Clone refs are passed per-request as an absolute path and aren't listed. ## Config (env) | var | default | meaning | |---|---|---| | `CBF_MODEL_DEVICE` | `cuda` | `cuda` / `cuda:0` / `cpu` | | `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs | | `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) | | `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind | | `CBF_TF32` | `1` | TF32 matmul/cudnn (free; off with `0`) | | `CBF_SDPA_FLASH` | `1` | flash + mem-efficient SDPA backend | ### Perf notes (measured 2026-06-02, turbo on A6000) - Model loads in **float32** (not the fp16 older notes assumed). - **TF32 + SDPA do not move TTFA** (~0.5s): the first-sentence latency is bound by the sequential AR token decode (T3 Llama, batch-1), not matmul throughput. They stay on (free, help the larger chunks marginally). - **bf16 deferred:** the lever that *would* help batch-1 decode, but `from_pretrained()` has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make a clean cast nontrivial. Not worth the quality risk while ~0.5s TTFA is fine. - **torch.compile: deferred** (research flags a batch-1 regression). ## Dev / test on irv-ml1 ```bash # (from this dir) copy the server into the chatterbox image and run it on GPU 1: scp app.py scheduler.py bench.py lkraven@10.100.79.3:/tmp/cbf/ IMG=$(ssh lkraven@10.100.79.3 "docker images --format '{{.Repository}}:{{.Tag}}' | grep -i chatterbox | grep -v '' | head -1") ssh lkraven@10.100.79.3 "docker run --rm --gpus '\"device=1\"' -e NVIDIA_VISIBLE_DEVICES=1 \ -e HF_HOME=/app/hf_cache -e CBF_VOICES_DIR=/refs -e CBF_DEFAULT_VOICE=glados_25s \ -p 8197:8197 -v /worktank/chatterbox/cache:/app/hf_cache \ -v /worktank/chatterbox/reference_audio:/refs -v /tmp/cbf:/cbf \ $IMG python /cbf/app.py" # then, from the host (or anywhere on the WG net): python bench.py --host http://10.100.79.3:8197 --out /refs/_fast.wav python bench.py --host http://10.100.79.3:8197 --oneshot --out /refs/_oneshot.wav ``` Pull the samples to listen: `scp lkraven@10.100.79.3:/worktank/chatterbox/reference_audio/_*.wav ~/chatterbox-ab/`. ## Acceptance (plan §6) - **Latency:** first-audio < ~0.8s on the deployment GPU. - **No starvation:** `bench.py` reports "stayed ahead"; `test_scheduler.py` green. - **Quality:** operator ear-A/B the streamed output vs the one-shot — join-context loss should be ~imperceptible for multi-sentence text.