refactor: extract chatterbox-fast to its own repo (vh/chatterbox-fast)
chatterbox-fast is authored software with a test suite, not a config-mirror stack — so it moves to its own MIT-licensed, versioned, CI'd repo (gitea vh/chatterbox-fast, v0.1.0) following the sister-repo pattern. Replace stacks/chatterbox-fast/ with a pointer README; the moved code (scheduler/app/bench/tests/Dockerfile/compose) now lives in the new repo. The deployed :8197 service is unaffected (still runs the legacy devnen-based image; self-contained-image migration is an optional follow-up). The fleet catalog entry stays in docs/asset-engine/services.yaml.
This commit is contained in:
@@ -1,129 +1,27 @@
|
||||
# chatterbox-fast — streaming TTS engine
|
||||
# chatterbox-fast — moved to its own repository
|
||||
|
||||
Custom streaming server on top of `ChatterboxTurboTTS` that delivers
|
||||
**sub-second time-to-first-audio** while keeping turbo's full quality. Workload:
|
||||
**single-stream interactive**. Deployed (Phase 3) **alongside** the live
|
||||
`chatterbox` (:8196) on irv-ml1, burned in, then catalog-flipped.
|
||||
The chatterbox-fast streaming TTS engine now lives at:
|
||||
|
||||
Design: [`docs/design/chatterbox-fast-plan.md`](../../docs/design/chatterbox-fast-plan.md)
|
||||
(canonical plan). The abandoned native-frame-streaming arc is recorded in
|
||||
`persistent-memory.md` → *Tried and abandoned*.
|
||||
> **https://gitea.phasefinal.com/vh/chatterbox-fast** — MIT, open-source (`v0.1.0`)
|
||||
|
||||
## How it works — adaptive buffer-ratchet chunking
|
||||
Extracted from this workspace on 2026-06-02. Unlike the other entries under
|
||||
`stacks/` (thin compose+conf wrappers around upstream images), chatterbox-fast is
|
||||
**authored software with a test suite** — so it gets its own versioned, CI'd repo
|
||||
following the sister-repo pattern (task-board / vor / asset-engine / …). The new
|
||||
repo owns the code, the self-contained Dockerfile, the tests, and three
|
||||
public-domain LibriVox starter voices.
|
||||
|
||||
The engine never splits mid-sentence (keeps each chunk prosodically coherent).
|
||||
Instead it rides Chatterbox's faster-than-realtime generation (RTF ~3.4–3.8×):
|
||||
## Deployed service
|
||||
|
||||
1. **Chunk 1 = first sentence**, generated and emitted immediately (~0.66s
|
||||
first-audio). Latency-critical.
|
||||
2. **While chunk N plays, generate chunk N+1** by greedily accumulating whole
|
||||
sentences until the next would exceed `margin × audio_buffered_remaining`.
|
||||
3. Each chunk's playback buys wall-clock for a ~3× bigger next chunk, so after
|
||||
2-3 joins the rest of the paragraph is one big near-full-context chunk.
|
||||
Context loss is confined to those few sentence-boundary joins.
|
||||
4. Driven off **measured** RTF + sec-per-char (EMA), not constants.
|
||||
5. **Starvation relief:** if a mid-stream sentence is too long to generate
|
||||
within the current buffer, its *clause* boundaries are exposed so chunks pack
|
||||
to commas (natural pauses) — never a mid-clause split. A long *comma-less*
|
||||
sentence after a short opener is the one unavoidable case: the rule is
|
||||
honored and the brief gap is **flagged** (`drained > 0`), never hidden.
|
||||
The live service on **irv-ml1:8197** (catalog entry `chatterbox-fast`, status
|
||||
`ready`) currently runs the original devnen-based image from before the extraction.
|
||||
Migrating it to the self-contained image from the new repo is an optional
|
||||
follow-up — note that image needs an `HF_TOKEN` at runtime (the Chatterbox-Turbo
|
||||
model is public + MIT, but the `chatterbox-tts` package requires a token to
|
||||
download it; the legacy devnen image sidesteps this).
|
||||
|
||||
This only works because RTF > 1 — a sub-realtime model (e.g. Fish) would starve
|
||||
regardless of chunking. That is why this is the chatterbox-specific answer.
|
||||
## Catalog
|
||||
|
||||
## Files
|
||||
|
||||
| file | role |
|
||||
|---|---|
|
||||
| `scheduler.py` | The adaptive-chunk scheduler. **GPU-free, pure logic** — the meat. |
|
||||
| `test_scheduler.py` | GPU-free simulation: asserts no-starvation + ratchet. `python test_scheduler.py` or `pytest`. |
|
||||
| `app.py` | FastAPI server: model holder + `POST /tts` (StreamingResponse) + `GET /health`. |
|
||||
| `bench.py` | Client: ground-truth TTFB + real 1×-consumer starvation check; saves `.wav` for A/B. |
|
||||
| `Dockerfile` | Thin overlay: `FROM local/chatterbox:v1` + our two modules. |
|
||||
| `compose.yaml` · `.env.example` | Deploy on irv-ml1 alongside the live `chatterbox`. |
|
||||
|
||||
## Deploy (Phase 3)
|
||||
|
||||
```bash
|
||||
scripts/deploy-stack.sh irv-ml1 chatterbox-fast # push compose+code to the host
|
||||
# then on irv-ml1, in /opt/docker/compose/chatterbox-fast/ (after copying .env):
|
||||
docker compose build && docker compose up -d
|
||||
```
|
||||
|
||||
`Dockerfile` is `FROM local/chatterbox:v1` (the sibling stack's image — must exist
|
||||
on irv-ml1) + `COPY scheduler.py app.py`. GPU pin and voices/cache paths come from
|
||||
`.env` (see `.env.example`). GPU is **device 1 (A6000)** — measured footprint is
|
||||
**5.34 GB** (turbo loads fp32), so the 3090's ~3.8 GB free does **not** fit it.
|
||||
|
||||
**Deployed 2026-06-02** alongside the live `chatterbox` (:8196): healthy on
|
||||
:8197, TTFB ~0.5s, no starvation, ~7 GB free left on the A6000.
|
||||
|
||||
## API
|
||||
|
||||
`POST /tts` → streamed audio. Body:
|
||||
|
||||
```json
|
||||
{ "text": "...", "voice": "glados_25s", "format": "pcm",
|
||||
"stream": true, "exaggeration": 0.5, "temperature": 0.8,
|
||||
"top_p": 0.95, "top_k": 1000, "repetition_penalty": 1.2 }
|
||||
```
|
||||
|
||||
- `format`: `pcm` (raw s16le @ 24 kHz, lowest latency, default) or `wav`.
|
||||
- `stream: false` → whole-text one-shot (the A/B quality baseline).
|
||||
- `voice`: predefined name (a `*.wav` in `CBF_VOICES_DIR`) or an absolute path
|
||||
to a clone reference. Omit → server default.
|
||||
- `margin` / `margin_first` / `rtf_prior`: optional scheduler overrides.
|
||||
|
||||
`GET /health` → `{status, sr, device, default_voice, voices_dir}`.
|
||||
|
||||
`GET /voices` → `{voices: [stem…], default}` — predefined `*.wav` stems in
|
||||
`CBF_VOICES_DIR` (`_`-prefixed scratch/A-B files excluded). Clone refs are passed
|
||||
per-request as an absolute path and aren't listed.
|
||||
|
||||
## Config (env)
|
||||
|
||||
| var | default | meaning |
|
||||
|---|---|---|
|
||||
| `CBF_MODEL_DEVICE` | `cuda` | `cuda` / `cuda:0` / `cpu` |
|
||||
| `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs |
|
||||
| `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) |
|
||||
| `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind |
|
||||
| `CBF_TF32` | `1` | TF32 matmul/cudnn (free; off with `0`) |
|
||||
| `CBF_SDPA_FLASH` | `1` | flash + mem-efficient SDPA backend |
|
||||
|
||||
### Perf notes (measured 2026-06-02, turbo on A6000)
|
||||
|
||||
- Model loads in **float32** (not the fp16 older notes assumed).
|
||||
- **TF32 + SDPA do not move TTFA** (~0.5s): the first-sentence latency is bound by
|
||||
the sequential AR token decode (T3 Llama, batch-1), not matmul throughput. They
|
||||
stay on (free, help the larger chunks marginally).
|
||||
- **bf16 deferred:** the lever that *would* help batch-1 decode, but `from_pretrained()`
|
||||
has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make
|
||||
a clean cast nontrivial. Not worth the quality risk while ~0.5s TTFA is fine.
|
||||
- **torch.compile: deferred** (research flags a batch-1 regression).
|
||||
|
||||
## Dev / test on irv-ml1
|
||||
|
||||
```bash
|
||||
# (from this dir) copy the server into the chatterbox image and run it on GPU 1:
|
||||
scp app.py scheduler.py bench.py lkraven@10.100.79.3:/tmp/cbf/
|
||||
IMG=$(ssh lkraven@10.100.79.3 "docker images --format '{{.Repository}}:{{.Tag}}' | grep -i chatterbox | grep -v '<none>' | head -1")
|
||||
ssh lkraven@10.100.79.3 "docker run --rm --gpus '\"device=1\"' -e NVIDIA_VISIBLE_DEVICES=1 \
|
||||
-e HF_HOME=/app/hf_cache -e CBF_VOICES_DIR=/refs -e CBF_DEFAULT_VOICE=glados_25s \
|
||||
-p 8197:8197 -v /worktank/chatterbox/cache:/app/hf_cache \
|
||||
-v /worktank/chatterbox/reference_audio:/refs -v /tmp/cbf:/cbf \
|
||||
$IMG python /cbf/app.py"
|
||||
|
||||
# then, from the host (or anywhere on the WG net):
|
||||
python bench.py --host http://10.100.79.3:8197 --out /refs/_fast.wav
|
||||
python bench.py --host http://10.100.79.3:8197 --oneshot --out /refs/_oneshot.wav
|
||||
```
|
||||
|
||||
Pull the samples to listen: `scp lkraven@10.100.79.3:/worktank/chatterbox/reference_audio/_*.wav ~/chatterbox-ab/`.
|
||||
|
||||
## Acceptance (plan §6)
|
||||
|
||||
- **Latency:** first-audio < ~0.8s on the deployment GPU.
|
||||
- **No starvation:** `bench.py` reports "stayed ahead"; `test_scheduler.py` green.
|
||||
- **Quality:** operator ear-A/B the streamed output vs the one-shot — join-context
|
||||
loss should be ~imperceptible for multi-sentence text.
|
||||
The fleet catalog entry stays here in
|
||||
[`docs/asset-engine/services.yaml`](../../docs/asset-engine/services.yaml) — it's
|
||||
fleet-wide (consumed by asset-engine), not chatterbox-fast-specific.
|
||||
|
||||
Reference in New Issue
Block a user