feat(chatterbox-fast): Phase 2 parity + perf levers
- /voices endpoint lists predefined voice stems (excludes `_`-prefixed bench/A-B scratch wavs); shared _predefined_wavs() also feeds default-voice discovery. - Perf levers: TF32 matmul/cudnn + flash/mem-efficient SDPA, default ON, env-gated (CBF_TF32 / CBF_SDPA_FLASH). Startup logs model dtype. Measured on irv-ml1 (turbo, A6000): the model loads FLOAT32 (not the fp16 older notes assumed). TF32+SDPA do NOT move TTFA (489->514ms, noise) — first-sentence latency is bound by the sequential AR token decode at batch-1, not matmul throughput. bf16 (the lever that would help) is DEFERRED: from_pretrained() has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make a clean cast nontrivial; not worth the quality risk at ~0.5s TTFA. torch.compile also deferred (batch-1 regression). Findings recorded in README. Voice management parity (predefined dir + per-request clone refs) was already in the Phase-1 resolve path; /voices completes the surface.
This commit is contained in:
@@ -60,6 +60,10 @@ Phase 3 will add `compose.yaml`, `Dockerfile`, `.env.example`.
|
||||
|
||||
`GET /health` → `{status, sr, device, default_voice, voices_dir}`.
|
||||
|
||||
`GET /voices` → `{voices: [stem…], default}` — predefined `*.wav` stems in
|
||||
`CBF_VOICES_DIR` (`_`-prefixed scratch/A-B files excluded). Clone refs are passed
|
||||
per-request as an absolute path and aren't listed.
|
||||
|
||||
## Config (env)
|
||||
|
||||
| var | default | meaning |
|
||||
@@ -68,6 +72,19 @@ Phase 3 will add `compose.yaml`, `Dockerfile`, `.env.example`.
|
||||
| `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs |
|
||||
| `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) |
|
||||
| `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind |
|
||||
| `CBF_TF32` | `1` | TF32 matmul/cudnn (free; off with `0`) |
|
||||
| `CBF_SDPA_FLASH` | `1` | flash + mem-efficient SDPA backend |
|
||||
|
||||
### Perf notes (measured 2026-06-02, turbo on A6000)
|
||||
|
||||
- Model loads in **float32** (not the fp16 older notes assumed).
|
||||
- **TF32 + SDPA do not move TTFA** (~0.5s): the first-sentence latency is bound by
|
||||
the sequential AR token decode (T3 Llama, batch-1), not matmul throughput. They
|
||||
stay on (free, help the larger chunks marginally).
|
||||
- **bf16 deferred:** the lever that *would* help batch-1 decode, but `from_pretrained()`
|
||||
has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make
|
||||
a clean cast nontrivial. Not worth the quality risk while ~0.5s TTFA is fine.
|
||||
- **torch.compile: deferred** (research flags a batch-1 regression).
|
||||
|
||||
## Dev / test on irv-ml1
|
||||
|
||||
|
||||
Reference in New Issue
Block a user