feat(chatterbox-fast): Phase 2 parity + perf levers

- /voices endpoint lists predefined voice stems (excludes `_`-prefixed bench/A-B
  scratch wavs); shared _predefined_wavs() also feeds default-voice discovery.
- Perf levers: TF32 matmul/cudnn + flash/mem-efficient SDPA, default ON, env-gated
  (CBF_TF32 / CBF_SDPA_FLASH). Startup logs model dtype.

Measured on irv-ml1 (turbo, A6000): the model loads FLOAT32 (not the fp16 older
notes assumed). TF32+SDPA do NOT move TTFA (489->514ms, noise) — first-sentence
latency is bound by the sequential AR token decode at batch-1, not matmul
throughput. bf16 (the lever that would help) is DEFERRED: from_pretrained() has no
dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make a
clean cast nontrivial; not worth the quality risk at ~0.5s TTFA. torch.compile
also deferred (batch-1 regression). Findings recorded in README.

Voice management parity (predefined dir + per-request clone refs) was already in
the Phase-1 resolve path; /voices completes the surface.
This commit is contained in:
vh
2026-06-01 22:57:09 -07:00
parent 7cd39001b2
commit 3a92fcd943
2 changed files with 73 additions and 5 deletions
+17
View File
@@ -60,6 +60,10 @@ Phase 3 will add `compose.yaml`, `Dockerfile`, `.env.example`.
`GET /health` → `{status, sr, device, default_voice, voices_dir}`.
`GET /voices` → `{voices: [stem…], default}` — predefined `*.wav` stems in
`CBF_VOICES_DIR` (`_`-prefixed scratch/A-B files excluded). Clone refs are passed
per-request as an absolute path and aren't listed.
## Config (env)
| var | default | meaning |
@@ -68,6 +72,19 @@ Phase 3 will add `compose.yaml`, `Dockerfile`, `.env.example`.
| `CBF_VOICES_DIR` | `/refs` | dir of predefined voice wavs |
| `CBF_DEFAULT_VOICE` | first wav in dir | default reference wav (path or name) |
| `CBF_BIND` / `CBF_PORT` | `0.0.0.0` / `8197` | uvicorn bind |
| `CBF_TF32` | `1` | TF32 matmul/cudnn (free; off with `0`) |
| `CBF_SDPA_FLASH` | `1` | flash + mem-efficient SDPA backend |
### Perf notes (measured 2026-06-02, turbo on A6000)
- Model loads in **float32** (not the fp16 older notes assumed).
- **TF32 + SDPA do not move TTFA** (~0.5s): the first-sentence latency is bound by
the sequential AR token decode (T3 Llama, batch-1), not matmul throughput. They
stay on (free, help the larger chunks marginally).
- **bf16 deferred:** the lever that *would* help batch-1 decode, but `from_pretrained()`
has no dtype arg and turbo's fp32 conditioning path + dtype-sensitive vocoder make
a clean cast nontrivial. Not worth the quality risk while ~0.5s TTFA is fine.
- **torch.compile: deferred** (research flags a batch-1 regression).
## Dev / test on irv-ml1