feat(omnivoice): tune streaming defaults (16-step + aggressive packing)
Empirical follow-up to the streaming /tts smoke test on the 3090. OmniVoice is diffusion: a ~fixed per-call overhead (~1.5s at 32 steps, ~0.7s at 16) dominates regardless of chunk length, so the upstream-claimed 40x RTF does NOT hold here (measured ~2.8x/32-step, ~5.6x/16-step) and the chatterbox- tuned scheduler over-chunks and starves. - Streaming /tts defaults to num_step=16 (TTFA ~1.5s -> ~0.7s); batch /v1/audio/speech stays num_step=32 for quality. Per-request override intact. - Scheduler prior raised to rtf_prior=20 (env OMNIVOICE_STREAM_RTF_PRIOR, wired through compose + .env.example) so it packs whole-text-minus-first- sentence into a few chunks: validated ~3 chunks, no starvation, total wall ~= one-shot, less per-chunk silence padding. - Docs corrected: the "sub-second / 40x" claims were wrong; streaming has a diffusion TTFA floor (~0.7s) and wins mainly on long replies. chatterbox- fast (autoregressive, ~0.5s TTFA) stays the lowest-latency front-end; OmniVoice is the multilingual / voice-design complement.
This commit is contained in:
@@ -36,15 +36,32 @@ reference), so per-request latency is just generation. The full generation
|
||||
surface is exposed: zero-shot **clone** (`voice`) and/or voice-**design**
|
||||
(`instruct`), plus `language` / `speed` / `duration` and the diffusion knobs.
|
||||
|
||||
### Streaming — sub-second time-to-first-audio
|
||||
### Streaming — earlier first-audio (with a diffusion floor)
|
||||
|
||||
`POST /tts` (`stream=true`, default) runs the **adaptive buffer-ratchet
|
||||
scheduler** vendored from chatterbox-fast ([`scheduler.py`](scheduler.py)): it
|
||||
emits the first sentence immediately and ratchets chunk size up on OmniVoice's
|
||||
~40× realtime headroom, so a live consumer hears speech start in ~tens of ms
|
||||
instead of waiting for the whole utterance. `stream=false` is a whole-text
|
||||
one-shot for A/B. Scheduler tunables (`margin`, `margin_first`, `rtf_prior`,
|
||||
`sec_per_char_prior`) are per-request overrides.
|
||||
emits the first sentence immediately, then packs the rest into a few chunks so a
|
||||
live consumer hears speech start sooner than waiting for the whole utterance.
|
||||
`stream=false` is a whole-text one-shot for A/B.
|
||||
|
||||
**Measured reality (3090, not the upstream-claimed 40× RTF):** OmniVoice is a
|
||||
diffusion model, so each `generate()` call has a **~fixed per-call overhead**
|
||||
(~1.5 s at `num_step=32`, ~0.7 s at 16) that sets a **time-to-first-audio
|
||||
floor** — short and long chunks cost nearly the same. Server-side TTFA is
|
||||
therefore ~0.7 s (streaming default, 16 steps), **not** sub-second-at-full-
|
||||
quality. Effective RTF is ~2.8× (32 steps) / ~5.6× (16 steps). The win over
|
||||
one-shot is small for short replies and grows with length (one-shot TTFA scales
|
||||
with the whole utterance; streaming stays ~flat at the first-sentence cost).
|
||||
For absolute-lowest TTFA, **chatterbox-fast** (autoregressive, ~0.5 s) remains
|
||||
the better front-end; OmniVoice is the multilingual / voice-design complement.
|
||||
|
||||
Defaults tuned for this: **streaming `num_step=16`** (batch `/v1/audio/speech`
|
||||
stays 32 for quality), and an **aggressive packing prior** (`rtf_prior=20`, env
|
||||
`OMNIVOICE_STREAM_RTF_PRIOR`) — diffusion's fixed overhead makes the chatterbox
|
||||
default over-chunk and starve, so we pack whole-text-minus-first-sentence into a
|
||||
few chunks (validated: ~3 chunks, no starvation, total ≈ one-shot). Scheduler
|
||||
tunables (`margin`, `margin_first`, `rtf_prior`, `sec_per_char_prior`) and
|
||||
`num_step` are per-request overrides.
|
||||
|
||||
`scheduler.py` is a **vendored byte-faithful copy** (not a dependency) of
|
||||
chatterbox-fast's pure-Python, torch-free scheduler — see its header for the
|
||||
|
||||
Reference in New Issue
Block a user