feat(omnivoice): tune streaming defaults (16-step + aggressive packing)

Empirical follow-up to the streaming /tts smoke test on the 3090. OmniVoice
is diffusion: a ~fixed per-call overhead (~1.5s at 32 steps, ~0.7s at 16)
dominates regardless of chunk length, so the upstream-claimed 40x RTF does
NOT hold here (measured ~2.8x/32-step, ~5.6x/16-step) and the chatterbox-
tuned scheduler over-chunks and starves.

- Streaming /tts defaults to num_step=16 (TTFA ~1.5s -> ~0.7s); batch
  /v1/audio/speech stays num_step=32 for quality. Per-request override intact.
- Scheduler prior raised to rtf_prior=20 (env OMNIVOICE_STREAM_RTF_PRIOR,
  wired through compose + .env.example) so it packs whole-text-minus-first-
  sentence into a few chunks: validated ~3 chunks, no starvation, total wall
  ~= one-shot, less per-chunk silence padding.
- Docs corrected: the "sub-second / 40x" claims were wrong; streaming has a
  diffusion TTFA floor (~0.7s) and wins mainly on long replies. chatterbox-
  fast (autoregressive, ~0.5s TTFA) stays the lowest-latency front-end;
  OmniVoice is the multilingual / voice-design complement.
This commit is contained in:
vh
2026-06-19 22:58:55 -07:00
parent 288d085236
commit cd92b85157
4 changed files with 50 additions and 10 deletions
+5
View File
@@ -18,3 +18,8 @@ OMNIVOICE_VERSION=
# Persistent HF weight cache + reference-voice staging on /worktank.
OMNIVOICE_CACHE_DIR=/worktank/omnivoice/hf_cache
OMNIVOICE_VOICES_DIR=/worktank/omnivoice/voices
# Streaming /tts scheduler prior. High = pack aggressively (OmniVoice is diffusion
# with a ~fixed per-call overhead; low priors over-chunk and starve). 20 is
# validated clean on the 3090. Per-request `rtf_prior` overrides this.
OMNIVOICE_STREAM_RTF_PRIOR=20