feat(omnivoice): streaming /tts + language-safe sanitizer
Add a live-consumer streaming path and text sanitation to the OmniVoice wrapper, so it can front speech-to-speech chat engines (not just the asset-engine's batch WAV use). - POST /tts: chunked 24 kHz mono s16le PCM (or open-ended WAV), driven by the adaptive buffer-ratchet scheduler. Emits the first sentence immediately, then ratchets chunk size up on OmniVoice's ~40x realtime headroom -> sub-second time-to-first-audio. Wire-compatible with chatterbox-fast /tts (both 24 kHz mono PCM). Batch /v1/audio/speech is unchanged for asset/file callers. - scheduler.py: VENDORED byte-faithful copy of chatterbox-fast's pure- Python (torch-free) scheduler, pinned to commit 7631462 (v0.1.0/v0.1.1). Vendor-copy over a shared package (operator call 2026-06-19): the module has no GPU deps, so reuse it without dragging chatterbox-fast's torch tree into this image. Promote to a shared package only on a 3rd consumer or real drift. - sanitize.py: language-safe TTS sanitizer run on both endpoints. Strips markdown, <think> blocks, HTML, and model control tokens; deliberately SKIPS the fork's English-only number/phone normalization that would corrupt OmniVoice's 600-language input. Preserves [laughter]-style tags. - Refactor: shared GenParams base for SpeechRequest + TTSStreamRequest; single GEN_LOCK serializes generation (single-stream interactive). - Dockerfile/playbook: copy + upload the two new modules; build-time `import app` smoke; correct stale "Gradio demo / no FastAPI" comments.
This commit is contained in:
@@ -1,11 +1,12 @@
|
||||
# Deploy OmniVoice (https://github.com/k2-fsa/OmniVoice) to irv-ml1, GPU 0
|
||||
# (RTX 3090). Apache-2.0 zero-shot multilingual voice-cloning TTS, served
|
||||
# via upstream's own Gradio demo (no FastAPI wrapper).
|
||||
# behind our OWN FastAPI wrapper (app.py): batch /v1/audio/speech plus a
|
||||
# streaming /tts driven by the vendored buffer-ratchet scheduler.
|
||||
#
|
||||
# Builds the image locally from stacks/omnivoice/Dockerfile (CUDA 12.8 +
|
||||
# torch 2.8.0 + omnivoice from PyPI), stages the build context under
|
||||
# /opt/docker/compose/omnivoice/, brings it up, and waits for the Gradio
|
||||
# UI on :8199.
|
||||
# torch 2.8.0 + omnivoice from PyPI + vendored scheduler.py/sanitize.py),
|
||||
# stages the build context under /opt/docker/compose/omnivoice/, brings it
|
||||
# up, and waits for /healthz on :8199.
|
||||
#
|
||||
# First run is slow: ~5-10 min docker build + a one-time HF weight pre-warm
|
||||
# (k2-fsa/OmniVoice) on first container start (entrypoint.sh). The wait loop
|
||||
@@ -61,12 +62,24 @@ steps:
|
||||
dest: "{{ compose_dir }}/Dockerfile"
|
||||
mode: "0644"
|
||||
|
||||
- name: Upload app.py (asset-engine FastAPI wrapper)
|
||||
- name: Upload app.py (batch + streaming FastAPI wrapper)
|
||||
upload:
|
||||
src: stacks/omnivoice/app.py
|
||||
dest: "{{ compose_dir }}/app.py"
|
||||
mode: "0644"
|
||||
|
||||
- name: Upload scheduler.py (vendored buffer-ratchet streaming scheduler)
|
||||
upload:
|
||||
src: stacks/omnivoice/scheduler.py
|
||||
dest: "{{ compose_dir }}/scheduler.py"
|
||||
mode: "0644"
|
||||
|
||||
- name: Upload sanitize.py (language-safe TTS text sanitizer)
|
||||
upload:
|
||||
src: stacks/omnivoice/sanitize.py
|
||||
dest: "{{ compose_dir }}/sanitize.py"
|
||||
mode: "0644"
|
||||
|
||||
- name: Stage chatterbox reference voices for cloning (skip _*.wav artifacts)
|
||||
shell: |
|
||||
set -e
|
||||
|
||||
Reference in New Issue
Block a user