feat(omnivoice): streaming /tts + language-safe sanitizer
Add a live-consumer streaming path and text sanitation to the OmniVoice wrapper, so it can front speech-to-speech chat engines (not just the asset-engine's batch WAV use). - POST /tts: chunked 24 kHz mono s16le PCM (or open-ended WAV), driven by the adaptive buffer-ratchet scheduler. Emits the first sentence immediately, then ratchets chunk size up on OmniVoice's ~40x realtime headroom -> sub-second time-to-first-audio. Wire-compatible with chatterbox-fast /tts (both 24 kHz mono PCM). Batch /v1/audio/speech is unchanged for asset/file callers. - scheduler.py: VENDORED byte-faithful copy of chatterbox-fast's pure- Python (torch-free) scheduler, pinned to commit 7631462 (v0.1.0/v0.1.1). Vendor-copy over a shared package (operator call 2026-06-19): the module has no GPU deps, so reuse it without dragging chatterbox-fast's torch tree into this image. Promote to a shared package only on a 3rd consumer or real drift. - sanitize.py: language-safe TTS sanitizer run on both endpoints. Strips markdown, <think> blocks, HTML, and model control tokens; deliberately SKIPS the fork's English-only number/phone normalization that would corrupt OmniVoice's 600-language input. Preserves [laughter]-style tags. - Refactor: shared GenParams base for SpeechRequest + TTSStreamRequest; single GEN_LOCK serializes generation (single-stream interactive). - Dockerfile/playbook: copy + upload the two new modules; build-time `import app` smoke; correct stale "Gradio demo / no FastAPI" comments.
This commit is contained in:
@@ -4,8 +4,9 @@
|
||||
# languages) voice-cloning + voice-design TTS, diffusion-LM architecture,
|
||||
# Apache-2.0. Upstream ships a pip package + its own Gradio demo
|
||||
# (`omnivoice-demo`); there's no official image, so we build a thin CUDA
|
||||
# container around the pip package and run its Gradio server directly.
|
||||
# Unlike index-tts we DON'T need a FastAPI wrapper — OmniVoice serves itself.
|
||||
# container around the pip package and run our OWN FastAPI wrapper (app.py):
|
||||
# a batch OpenAI-style /v1/audio/speech plus a streaming /tts driven by the
|
||||
# vendored buffer-ratchet scheduler (scheduler.py) for live chat consumers.
|
||||
|
||||
ARG CUDA_BASE=nvidia/cuda:12.8.0-cudnn-runtime-ubuntu22.04
|
||||
FROM ${CUDA_BASE}
|
||||
@@ -41,9 +42,15 @@ RUN python -c "import omnivoice, fastapi, soundfile, uvicorn; print('omnivoice w
|
||||
|
||||
WORKDIR /app
|
||||
COPY app.py /app/app.py
|
||||
COPY scheduler.py /app/scheduler.py
|
||||
COPY sanitize.py /app/sanitize.py
|
||||
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
|
||||
RUN chmod +x /usr/local/bin/entrypoint.sh
|
||||
|
||||
# Fail the build if the wrapper (incl. the vendored scheduler + sanitizer) won't
|
||||
# import. Model load is lazy (startup event), so this is a cheap CPU-only check.
|
||||
RUN python -c "import app; print('omnivoice app import OK')"
|
||||
|
||||
EXPOSE 8001
|
||||
|
||||
# Our wrapper exposes /healthz (200 once model + >=1 voice are loaded). Generous
|
||||
|
||||
Reference in New Issue
Block a user