feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices
- app.py: thin FastAPI wrapper exposing OpenAI /v1/audio/speech (+ /v1/audio/voices, /healthz) around OmniVoice's Python API; precomputes a voice-clone prompt per voice at startup (loaded Whisper auto-transcribes each reference). Replaces the Gradio demo. - Dockerfile/compose: run the uvicorn wrapper, /healthz healthcheck, project name pinned to "omnivoice" so the asset-engine liveness probe matches. - deploy-omnivoice.yaml: stage chatterbox /refs/*.wav as clone voices (skip _* artifacts) + verify the API surface. - services.yaml: catalog entry (id omnivoice, :8199/v1/audio/speech, voice list sourced live from /v1/audio/voices) + reproducibility_audit row. Verified live on irv-ml1: /healthz ok, 33 voices loaded, test synth -> 24kHz PCM_16 WAV.
This commit is contained in:
@@ -1,9 +1,10 @@
|
||||
# OmniVoice (k2-fsa/OmniVoice) — zero-shot, massively-multilingual (600+
|
||||
# language) voice-cloning + voice-design TTS, diffusion-LM, Apache-2.0.
|
||||
# Served via upstream's own Gradio demo. NOTE: this exposes the Gradio UI
|
||||
# + Gradio API, NOT an OpenAI-compatible /v1/audio/speech endpoint — wrap
|
||||
# it later (à la stacks/index-tts/app.py) if asset-engine integration is
|
||||
# wanted. For now it's a "stand it up and try it" UI.
|
||||
# Served behind our OWN thin FastAPI wrapper (stacks/omnivoice/app.py) exposing
|
||||
# OpenAI-compatible /v1/audio/speech (+ /v1/audio/voices, /healthz) so the
|
||||
# asset-engine can consume it. Upstream ships only a Gradio demo; the wrapper
|
||||
# replaced it (2026-06-19). Voices are reference WAVs in ${OMNIVOICE_VOICES_DIR}
|
||||
# (the reused chatterbox /refs/*.wav); clone prompts are precomputed at startup.
|
||||
#
|
||||
# Build: local image from the Dockerfile in this dir. Weights download
|
||||
# from HF (k2-fsa/OmniVoice) on first boot into ${OMNIVOICE_CACHE_DIR}.
|
||||
@@ -14,6 +15,10 @@
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
# Compose project name MUST equal the catalog lifecycle.stack ("omnivoice") or the
|
||||
# asset-engine liveness probe (docker ps project-name match) shows it OFFLINE.
|
||||
name: omnivoice
|
||||
|
||||
services:
|
||||
omnivoice:
|
||||
image: local/omnivoice:${OMNIVOICE_TAG:-latest}
|
||||
@@ -34,15 +39,15 @@ services:
|
||||
- ${OMNIVOICE_CACHE_DIR}:/app/hf_cache
|
||||
- ${OMNIVOICE_VOICES_DIR}:/app/voices
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8001/ || exit 1"]
|
||||
test: ["CMD-SHELL", "wget -q -O /dev/null http://localhost:8001/healthz || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# First boot: weight pre-warm download (entrypoint) + CUDA warmup.
|
||||
start_period: 900s
|
||||
# First boot: OmniVoice + Whisper ASR pre-warm + cloning every staged voice.
|
||||
start_period: 1200s
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=OmniVoice
|
||||
- homepage.icon=mdi-account-voice
|
||||
- homepage.description=Zero-shot multilingual voice-cloning TTS (irv-ml1, 3090)
|
||||
- homepage.href=http://10.100.79.3:${OMNIVOICE_PORT}
|
||||
- homepage.description=Zero-shot multilingual voice-cloning TTS, OpenAI /v1/audio/speech (irv-ml1, 3090)
|
||||
- homepage.href=http://10.100.79.3:${OMNIVOICE_PORT}/docs
|
||||
|
||||
Reference in New Issue
Block a user