feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices
- app.py: thin FastAPI wrapper exposing OpenAI /v1/audio/speech (+ /v1/audio/voices, /healthz) around OmniVoice's Python API; precomputes a voice-clone prompt per voice at startup (loaded Whisper auto-transcribes each reference). Replaces the Gradio demo. - Dockerfile/compose: run the uvicorn wrapper, /healthz healthcheck, project name pinned to "omnivoice" so the asset-engine liveness probe matches. - deploy-omnivoice.yaml: stage chatterbox /refs/*.wav as clone voices (skip _* artifacts) + verify the API surface. - services.yaml: catalog entry (id omnivoice, :8199/v1/audio/speech, voice list sourced live from /v1/audio/voices) + reproducibility_audit row. Verified live on irv-ml1: /healthz ok, 33 voices loaded, test synth -> 24kHz PCM_16 WAV.
This commit is contained in:
@@ -17,14 +17,33 @@ RTF as low as ~0.025 (≈40× real-time). **Apache-2.0** — commercially clean
|
||||
|
||||
## How it's served
|
||||
|
||||
Upstream ships **its own Gradio demo** (`omnivoice-demo`), so this stack
|
||||
just runs that — no custom wrapper. That means the surface is the **Gradio
|
||||
UI + Gradio API**, *not* an OpenAI-compatible `/v1/audio/speech` endpoint.
|
||||
Behind our own thin **FastAPI wrapper** ([`app.py`](app.py)) — upstream ships
|
||||
only a Gradio demo, which we replaced (2026-06-19) so the **asset-engine** can
|
||||
consume it. Endpoints on `http://10.100.79.3:8199`:
|
||||
|
||||
- UI: `http://10.100.79.3:8199/`
|
||||
- Programmatic: the Gradio API under `/gradio_api/` (or `/config` to
|
||||
introspect). If you later want OpenAI-compat for asset-engine, add a thin
|
||||
FastAPI wrapper like [`stacks/index-tts/app.py`](../index-tts/app.py).
|
||||
| Endpoint | Purpose |
|
||||
|---|---|
|
||||
| `POST /v1/audio/speech` | OpenAI-style `{input, voice, response_format=wav}` → 24 kHz PCM_16 mono |
|
||||
| `GET /v1/audio/voices` | `{"voices": [...]}` — the staged clone targets |
|
||||
| `GET /healthz` | readiness (200 once model + ≥1 voice loaded) |
|
||||
|
||||
The wrapper loads OmniVoice + a Whisper ASR and **precomputes a voice-clone
|
||||
prompt per staged reference WAV at startup** (Whisper auto-transcribes each
|
||||
reference), so per-request latency is just generation. v1 is **clone-only** —
|
||||
OmniVoice's voice-*design* / language / instruct controls aren't exposed yet.
|
||||
|
||||
### Voices — reused from chatterbox
|
||||
|
||||
The clone references are chatterbox-fast's `/refs/*.wav`, staged into
|
||||
`/worktank/omnivoice/voices/` by the deploy playbook (33 named voices at deploy;
|
||||
`_*.wav` test artifacts skipped). Add more by dropping WAVs there and restarting.
|
||||
|
||||
### asset-engine
|
||||
|
||||
Catalogued in [`docs/asset-engine/services.yaml`](../../docs/asset-engine/services.yaml)
|
||||
(`id: omnivoice`, `lifecycle.stack: omnivoice`, `voice` field sourced live from
|
||||
`/v1/audio/voices`). The compose **project name is pinned to `omnivoice`** so the
|
||||
liveness probe (docker-ps project-name match) sees it online.
|
||||
|
||||
## Placement
|
||||
|
||||
|
||||
Reference in New Issue
Block a user