feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices

- app.py: thin FastAPI wrapper exposing OpenAI /v1/audio/speech (+ /v1/audio/voices,
  /healthz) around OmniVoice's Python API; precomputes a voice-clone prompt per voice
  at startup (loaded Whisper auto-transcribes each reference). Replaces the Gradio demo.
- Dockerfile/compose: run the uvicorn wrapper, /healthz healthcheck, project name pinned
  to "omnivoice" so the asset-engine liveness probe matches.
- deploy-omnivoice.yaml: stage chatterbox /refs/*.wav as clone voices (skip _* artifacts)
  + verify the API surface.
- services.yaml: catalog entry (id omnivoice, :8199/v1/audio/speech, voice list sourced
  live from /v1/audio/voices) + reproducibility_audit row.

Verified live on irv-ml1: /healthz ok, 33 voices loaded, test synth -> 24kHz PCM_16 WAV.
This commit is contained in:
vh
2026-06-18 23:03:20 -07:00
parent 984b72757f
commit 06eb487a26
6 changed files with 255 additions and 37 deletions
+26 -7
View File
@@ -17,14 +17,33 @@ RTF as low as ~0.025 (≈40× real-time). **Apache-2.0** — commercially clean
## How it's served
Upstream ships **its own Gradio demo** (`omnivoice-demo`), so this stack
just runs that — no custom wrapper. That means the surface is the **Gradio
UI + Gradio API**, *not* an OpenAI-compatible `/v1/audio/speech` endpoint.
Behind our own thin **FastAPI wrapper** ([`app.py`](app.py)) — upstream ships
only a Gradio demo, which we replaced (2026-06-19) so the **asset-engine** can
consume it. Endpoints on `http://10.100.79.3:8199`:
- UI: `http://10.100.79.3:8199/`
- Programmatic: the Gradio API under `/gradio_api/` (or `/config` to
introspect). If you later want OpenAI-compat for asset-engine, add a thin
FastAPI wrapper like [`stacks/index-tts/app.py`](../index-tts/app.py).
| Endpoint | Purpose |
|---|---|
| `POST /v1/audio/speech` | OpenAI-style `{input, voice, response_format=wav}` → 24 kHz PCM_16 mono |
| `GET /v1/audio/voices` | `{"voices": [...]}` — the staged clone targets |
| `GET /healthz` | readiness (200 once model + ≥1 voice loaded) |
The wrapper loads OmniVoice + a Whisper ASR and **precomputes a voice-clone
prompt per staged reference WAV at startup** (Whisper auto-transcribes each
reference), so per-request latency is just generation. v1 is **clone-only** —
OmniVoice's voice-*design* / language / instruct controls aren't exposed yet.
### Voices — reused from chatterbox
The clone references are chatterbox-fast's `/refs/*.wav`, staged into
`/worktank/omnivoice/voices/` by the deploy playbook (33 named voices at deploy;
`_*.wav` test artifacts skipped). Add more by dropping WAVs there and restarting.
### asset-engine
Catalogued in [`docs/asset-engine/services.yaml`](../../docs/asset-engine/services.yaml)
(`id: omnivoice`, `lifecycle.stack: omnivoice`, `voice` field sourced live from
`/v1/audio/voices`). The compose **project name is pinned to `omnivoice`** so the
liveness probe (docker-ps project-name match) sees it online.
## Placement