Files
esh-pfi-infrastructure/stacks/omnivoice/README.md
T
vh 06eb487a26 feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices
- app.py: thin FastAPI wrapper exposing OpenAI /v1/audio/speech (+ /v1/audio/voices,
  /healthz) around OmniVoice's Python API; precomputes a voice-clone prompt per voice
  at startup (loaded Whisper auto-transcribes each reference). Replaces the Gradio demo.
- Dockerfile/compose: run the uvicorn wrapper, /healthz healthcheck, project name pinned
  to "omnivoice" so the asset-engine liveness probe matches.
- deploy-omnivoice.yaml: stage chatterbox /refs/*.wav as clone voices (skip _* artifacts)
  + verify the API surface.
- services.yaml: catalog entry (id omnivoice, :8199/v1/audio/speech, voice list sourced
  live from /v1/audio/voices) + reproducibility_audit row.

Verified live on irv-ml1: /healthz ok, 33 voices loaded, test synth -> 24kHz PCM_16 WAV.
2026-06-18 23:03:20 -07:00

78 lines
3.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OmniVoice
[k2-fsa/OmniVoice](https://github.com/k2-fsa/OmniVoice) — zero-shot,
massively-multilingual (**600+ languages**) voice-cloning + voice-design
TTS from the Next-gen Kaldi / k2-fsa team. Diffusion-LM architecture,
RTF as low as ~0.025 (≈40× real-time). **Apache-2.0** — commercially clean
(unlike Voxtral's CC BY-NC).
## What it does
| Capability | Notes |
|---|---|
| Zero-shot voice cloning | Clone from a short reference clip |
| Voice **design** | Synthesize a voice from attributes (gender, age, pitch, accent, whisper, …) — no reference needed |
| 600+ languages | Broadest coverage of any zero-shot TTS |
| Fine control | Non-verbal symbols + pronunciation correction |
## How it's served
Behind our own thin **FastAPI wrapper** ([`app.py`](app.py)) — upstream ships
only a Gradio demo, which we replaced (2026-06-19) so the **asset-engine** can
consume it. Endpoints on `http://10.100.79.3:8199`:
| Endpoint | Purpose |
|---|---|
| `POST /v1/audio/speech` | OpenAI-style `{input, voice, response_format=wav}` → 24 kHz PCM_16 mono |
| `GET /v1/audio/voices` | `{"voices": [...]}` — the staged clone targets |
| `GET /healthz` | readiness (200 once model + ≥1 voice loaded) |
The wrapper loads OmniVoice + a Whisper ASR and **precomputes a voice-clone
prompt per staged reference WAV at startup** (Whisper auto-transcribes each
reference), so per-request latency is just generation. v1 is **clone-only** —
OmniVoice's voice-*design* / language / instruct controls aren't exposed yet.
### Voices — reused from chatterbox
The clone references are chatterbox-fast's `/refs/*.wav`, staged into
`/worktank/omnivoice/voices/` by the deploy playbook (33 named voices at deploy;
`_*.wav` test artifacts skipped). Add more by dropping WAVs there and restarting.
### asset-engine
Catalogued in [`docs/asset-engine/services.yaml`](../../docs/asset-engine/services.yaml)
(`id: omnivoice`, `lifecycle.stack: omnivoice`, `voice` field sourced live from
`/v1/audio/voices`). The compose **project name is pinned to `omnivoice`** so the
liveness probe (docker-ps project-name match) sees it online.
## Placement
- **irv-ml1, GPU 0 (RTX 3090)** — pinned via `OMNIVOICE_GPU_DEVICES=0`.
The A6000 (device 1) is ComfyUI-exclusive after the 2026-06-18 VRAM
consolidation. OmniVoice fits in <5 GB; the 3090 had ~18 GB free.
- Port **8199** (8001 inside the container).
## Deploy
```bash
scripts/elway irv-ml1 --playbook playbooks/deploy-omnivoice.yaml
```
Builds the image locally (CUDA 12.8 + torch 2.8.0 + `omnivoice` from PyPI),
stages the build context under `/opt/docker/compose/omnivoice/`, brings it
up, and waits for the Gradio UI on `:8199`. First boot is slow: ~5-10 min
docker build + a one-time HF weight pre-warm (`k2-fsa/OmniVoice`, entrypoint
pre-download into `${OMNIVOICE_CACHE_DIR}`).
## Tunables
All in `.env` (see `.env.example`): `OMNIVOICE_PORT`, `OMNIVOICE_GPU_DEVICES`,
`OMNIVOICE_TAG`, `OMNIVOICE_VERSION` (optional PyPI pin),
`OMNIVOICE_CACHE_DIR`, `OMNIVOICE_VOICES_DIR`. Drop reference WAV/FLAC into
`/worktank/omnivoice/voices/` to stage cloning sources.
## Footprint
- **Disk**: HF weight cache under `/worktank/omnivoice/hf_cache`.
- **VRAM**: <5 GB (docs cite 4 GB+ GPUs).