feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices

- app.py: thin FastAPI wrapper exposing OpenAI /v1/audio/speech (+ /v1/audio/voices,
  /healthz) around OmniVoice's Python API; precomputes a voice-clone prompt per voice
  at startup (loaded Whisper auto-transcribes each reference). Replaces the Gradio demo.
- Dockerfile/compose: run the uvicorn wrapper, /healthz healthcheck, project name pinned
  to "omnivoice" so the asset-engine liveness probe matches.
- deploy-omnivoice.yaml: stage chatterbox /refs/*.wav as clone voices (skip _* artifacts)
  + verify the API surface.
- services.yaml: catalog entry (id omnivoice, :8199/v1/audio/speech, voice list sourced
  live from /v1/audio/voices) + reproducibility_audit row.

Verified live on irv-ml1: /healthz ok, 33 voices loaded, test synth -> 24kHz PCM_16 WAV.
This commit is contained in:
vh
2026-06-18 23:03:20 -07:00
parent 984b72757f
commit 06eb487a26
6 changed files with 255 additions and 37 deletions
+28 -11
View File
@@ -61,6 +61,25 @@ steps:
dest: "{{ compose_dir }}/Dockerfile"
mode: "0644"
- name: Upload app.py (asset-engine FastAPI wrapper)
upload:
src: stacks/omnivoice/app.py
dest: "{{ compose_dir }}/app.py"
mode: "0644"
- name: Stage chatterbox reference voices for cloning (skip _*.wav artifacts)
shell: |
set -e
mkdir -p {{ voices_dir }}
docker exec chatterbox-fast sh -c 'ls /refs/*.wav' | while read -r f; do
b=$(basename "$f")
case "$b" in _*) continue;; esac
docker cp "chatterbox-fast:$f" "{{ voices_dir }}/$b"
done
echo "staged:"; ls {{ voices_dir }}
# Skip if already staged (Emily.wav is a proxy for "voices present").
when: "[ ! -f {{ voices_dir }}/Emily.wav ]"
- name: Upload entrypoint.sh
upload:
src: stacks/omnivoice/entrypoint.sh
@@ -85,26 +104,24 @@ steps:
- name: docker compose up -d
shell: cd {{ compose_dir }} && docker compose up -d
- name: Wait for the Gradio UI to respond (allow ~20 min for weight pre-warm)
- name: Wait for /healthz (allow ~25 min for weight + Whisper pre-warm + voice cloning)
shell: |
for i in $(seq 1 240); do
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/ && exit 0
for i in $(seq 1 300); do
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/healthz && exit 0
sleep 5
done
exit 1
changed_when: "false"
verify:
- name: Gradio UI returns 200
shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/
- name: /healthz returns 200
shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/healthz
changed_when: "false"
- name: /v1/audio/voices lists the reused chatterbox voices
shell: curl -sf http://localhost:{{ host_port }}/v1/audio/voices | grep -q '"voices"'
changed_when: "false"
- name: Container is running
shell: docker inspect omnivoice --format '{{.State.Status}}' | grep -q running
changed_when: "false"
- name: Container is healthy (or still starting weights)
shell: |
s=$(docker inspect omnivoice --format '{{.State.Health.Status}}' 2>/dev/null)
echo "health: $s"; [ "$s" = healthy ] || [ "$s" = starting ]
changed_when: "false"