feat(zonos): OpenAI-ish REST adapter for asset-engine routing

Upstream Zonos ships only Gradio + Python SDK — no REST surface — so
asset-engine (which routes a clean JSON POST to /v1/audio/speech) can't
target it directly. Add a thin FastAPI adapter (stacks/zonos/adapter/):
POST /v1/audio/speech in front of the Zonos SDK, built FROM local/zonos
to reuse torch/CUDA/SDK. Returns a JSON envelope {audio, audio_format,
seed} — the seed rides back so asset-engine regenerate/fork can pin it
(Zonos is the fleet's first genuinely seedable TTS). compose gains a
zonos-api service on 8201; .env.example gains the port + voices dir.
This commit is contained in:
vh
2026-05-31 13:59:10 -07:00
parent 71df6f7474
commit 81efa8da96
7 changed files with 283 additions and 10 deletions
+29 -9
View File
@@ -13,7 +13,7 @@ Models ([`Zonos.from_pretrained()`](https://github.com/Zyphra/Zonos)):
| hybrid | [Zyphra/Zonos-v0.1-hybrid](https://huggingface.co/Zyphra/Zonos-v0.1-hybrid) | Mamba-SSM; needs Ampere+ GPU + extra build deps |
**Server:** irv-ml1 (Irvine, WireGuard-only)
**Port:** 8199 (container Gradio listens on 7860)
**Ports:** 8199 Gradio eval UI (container 7860) · 8201 REST adapter (container 8000)
**GPUs:** pins to device 0 (RTX 3090) by default; ~6 GB VRAM
**Image:** `local/zonos:v1` — built locally from a pinned git SHA of the
upstream repo via docker buildx's git URL context
@@ -29,17 +29,37 @@ natural-language paralinguistic tags. Worth A/B-ing by ear against
Chatterbox-Turbo (the research that prompted this stack explicitly said
"benchmark Zonos against Chatterbox before choosing").
## ⚠️ Audition surface, not skaldsong-pluggable (yet)
## Two surfaces: Gradio eval (8199) + REST adapter (8201)
The official repo ships a **Gradio WebUI + Python SDK only** — there is
**no OpenAI-compatible `/v1/audio/speech` endpoint**. So this stack is
for *auditioning quality*, not for wiring into skaldsong's engine
router as-is. To promote Zonos to a real engine slot we'd need either:
The official repo ships a **Gradio WebUI + Python SDK only** — no REST
endpoint. So this stack runs two services:
- the community FastAPI fork ([Zyphra/Zonos PR #73](https://github.com/Zyphra/Zonos/pull/73), adds REST + basic streaming), or
- a thin OpenAI-compat adapter in front of the Python SDK.
- **`zonos`** (8199) — upstream Gradio UI, for *auditioning quality by ear*.
- **`zonos-api`** (8201) — a thin OpenAI-ish adapter we built
(`adapter/server.py`) exposing `POST /v1/audio/speech` so **asset-engine**
can route to Zonos like every other TTS in the catalog. Returns a JSON
envelope `{audio: <base64>, audio_format, seed}` — the `seed` rides back
so regenerate/fork can pin it (Zonos is the fleet's first seedable TTS).
Both are a follow-up if Zonos earns a slot in the ear test.
The adapter is built `FROM local/zonos:<tag>` (reuses torch/CUDA/SDK) and
loads its own copy of the model (~6 GB on top of the Gradio service). Once
Zonos earns a permanent slot, drop the Gradio service and keep the adapter.
The community FastAPI fork ([PR #73](https://github.com/Zyphra/Zonos/pull/73))
was the alternative; we chose the self-owned adapter over pinning to an
unmerged fork.
### Adapter request (example)
```bash
curl -sS http://10.100.79.3:8201/v1/audio/speech \
-H 'content-type: application/json' \
-d '{"input":"Hello from Zonos.","language":"en-us","seed":420}' \
| python3 -c 'import sys,json,base64; d=json.load(sys.stdin); open("out.wav","wb").write(base64.b64decode(d["audio"])); print("seed",d["seed"])'
```
Cloning: drop a 10–30s WAV in `/worktank/zonos/voices/` and pass its
filename as `"voice"`. `GET /v1/audio/voices` lists what's available.
## Deploy