feat(zonos): OpenAI-ish REST adapter for asset-engine routing
Upstream Zonos ships only Gradio + Python SDK — no REST surface — so
asset-engine (which routes a clean JSON POST to /v1/audio/speech) can't
target it directly. Add a thin FastAPI adapter (stacks/zonos/adapter/):
POST /v1/audio/speech in front of the Zonos SDK, built FROM local/zonos
to reuse torch/CUDA/SDK. Returns a JSON envelope {audio, audio_format,
seed} — the seed rides back so asset-engine regenerate/fork can pin it
(Zonos is the fleet's first genuinely seedable TTS). compose gains a
zonos-api service on 8201; .env.example gains the port + voices dir.
This commit is contained in:
+29
-9
@@ -13,7 +13,7 @@ Models ([`Zonos.from_pretrained()`](https://github.com/Zyphra/Zonos)):
|
||||
| hybrid | [Zyphra/Zonos-v0.1-hybrid](https://huggingface.co/Zyphra/Zonos-v0.1-hybrid) | Mamba-SSM; needs Ampere+ GPU + extra build deps |
|
||||
|
||||
**Server:** irv-ml1 (Irvine, WireGuard-only)
|
||||
**Port:** 8199 (container Gradio listens on 7860)
|
||||
**Ports:** 8199 Gradio eval UI (container 7860) · 8201 REST adapter (container 8000)
|
||||
**GPUs:** pins to device 0 (RTX 3090) by default; ~6 GB VRAM
|
||||
**Image:** `local/zonos:v1` — built locally from a pinned git SHA of the
|
||||
upstream repo via docker buildx's git URL context
|
||||
@@ -29,17 +29,37 @@ natural-language paralinguistic tags. Worth A/B-ing by ear against
|
||||
Chatterbox-Turbo (the research that prompted this stack explicitly said
|
||||
"benchmark Zonos against Chatterbox before choosing").
|
||||
|
||||
## ⚠️ Audition surface, not skaldsong-pluggable (yet)
|
||||
## Two surfaces: Gradio eval (8199) + REST adapter (8201)
|
||||
|
||||
The official repo ships a **Gradio WebUI + Python SDK only** — there is
|
||||
**no OpenAI-compatible `/v1/audio/speech` endpoint**. So this stack is
|
||||
for *auditioning quality*, not for wiring into skaldsong's engine
|
||||
router as-is. To promote Zonos to a real engine slot we'd need either:
|
||||
The official repo ships a **Gradio WebUI + Python SDK only** — no REST
|
||||
endpoint. So this stack runs two services:
|
||||
|
||||
- the community FastAPI fork ([Zyphra/Zonos PR #73](https://github.com/Zyphra/Zonos/pull/73), adds REST + basic streaming), or
|
||||
- a thin OpenAI-compat adapter in front of the Python SDK.
|
||||
- **`zonos`** (8199) — upstream Gradio UI, for *auditioning quality by ear*.
|
||||
- **`zonos-api`** (8201) — a thin OpenAI-ish adapter we built
|
||||
(`adapter/server.py`) exposing `POST /v1/audio/speech` so **asset-engine**
|
||||
can route to Zonos like every other TTS in the catalog. Returns a JSON
|
||||
envelope `{audio: <base64>, audio_format, seed}` — the `seed` rides back
|
||||
so regenerate/fork can pin it (Zonos is the fleet's first seedable TTS).
|
||||
|
||||
Both are a follow-up if Zonos earns a slot in the ear test.
|
||||
The adapter is built `FROM local/zonos:<tag>` (reuses torch/CUDA/SDK) and
|
||||
loads its own copy of the model (~6 GB on top of the Gradio service). Once
|
||||
Zonos earns a permanent slot, drop the Gradio service and keep the adapter.
|
||||
|
||||
The community FastAPI fork ([PR #73](https://github.com/Zyphra/Zonos/pull/73))
|
||||
was the alternative; we chose the self-owned adapter over pinning to an
|
||||
unmerged fork.
|
||||
|
||||
### Adapter request (example)
|
||||
|
||||
```bash
|
||||
curl -sS http://10.100.79.3:8201/v1/audio/speech \
|
||||
-H 'content-type: application/json' \
|
||||
-d '{"input":"Hello from Zonos.","language":"en-us","seed":420}' \
|
||||
| python3 -c 'import sys,json,base64; d=json.load(sys.stdin); open("out.wav","wb").write(base64.b64decode(d["audio"])); print("seed",d["seed"])'
|
||||
```
|
||||
|
||||
Cloning: drop a 10–30s WAV in `/worktank/zonos/voices/` and pass its
|
||||
filename as `"voice"`. `GET /v1/audio/voices` lists what's available.
|
||||
|
||||
## Deploy
|
||||
|
||||
|
||||
Reference in New Issue
Block a user