Files
esh-pfi-infrastructure/stacks/zonos/README.md
T
vh 81efa8da96 feat(zonos): OpenAI-ish REST adapter for asset-engine routing
Upstream Zonos ships only Gradio + Python SDK — no REST surface — so
asset-engine (which routes a clean JSON POST to /v1/audio/speech) can't
target it directly. Add a thin FastAPI adapter (stacks/zonos/adapter/):
POST /v1/audio/speech in front of the Zonos SDK, built FROM local/zonos
to reuse torch/CUDA/SDK. Returns a JSON envelope {audio, audio_format,
seed} — the seed rides back so asset-engine regenerate/fork can pin it
(Zonos is the fleet's first genuinely seedable TTS). compose gains a
zonos-api service on 8201; .env.example gains the port + voices dir.
2026-05-31 13:59:10 -07:00

3.8 KiB
Raw Blame History

Zonos

Zyphra's expressive multilingual open-weight TTS (Zyphra/Zonos) — 44 kHz output, zero-shot voice cloning, and explicit emotion/conditioning controls — served via the official repo's Gradio interface.

Models (Zonos.from_pretrained()):

Variant HF repo Notes
transformer Zyphra/Zonos-v0.1-transformer ~3.6 GB, no special kernels (default)
hybrid Zyphra/Zonos-v0.1-hybrid Mamba-SSM; needs Ampere+ GPU + extra build deps

Server: irv-ml1 (Irvine, WireGuard-only) Ports: 8199 Gradio eval UI (container 7860) · 8201 REST adapter (container 8000) GPUs: pins to device 0 (RTX 3090) by default; ~6 GB VRAM Image: local/zonos:v1 — built locally from a pinned git SHA of the upstream repo via docker buildx's git URL context Upstream: Zyphra/Zonos (Apache-2.0), trained on 200k+ hours; multilingual EN/JA/ZH/FR/DE

Why this stack exists

A different control surface from the rest of the bench: Zonos exposes emotion/conditioning sliders plus pitch/rate conditioning and 44 kHz output, where Chatterbox/Dia use inline tags and Fish uses natural-language paralinguistic tags. Worth A/B-ing by ear against Chatterbox-Turbo (the research that prompted this stack explicitly said "benchmark Zonos against Chatterbox before choosing").

Two surfaces: Gradio eval (8199) + REST adapter (8201)

The official repo ships a Gradio WebUI + Python SDK only — no REST endpoint. So this stack runs two services:

  • zonos (8199) — upstream Gradio UI, for auditioning quality by ear.
  • zonos-api (8201) — a thin OpenAI-ish adapter we built (adapter/server.py) exposing POST /v1/audio/speech so asset-engine can route to Zonos like every other TTS in the catalog. Returns a JSON envelope {audio: <base64>, audio_format, seed} — the seed rides back so regenerate/fork can pin it (Zonos is the fleet's first seedable TTS).

The adapter is built FROM local/zonos:<tag> (reuses torch/CUDA/SDK) and loads its own copy of the model (~6 GB on top of the Gradio service). Once Zonos earns a permanent slot, drop the Gradio service and keep the adapter.

The community FastAPI fork (PR #73) was the alternative; we chose the self-owned adapter over pinning to an unmerged fork.

Adapter request (example)

curl -sS http://10.100.79.3:8201/v1/audio/speech \
  -H 'content-type: application/json' \
  -d '{"input":"Hello from Zonos.","language":"en-us","seed":420}' \
  | python3 -c 'import sys,json,base64; d=json.load(sys.stdin); open("out.wav","wb").write(base64.b64decode(d["audio"])); print("seed",d["seed"])'

Cloning: drop a 1030s WAV in /worktank/zonos/voices/ and pass its filename as "voice". GET /v1/audio/voices lists what's available.

Deploy

# from this workstation (irv-ml1 is WG-only — routes via ana-wg):
scripts/deploy-stack.sh irv-ml1 zonos
# then on irv-ml1, first run builds the image from the pinned SHA:
#   docker compose up -d --build

Open http://10.100.79.3:8199/ for the Gradio UI. First boot pulls the model (~3.6 GB) into ZONOS_CACHE_DIR; the 600 s start_period covers it.

Notes

  • Pin ZONOS_SHA to a full 40-char commit before relying on this — .env.example ships main, which is not reproducible.
  • The upstream image bundles espeak-ng (required for Zonos's phonemization) via its Dockerfile.
  • GRADIO_SERVER_NAME=0.0.0.0 is set so the in-container Gradio binds all interfaces and the host port-map reaches it. If a future upstream hardcodes server_name, the API fork above is the cleaner path.