docs(zonos-gateway): capture the live :8890 gateway stack + emotion-preset system

Captures the production Zonos TTS path that the repo was blind to: stock ZONOS2
:1920 engine (stacks/zonos-engine) fronted by zonos-gateway:0.2.0 :8890, reached
via the LiteLLM ext-tts alias. Documents the emotion-preset system (neutral/warm/
excited/sad/intense/whisper, the simple preset: caller path), the API, and the
measured real-time cost (calibrated steering free at RTF~0.52, cfg1.5 ~0.625 —
still realtime). Flags the gateway source (~/zonos-gateway on irv-ml1) as not yet
in gitea. Marks stacks/zonos (v0.1 Gradio) dead/superseded. Records the emotion-
lever finding (text-priming flat -> native steering works) in persistent memory.
This commit is contained in:
vh
2026-07-17 15:56:11 -07:00
parent 196a0c1e4c
commit ba0ec64ac3
5 changed files with 157 additions and 0 deletions
+101
View File
@@ -0,0 +1,101 @@
# zonos-gateway — OpenAI-compatible TTS gateway over stock ZONOS2 (`:8890`, irv-ml1)
The production entry point for Zonos TTS. A thin FastAPI service that exposes a
clean OpenAI `POST /v1/audio/speech` surface and drives the **stock ZONOS2
engine's** correct native `POST /tts/generate` path — while adding the
**emotion-preset system** and voice defaults. Fronts `stacks/zonos-engine`
(the ZONOS2 engine at `:1920`); consumers reach it through a LiteLLM alias.
- **Server:** irv-ml1 (Irvine, WireGuard-only), 3090.
- **Deploy:** `/opt/docker/compose/zonos-gateway/compose.yaml` (fleet convention).
- **Image:** `zonos-gateway:0.2.0`, `build: .` from the source tree.
- **⚠️ Source:** lives at `~/zonos-gateway` on irv-ml1 (`src/zonos_gateway/`,
its own `pyproject.toml` + `ROADMAP.md` + `persistent-memory.md`) and is **NOT
yet in gitea** — a real version-control/backup gap. This eshpfi dir mirrors
the *stack* (compose + env + docs), not the source; the source needs its own
gitea repo (follow-up).
## Why it exists
ZONOS2 ships its own OpenAI `/v1/audio/speech`, but that path has two bugs:
1. **Repetition-penalty off by default** → short prompts pad **~70 s of silence**
instead of stopping at end-of-speech.
2. **The `voice` field is ignored** → a random speaker every call.
ZONOS2's **native `POST /tts/generate` is correct** (stock rep-penalty window
50 / penalty 1.2 / codebooks 8 stop the model at end-of-speech; speaker honored;
streams float32 PCM @ 44.1 kHz). This gateway drives that good native path and
leaves the vendored engine **100% stock** — no upstream patch to carry.
## Architecture — the `ext-tts` alias is the swap point
```
consumers ──▶ LiteLLM alias ext-tts ──▶ zonos-gateway ──▶ ZONOS2 /tts/generate (stock)
(generic, engine-neutral) (:8890, this stack) (:1920, stacks/zonos-engine)
```
Consumers **never** call this service by name — they call the generic LiteLLM
alias **`ext-tts`**, so the backing engine can be hot-swapped later with a
one-line alias repoint, invisible to every consumer. This repo stays
Zonos-specific (voice registry, correct native defaults, presets); the swap
point is the alias, not the code.
## Emotion control — the simple caller path (presets)
Emotion is exposed as **named presets** that bundle the engine's steering knobs
(`emotion_sliders` / `emotion_valence` / `emotion_arousal` / `emotion_strength`
/ `emotion_cfg_scale` + `speaking_rate`), defined in `src/zonos_gateway/dials.py`.
The simplest caller request:
```bash
curl -s http://<host>:8890/v1/audio/speech -H 'Content-Type: application/json' \
-d '{"input":"All the world'\''s a stage.","voice":"Cora","preset":"excited"}'
```
- **Presets (2026-07-17):** `neutral`, `warm`, `excited`, `sad`, `intense`, `whisper`.
- **Discovery:** `GET /v1/dials` self-describes voices / emotions
(`happy` `sad` `angry` `surprised`) / axes (`valence` `arousal`) / presets.
- **Full manual control** is also accepted on `/v1/audio/speech`
(`emotion_enabled`, `emotion_sliders`, `emotion_valence/arousal`,
`emotion_strength`, `emotion_cfg_scale`, `accurate_mode`, `speaking_rate*`,
`seed`, samplers) for fine-tuning or new presets.
### Real-time cost (measured on the 3090, 2026-07-17)
| Mode | RTF | Note |
|---|---|---|
| neutral (no steering) | ~0.52 | baseline |
| calibrated steering (`emotion_cfg_scale` 1.0) | ~0.52 | **free** — additive direction vectors, no extra forward pass |
| amplified (`emotion_cfg_scale` 1.5) | ~0.625 | ~+20% wall (CFG doubles the *decode* pass only), **still realtime** |
Emotion is effectively free at cfg 1.0 and cheap (~+20%, RTF still ~0.6) at
cfg 1.5–1.6 (where `warm`/`excited`/`intense` sit). No loss of real-time; the
only cost is ~20% less aggregate throughput under heavy concurrency at high cfg.
## API
| Method + path | Purpose |
|---|---|
| `POST /v1/audio/speech` | OpenAI-style TTS; accepts `preset` or manual emotion/prosody knobs |
| `GET /v1/dials` | Self-describe voices / emotions / axes / presets |
| `GET /v1/voices` | Available speaker voices |
| `GET /health` | Liveness |
## Relationship to the other Zonos stacks
- **`stacks/zonos-engine`** — the ZONOS2 engine (`:1920`) this gateway fronts.
- **`stacks/zonos`** — **DEAD / superseded.** That stack is Zonos **v0.1**
(Gradio UI + a v0.1 SDK adapter); it is no longer deployed. The live path is
ZONOS2 `:1920` (engine) + this gateway `:8890`.
## Emotion research (2026-07-17)
Text-based prosody-priming (prime → generate → excise) was **flat** on this
checkpoint. The **native emotion-steering subsystem** (this gateway's presets)
is the working lever — clearly moves delivery (sad → slow/quiet, excited →
fast/bright, etc.). `whisper` is a composed preset (no `whisper` *direction*
exists; the stock directions are happy/sad/angry/surprised + valence/arousal).
Escalation path for affects the 4 directions can't reach = custom directions
(`scripts/build_emotion_directions.py` in the engine repo) or emotional
reference-audio cloning (`speaker_audio_base64`).