Files
esh-pfi-infrastructure/stacks/zonos-gateway/EMOTION-DIALS-SPEC.md
T
vh 4bdf01001c docs(zonos-gateway): sync emotion-presets spec + memory (0.2.1 bake)
Mirror the canonical EMOTION-DIALS-SPEC.md from vh/zonos-gateway (now carries
the provisional per-voice emotion presets baked as gateway 0.2.1) and capture
the axes-sweep → bake arc in persistent memory.
2026-07-18 01:37:27 -07:00

12 KiB
Raw Blame History

Zonos Emotion Control — Dials-First Spec

Status: canonical direction (operator ruling 2026-07-17) + empirically- measured per-voice emotion presets baked provisional (2026-07-18). Applies to: zonos-gateway (:8890) + its consumers (asset-engine, gateway-chat, any future client). Supersedes: the preset-centric usage pattern.

Two layers, not a contradiction. The philosophy is dials-first (§1): twist the raw dials per utterance. The 2026-07-18 presets (§5) are the measured answer for the emotions naive twisting can't hit — anger (whose named direction misfires) and surprise (dead as a class). They are empirically-tuned dial-sets exposed under a name, per voice; you can still twist your own.

1. Philosophy — twist the dials, don't pick a mood

Emotion is set per-utterance by twisting the raw dials, not by choosing from a fixed menu of named presets. The dials ARE the engine's canonical interface (Zyphra/ZONOS2 exposes exactly these; presets were a PFI-only convenience layer, not upstream). Rationale:

  • Expressiveness. A scene needs "wry, tired, a little fond" — not sad. An LLM client picks fine-grained affect per line far better than N fixed moods.
  • Explicit over implicit. A dial-set states the actual emotion in the call; a preset name hides it behind an opaque label we alone define.
  • No curation burden. Presets are a layer we'd have to keep tuning; they already drifted overwrought (excited/intense cfg 1.61.8 chipmunked with nobody re-tuning them). Dials have no drift and no maintenance surface.

The flavor presets (warm/excited/intense/whisper) stay demoted to optional copy-and-tweak examples (§6). The emotion presets in §5 are a different thing — an empirical result, not a hand-tuned convenience.

2. The dials (POST /v1/audio/speech)

All emotion dials require emotion_enabled: true — it is the master switch; without it every emotion_* field is inert.

Dial Range Default What it does Guidance
emotion_enabled bool false Master switch for emotion conditioning. Set true to use any dial below.
emotion_sliders {happy|sad|angry|surprised: w} null Per-emotion weights (the 4 named directions). Blend by setting several. Weights ~0.40.8. surprised raises pitch — the chipmunk driver; use sparingly or omit. Named angry is WEAK/misfires — prefer the axes preset (§5).
emotion_valence 1.0 … +1.0 0.0 Pleasant ↔ unpleasant axis. Continuous affect for states between the 4 named emotions. The lever behind §5.
emotion_arousal 1.0 … +1.0 0.0 Calm ↔ energized axis. + = excited/anxious, = tired/subdued. High arousal + surprised = chipmunk.
emotion_strength float 1.0 Multiplier on the per-speaker-calibrated direction. 1.0 = calibrated; >1 exaggerates. Keep ~0.71.1. Above ~1.3 strains most voices; NOT a smooth knob (§5 caveat 4).
emotion_cfg_scale ≥ 1.0 1.0 Classifier-free guidance on emotion. 1.0 = off. >1 amplifies. See §3 — the "deaf by 1.5" dial.
accurate_mode bool true true = faithful to the reference voice; false = looser / more expressive. false is REQUIRED for emotion to land (accurate mode suppresses it); the cost is some identity drift.
speaking_rate_enabled bool false Master switch for rate conditioning. Needed for speaking_rate / speaking_rate_bucket / speed to bite.
speaking_rate_bucket int 07 null Exact rate bucket (0=slowest … 7=fastest; wps bands from /tts/capabilities). Fast = excited/anxious; slow = sad/tired. Pair with the emotion, don't overdrive.

Discovery is live: GET /v1/dials self-describes emotions, axes, rate buckets, per-dial ranges, and the per-voice emotion presets (the voice_emotion_presets block).

3. emotion_cfg_scale — the knob goes to 30, but you're deaf by 1.5

There is deliberately NO enforced cap. The dial will take large values — set it to 30 if you want. You will regret it. Emotional realism tops out around 1.5; past ~1.4 the voice is overwrought and distorted, and by 1.5 you're "deaf" — pushing harder only adds artifact, not feeling.

  • Usable range: 1.01.5. Recommended ceiling 1.4. Never above 1.5.
  • 1.0 = off: the additive emotion directions alone (already calibrated). Plenty expressive for most lines.
  • ~1.251.4 = the sweet spot when you want a line to land harder.
  • > 1.5 = do not. It's the amp that goes to 11 — the number is bigger, the sound is worse.

The §5 emotion presets use cfg 1.5 deliberately — expressive mode needs the push for the axes emotion to land; it's the top of the usable band, not past it.

4. Cost / real-time — a second reason to stay ≤ 1.5

Measured on the 3090 (irv-ml1):

emotion_cfg_scale RTF Note
1.0 (off) ~0.52 free — same as neutral; additive vectors, no extra forward pass
1.5 ~0.625 ~+20% wall (CFG doubles the decode pass), still comfortably real-time
> 1.5 worse more compute AND worse sound — strictly dominated

5. Empirically-measured per-voice emotion presets (PROVISIONAL, 2026-07-18)

An empirical sweep (valence×arousal grid, scored by an emotion classifier + speaker-identity retention) established that the four named emotions do not respond uniformly to naive dial-twisting:

  • happy / sad — work via the named sliders (already usable).
  • angry — the named angry slider is weak-to-broken (on BritishFemale it misfires as fear). Driving the valence/arousal axes instead (negative valence + high arousal) rescues it.
  • surpriseddead as an emotion class on this engine (never activates, any dial). The usable stand-in is startled-happy (high arousal + positive valence), which reads as bright surprise.

The winning cell is different per voice, so a single global preset is unsafe. The gateway therefore exposes emotion presets that resolve against the voice:

POST /v1/audio/speech  { "input": "...", "voice": "BritishFemale", "preset": "angry" }

Preset names

preset resolves to aliases
angry axes anger cell for the voice
happy axes happy cell (better identity than the named happy slider)
startled_happy the "surprised" product stand-in surprised, startled
sad named-slider preset (unchanged; not axes-tuned yet)
neutral/warm/excited/intense/whisper voice-independent flavor presets (§6)

Calibrated cells (the 3 default voices)

All expressive (accurate_mode:false), emotion_cfg_scale:1.5, pure-axes (no sliders). emo = emotion-classifier target prob, id = speaker-identity retention (floor 0.65; neutral ~0.85).

Voice preset valence arousal strength emo id
AmericanFemale angry 0.4 +1.0 1.0 0.53 0.69
AmericanFemale happy / startled_happy +0.6 +0.8 1.0 1.0 0.80
AmericanMale angry (drama) 0.6 +0.8 1.2 1.0 0.62
AmericanMale happy / startled_happy +0.3 +1.0 1.0 0.98 0.76
BritishFemale angry 0.4 +0.8 1.0 0.99 0.73
BritishFemale happy / startled_happy +0.6 +1.0 1.0 1.0 0.74

Uncalibrated voices (Cora + cloned characters) fall back to a mid-region default (angry = v0.5/a+0.9/str1.1; happy/startled_happy = v+0.5/a+0.9) until they earn a measured row. The live table is in GET /v1/dialsvoice_emotion_presets.

Caveats (why "provisional")

  1. Non-monotonic surface — the exact (valence, arousal, strength) is pinned; do NOT interpolate a preset's neighborhood in a UI.
  2. Anger costs identity — retention ~0.620.73 vs ~0.85 neutral. Acceptable for a drama beat, not for identity-critical dialogue.
  3. Surprised-as-class is dead — never expose a "surprised" slider promise; map user-intent surprised/startled/shocked to the startled_happy preset.
  4. emotion_strength is not a smooth knob (esp. AmericanMale: 1.0 = mild, 1.2 = full anger, higher flips to disgust). The angry preset pins the voice's tuned strength; send an explicit emotion_strength to shift tier (e.g. 1.0 on AmericanMale for a softer anger).
  5. Two-regime policy — identity-critical lines: accurate_mode:true, emotion off or soft. Tagged drama beats: these presets. Don't stack a high named slider AND a high axes preset (single-condition wins).

Baked provisional pending confident ear-validation on emotion-congruent text (neutral-sentence auditing was inconclusive — the dials clearly differ, but a flat line understates landing). Re-tune when we revisit.

6. Starting-point flavor dial-sets (examples — copy and tweak)

Tuned by ear on BritishFemale. Departure points, not a menu. (Distinct from §5: these are hand-tuned flavors, not measured emotion cells.)

// warm  — friendly, unhurried
{ "emotion_enabled": true, "emotion_sliders": {"happy": 0.4},
  "emotion_valence": 0.4, "emotion_strength": 0.8, "emotion_cfg_scale": 1.3 }

// sad — slow, quiet
{ "emotion_enabled": true, "emotion_sliders": {"sad": 0.8},
  "emotion_valence": -0.6, "emotion_arousal": -0.5, "emotion_strength": 1.0,
  "emotion_cfg_scale": 1.5, "speaking_rate_enabled": true, "speaking_rate_bucket": 1 }

// whisper — hushed, breathy
{ "emotion_enabled": true, "emotion_sliders": {"sad": 0.2},
  "emotion_arousal": -0.7, "emotion_strength": 0.9, "emotion_cfg_scale": 1.4,
  "speaking_rate_enabled": true, "speaking_rate_bucket": 2 }

7. LLM client system-prompt snippet

Paste into an LLM client that drives TTS, so it emits a dial-set per utterance:

You control speech emotion per line. For anger or surprise/startle, use a preset: {"preset": "angry"} or {"preset": "startled_happy"} (the named angry/surprised sliders misfire — don't use them). For everything else, emit a dial-set: set emotion_enabled: true, then shape delivery with emotion_sliders (happy/sad ~00.8), emotion_valence/emotion_arousal (1..+1), emotion_strength (~0.71.1), and emotion_cfg_scale (1.01.4, never above 1.5). Add speaking_rate_enabled: true + speaking_rate_bucket (0 slow … 7 fast) for pacing. Subtle by default, bold only when the moment earns it.

8. Voice cloning (reference audio)

Cloning is text-independent (Qwen3 speaker embedding) — no transcript needed. Give it ~1524s of clean, single-speaker audio (Zyphra's own default voices run 8.723.9s). Shorter clips (≤5s) under-condition the embedding and audibly degrade, worse under emotion steering. A cloned voice has no calibrated emotion row yet — it uses the §5 fallback; add a measured row (or map it to the nearest default) when its emotions matter.

9. Follow-ups (not done here)

  • Ear-validation on congruent text — the §5 cells were baked provisional off an inconclusive neutral-text audition; confirm on emotion-appropriate lines and promote from provisional (or re-tune).
  • Sad axes pass — sad was not in the axes sweep; still the named-slider preset. Owes a per-voice text/strength pass on the 3 defaults.
  • AmericanMale angry sweet-spot — no single cell clears both the emo and identity floors; it ships as drama (str 1.2) with a documented soft override (str 1.0). A finer strength ladder could find a middle tier.
  • Clone-character rows — Emmie/Penny/Natalie/Miranda use the fallback; run a 1-row confirm at the baked cell per voice, or map to the nearest default.
  • CI / deploy wiring — the deployed irv-ml1 tree is a hand-updated build context (not git / not CI-deployed). Git-connect + build-on-push with a deploy key, like the other sister services.