docs(zonos-gateway): sync emotion-presets spec + memory (0.2.1 bake)

Mirror the canonical EMOTION-DIALS-SPEC.md from vh/zonos-gateway (now carries
the provisional per-voice emotion presets baked as gateway 0.2.1) and capture
the axes-sweep → bake arc in persistent memory.
This commit is contained in:
2026-07-18 01:37:27 -07:00
parent 89611eb06b
commit 4bdf01001c
2 changed files with 134 additions and 57 deletions
+12 -4
View File
@@ -117,10 +117,16 @@ _As of 2026-07-18 — active work = Zonos2 emotion-tuning + voice-cloning for th
- **Emotion CANONICAL established** (empirical sweep, `[2026-07-18]` entry): single-emotion only; two-regime accurate(identity)/expressive(drama) policy; happy/sad usable, angry weak, surprised dead on the named directions. - **Emotion CANONICAL established** (empirical sweep, `[2026-07-18]` entry): single-emotion only; two-regime accurate(identity)/expressive(drama) policy; happy/sad usable, angry weak, surprised dead on the named directions.
**Open loops for the fresh session:** **Open loops for the fresh session:**
- **yt-voice-clipper yields test** — job `f3ff746dbae9494d` running on irv-ml1 (submitted to verify v0.3.3's `max_gap` fix). Check `ssh irv-ml1 'curl -s :8000/jobs/f3ff746dbae9494d'` (or /diagnostics) for ~14 segments (not 2), then **reply to yt-voice-clipper-dev** (thread `01KXT0T6GYHB`) confirming verify #3. (Redeploy #1 A6000 + #2 version-0.3.3 already confirmed in `01KXT1BFTA5J`.) - **yt-voice-clipper yields test CLOSED** (2026-07-18) — job `f3ff746dbae9494d` returned **13 segments / 203.5s** (was 2 pre-fix), `max_gap=1.2` confirmed live; replied verify #3 to yt-voice-clipper-dev (thread `01KXT0T6GYHB61EKJ81PZAMA7W`, msg `01KXT25D…`). All 3 verifies green (A6000 / v0.3.3 / yields). Redeploy done + validated.
- **dvalin owes-me / I-owe-dvalin the axes-sweep numbers** (thread `01KXT12FN0AS5A3WMKEK06BVPS`) — I committed to run the axes sweep and send results. - **Axes sweep DONE + dvalin loop CLOSED** (2026-07-18) — ran the valence×arousal sweep on the 3 calibrated defaults, sent dvalin the graded numbers (thread `01KXT12FN0AS5A3WMKEK06BVPS`, msg `01KXT2ZB8G…`). Result in the `[2026-07-18] axes sweep` Recent-decisions entry below. **Open Qs posed to dvalin** (awaiting reply): (a) strength-ladder at ship cells to find AmericanMale's missing angry sweet spot? (b) emotion-congruent text next, or bake happy/sad/angry canonical first? — monitor armed for the reply.
- **NEXT experiment (operator to green-light): axes sweep** for angry/surprised — valence/arousal grid (angry ≈ val/+aro; surprised ≈ +aro), the only path to rescue the two broken named emotions; then strength ladder + emotion-congruent text. Reuse `~/development/zonos-tools/emotion_sweep.py` (scoring env: `uv run --with resemblyzer --with funasr --with "numpy<2" --with soundfile --with requests --with "setuptools<80" --with torchaudio`). Then bake happy/sad canonical into gateway presets. - **BAKED — emotion presets shipped as zonos-gateway 0.2.1** (2026-07-18). Operator ear-checked (neutral-text audition inconclusive: "they all sound different, hard to tell" → "call it, provisional, we'll deal with it"), then green-lit bake+docs. Implemented **voice-resolved** emotion presets (`resolve_preset(name, voice)`): `angry`/`happy`/`startled_happy` (+`surprised`/`startled` aliases) resolve to the per-voice measured cell — NOT a global preset (dvalin ruling; BrF named-angry→fear is why). Calibrated for AmF/AmM/BrF; mid-region fallback for Cora+clones. `/v1/dials` exposes `voice_emotion_presets`; `/docs` FastAPI description documents it (the web surface); durable spec at `docs/EMOTION-DIALS-SPEC.md` (moved into vh/zonos-gateway repo — was mirror-only) + README table. 44 tests green; rebuilt + verified live on irv-ml1 (BrF angry→200, Cora surprised→fallback 200). Commit `a92cc40` + tag `v0.2.1` on vh/zonos-gateway clone `~/development/zonos-gateway` (UNPUSHED). Audition tooling: `~/development/zonos-tools/{gen_auditions,axes_sweep,strength_ladder}.py`; audition server was nohup :8901 on nh3-dev (throwaway, tear down). **Ship cells baked (per-voice, exp/cfg1.5, no sliders):**
- **Re-arm the althing monitor** (`/althing:monitor`, handle `infra-ops`) — wake-listener dies on /clear; "read mail after" per operator. Open watch: worldtree-dev (#368 memory-leak, my read-only forensics done; #363 research-wing ingest PARKED). - **angry** — BritishFemale v-0.4/a+0.8 str1.0 (0.99/id0.725, proxy fear-CLEAN); AmericanFemale v-0.4/a+1.0 str1.0 (0.53/id0.685); AmericanMale **two-tier** (no single ship cell) → soft v-0.6/a+0.8 str1.0 (0.23/id0.654) / drama v-0.6/a+0.8 str1.2 (1.0/id0.616, clean).
- **startled-happy** (="surprised" product alias; never set surprised slider) — AmF v+0.6/a+0.8, AmM v+0.3/a+1.0, BrF v+0.6/a+1.0 (all happy~1.0, id 0.74-0.80).
- **happy** — axes-preferred (same cells as startled-happy; +0.15 id over named-happy); named-happy = fallback. **sad** = unchanged (named exp cfg1.5, not re-tested on these 3).
- dvalin ruling: store as **per-voice dial tables, NOT a global preset** (BrF named-angry→fear is the exhibit). Bake provisional + ear-spot-check 1 clip/cell.
- Follow-ups: emotion-congruent-text pass (validates intensity, may rescue sad id); clone-char voices map to nearest-default table or 1-row confirm. Tools: `~/development/zonos-tools/{axes_sweep,strength_ladder}.py` (scoring env: `uv run --with resemblyzer --with funasr --with "numpy<2" --with soundfile --with requests --with "setuptools<80" --with torchaudio`; run ON irv-ml1 — localhost:8890 + REFDIR).
- **soong-lab containerize cutover on corviduo-dev — AUTHORIZED/slotted** (2026-07-18, operator "slot the soong"; soong-dev thread `01KXT3A6C3908TA4V9THV3AMH7`). In-repo half landed (soong-lab `0e49391`: Dockerfile/compose/`.gitea/workflows/build-and-push.yml`/`docs/DEPLOY.md` = source-of-truth checklist). My side to drive: (1) creds — internal Gitea PyPI read (`bifrost[reference-server]==1.1.0`, BuildKit secret `id=gitea_token`) + container-registry push; **source via claude-bot service account first** (migrate-off-personal-creds), vh fallback only on scope gap; (2) confirm a Gitea Actions runner for the tag/manual build; (3) propose a quiet cutover WINDOW to vh. HARD: migrate `/data/library` (incl. Sindra) + `/data/portraits` into volumes BEFORE first container run; verify Bifrost callback path (`SOONG_LAB_BIFROST_ENDPOINT_URL`, proxy→container + WT host-allowlist) + `/api/version` + one live Soong tool-call round-trip BEFORE retiring `soong-lab-studio.service` + the webhook git-pull deploy. Not started (queued behind Zonos ear-check). ⚠️ corviduo-dev = WT-team-managed, data-affecting → coordinate. [[reference_corviduo_dev_emergency_ops]] [[reference_claude_bot_gitea_creds]]
- **althing monitor ARMED this session** (handle `infra-ops`, wake-listener; herald up). Re-arm after each fire (dies on /clear). Open watch: worldtree-dev (#368 memory-leak forensics done; #363 research-wing ingest PARKED).
- **eshpfi commits UNPUSHED** — the zonos-gateway capture + memory (`14a0004`..`438cd35`) are committed; push is the operator's call. `stacks/heretic2-charrp-reasoning/` still UNTRACKED; `graphify-out/GRAPH_REPORT.md` modified. - **eshpfi commits UNPUSHED** — the zonos-gateway capture + memory (`14a0004`..`438cd35`) are committed; push is the operator's call. `stacks/heretic2-charrp-reasoning/` still UNTRACKED; `graphify-out/GRAPH_REPORT.md` modified.
- **irv-ml1 3090 oversubscription** (kokoro `:8193` + vibevoicefusion `:9527` idle-pinned; zonos engine :1920 also on 3090) — carried; operator declined to fix. - **irv-ml1 3090 oversubscription** (kokoro `:8193` + vibevoicefusion `:9527` idle-pinned; zonos engine :1920 also on 3090) — carried; operator declined to fix.
@@ -130,6 +136,8 @@ _As of 2026-07-18 — active work = Zonos2 emotion-tuning + voice-cloning for th
## Recent decisions ## Recent decisions
- `[2026-07-18]` **Axes sweep RESCUED angry; surprised-class dead but startled-happy ships.** Valence×arousal grid on the 3 calibrated defaults (AmericanFemale/Male, BritishFemale), exp/cfg1.5/strength1.0, 84 clips, emotion2vec + resemblyzer scored, graded vs dvalin's floor. **ANGRY rescued** (named direction was 0.0040.15, British named-angry even misfired as fear 0.89): axes ship cells at **negative valence (0.4..0.8) + high arousal (+0.8..+1.0)** — BritishFemale v-0.4/a+0.8 angry=0.99/id0.725 SHIP, AmericanFemale v-0.4/a+1.0 angry=0.53/id0.685 SHIP; AmericanMale two-tier post-ladder (no single ship cell — best drama = v-0.6/a+0.8 str1.2 angry=1.0/id0.616 clean, soft = same cell str1.0 angry0.23/id0.654; cell A v-0.6/a+1.0 is a non-monotonic minefield, skip). BrF ship cell proxy-CLEAN of fear (str<1.0 just kills anger). **SURPRISED-class DEAD** (max 0.047 across all 84 cells) but **startled-happy** (happy-proxy) ships all 3 at high arousal + neutral/positive valence, with a **+0.170.20 identity LIFT** over the named-surprised route (named hits happy~1.0 but at id0.570.61, under floor; axes hits happy~1.0 at id0.740.80). Bonus: axes-happy retains ~0.100.15 more identity than the named happy slider too. Caveats: response surface non-monotonic/sharp-thresholded; angry region borders fear/disgust (bleed); emotion2vec saturates at 1.0 (needs ear-confirm); neutral text understates. Tooling `~/development/zonos-tools/axes_sweep.py`; per-clip JSON was `irv-ml1:/tmp/axes_sweep_results.json` (ephemeral). Sent dvalin msg `01KXT2ZB8G…`. NEXT = operator ear-confirm → bake presets. [[reference_zonos_tts_stack]]
- `[2026-07-18]` **Zonos2 emotion CANONICAL from an empirical sweep + the voice-cloning pipeline** — 4 chars cloned (Emmie/Penny/Natalie/Miranda), host-managed gateway voices, two-regime accurate/expressive policy, happy/sad usable + angry-weak/surprised-dead on named directions, dvalin-synthesized; axes sweep is the NEXT experiment. Studio + sweep tooling at `~/development/zonos-tools/`. → `persistent-memory.d/2026-07-18-zonos-emotion-canonical.md` - `[2026-07-18]` **Zonos2 emotion CANONICAL from an empirical sweep + the voice-cloning pipeline** — 4 chars cloned (Emmie/Penny/Natalie/Miranda), host-managed gateway voices, two-regime accurate/expressive policy, happy/sad usable + angry-weak/surprised-dead on named directions, dvalin-synthesized; axes sweep is the NEXT experiment. Studio + sweep tooling at `~/development/zonos-tools/`. → `persistent-memory.d/2026-07-18-zonos-emotion-canonical.md`
- `[2026-07-18]` **yt-voice-clipper: A6000-pin fix + v0.3.3 redeploy.** Fixed a latent misconfig — the host override *said* "pin worker to A6000" but `NVIDIA_VISIBLE_DEVICES` was `"0"` (the 3090); re-pinned worker+api to the A6000 by UUID (`GPU-9672f0d5`, 3090 is zonos2's). Then redeployed api+worker to v0.3.3 (`docker compose up -d --build`; SPA+Python; `max_gap` 0.6→1.2s; stderr surfaced in job.log). A6000 + version verified; yields test in-flight (job `f3ff746dbae9494d`). yt-voice-clipper-dev thread `01KXT0T6GYHB`. [[reference_ytvc_autodeploy]] - `[2026-07-18]` **yt-voice-clipper: A6000-pin fix + v0.3.3 redeploy.** Fixed a latent misconfig — the host override *said* "pin worker to A6000" but `NVIDIA_VISIBLE_DEVICES` was `"0"` (the 3090); re-pinned worker+api to the A6000 by UUID (`GPU-9672f0d5`, 3090 is zonos2's). Then redeployed api+worker to v0.3.3 (`docker compose up -d --build`; SPA+Python; `max_gap` 0.6→1.2s; stderr surfaced in job.log). A6000 + version verified; yields test in-flight (job `f3ff746dbae9494d`). yt-voice-clipper-dev thread `01KXT0T6GYHB`. [[reference_ytvc_autodeploy]]
+122 -53
View File
@@ -1,9 +1,16 @@
# Zonos Emotion Control — Dials-First Spec # Zonos Emotion Control — Dials-First Spec
**Status:** canonical direction (operator ruling 2026-07-17). **Applies to:** **Status:** canonical direction (operator ruling 2026-07-17) + **empirically-
`zonos-gateway` (`:8890`) + its consumers (asset-engine, gateway-chat, any measured per-voice emotion presets baked provisional (2026-07-18).** **Applies
to:** `zonos-gateway` (`:8890`) + its consumers (asset-engine, gateway-chat, any
future client). **Supersedes:** the preset-centric usage pattern. future client). **Supersedes:** the preset-centric usage pattern.
> **Two layers, not a contradiction.** The philosophy is dials-first (§1): twist
> the raw dials per utterance. The 2026-07-18 presets (§5) are the *measured
> answer* for the emotions naive twisting can't hit — anger (whose named
> direction misfires) and surprise (dead as a class). They are empirically-tuned
> dial-sets exposed under a name, per voice; you can still twist your own.
## 1. Philosophy — twist the dials, don't pick a mood ## 1. Philosophy — twist the dials, don't pick a mood
Emotion is set **per-utterance by twisting the raw dials**, not by choosing from Emotion is set **per-utterance by twisting the raw dials**, not by choosing from
@@ -19,8 +26,9 @@ not upstream). Rationale:
already drifted overwrought (`excited`/`intense` cfg 1.61.8 chipmunked with already drifted overwrought (`excited`/`intense` cfg 1.61.8 chipmunked with
nobody re-tuning them). Dials have no drift and no maintenance surface. nobody re-tuning them). Dials have no drift and no maintenance surface.
**Presets are demoted to optional copy-and-tweak examples** (see §5), not the **The flavor presets (`warm`/`excited`/`intense`/`whisper`) stay demoted to
primary interface and not a maintained product. optional copy-and-tweak examples** (§6). The *emotion* presets in §5 are a
different thing — an empirical result, not a hand-tuned convenience.
## 2. The dials (`POST /v1/audio/speech`) ## 2. The dials (`POST /v1/audio/speech`)
@@ -30,17 +38,18 @@ without it every `emotion_*` field is inert.
| Dial | Range | Default | What it does | Guidance | | Dial | Range | Default | What it does | Guidance |
|---|---|---|---|---| |---|---|---|---|---|
| `emotion_enabled` | bool | `false` | Master switch for emotion conditioning. | Set `true` to use any dial below. | | `emotion_enabled` | bool | `false` | Master switch for emotion conditioning. | Set `true` to use any dial below. |
| `emotion_sliders` | `{happy\|sad\|angry\|surprised: w}` | `null` | Per-emotion weights (the 4 named directions). Blend by setting several. | Weights ~`0.40.8`. **`surprised` raises pitch — the chipmunk driver; use sparingly or omit.** | | `emotion_sliders` | `{happy\|sad\|angry\|surprised: w}` | `null` | Per-emotion weights (the 4 named directions). Blend by setting several. | Weights ~`0.40.8`. **`surprised` raises pitch — the chipmunk driver; use sparingly or omit.** Named `angry` is WEAK/misfires — prefer the axes preset (§5). |
| `emotion_valence` | `1.0 … +1.0` | `0.0` | Pleasant ↔ unpleasant axis. | Continuous affect for states between the 4 named emotions. | | `emotion_valence` | `1.0 … +1.0` | `0.0` | Pleasant ↔ unpleasant axis. | Continuous affect for states between the 4 named emotions. The lever behind §5. |
| `emotion_arousal` | `1.0 … +1.0` | `0.0` | Calm ↔ energized axis. | `+` = excited/anxious, `` = tired/subdued. High arousal + `surprised` = chipmunk. | | `emotion_arousal` | `1.0 … +1.0` | `0.0` | Calm ↔ energized axis. | `+` = excited/anxious, `` = tired/subdued. High arousal + `surprised` = chipmunk. |
| `emotion_strength` | float | `1.0` | Multiplier on the **per-speaker-calibrated** direction. `1.0` = calibrated; `>1` exaggerates. | Keep `~0.71.1`. Above `~1.3` strains the voice. | | `emotion_strength` | float | `1.0` | Multiplier on the **per-speaker-calibrated** direction. `1.0` = calibrated; `>1` exaggerates. | Keep `~0.71.1`. Above `~1.3` strains most voices; NOT a smooth knob (§5 caveat 4). |
| `emotion_cfg_scale` | `≥ 1.0` | `1.0` | Classifier-free guidance on emotion. `1.0` = off. `>1` amplifies. | **See §3 — the "deaf by 1.5" dial.** | | `emotion_cfg_scale` | `≥ 1.0` | `1.0` | Classifier-free guidance on emotion. `1.0` = off. `>1` amplifies. | **See §3 — the "deaf by 1.5" dial.** |
| `accurate_mode` | bool | `true` | `true` = faithful to the reference voice; `false` = looser / more expressive. | Leave `true` for identity-preserving; flip `false` only when you want the voice to bend. | | `accurate_mode` | bool | `true` | `true` = faithful to the reference voice; `false` = looser / more expressive. | `false` is REQUIRED for emotion to land (accurate mode suppresses it); the cost is some identity drift. |
| `speaking_rate_enabled` | bool | `false` | Master switch for rate conditioning. | Needed for `speaking_rate` / `speaking_rate_bucket` / `speed` to bite. | | `speaking_rate_enabled` | bool | `false` | Master switch for rate conditioning. | Needed for `speaking_rate` / `speaking_rate_bucket` / `speed` to bite. |
| `speaking_rate_bucket` | int `07` | `null` | Exact rate bucket (`0`=slowest … `7`=fastest; wps bands from `/tts/capabilities`). | Fast = excited/anxious; slow = sad/tired. Pair with the emotion, don't overdrive. | | `speaking_rate_bucket` | int `07` | `null` | Exact rate bucket (`0`=slowest … `7`=fastest; wps bands from `/tts/capabilities`). | Fast = excited/anxious; slow = sad/tired. Pair with the emotion, don't overdrive. |
Discovery is live: `GET /v1/dials` self-describes emotions, axes, rate buckets, Discovery is live: `GET /v1/dials` self-describes emotions, axes, rate buckets,
and per-dial ranges. per-dial ranges, **and the per-voice emotion presets** (the `voice_emotion_presets`
block).
## 3. `emotion_cfg_scale` — the knob goes to 30, but you're deaf by 1.5 ## 3. `emotion_cfg_scale` — the knob goes to 30, but you're deaf by 1.5
@@ -56,8 +65,8 @@ it to 30 if you want. You will regret it. Emotional realism tops out around
- `> 1.5` = do not. It's the amp that goes to 11 — the number is bigger, the - `> 1.5` = do not. It's the amp that goes to 11 — the number is bigger, the
sound is worse. sound is worse.
This is documented, not clamped, on purpose (explicit over implicit): the client The §5 emotion presets use cfg `1.5` deliberately — expressive mode needs the
owns the choice; the spec owns the warning. push for the axes emotion to land; it's the top of the usable band, not past it.
## 4. Cost / real-time — a second reason to stay ≤ 1.5 ## 4. Cost / real-time — a second reason to stay ≤ 1.5
@@ -69,31 +78,91 @@ Measured on the 3090 (irv-ml1):
| `1.5` | ~0.625 | ~+20% wall (CFG doubles the *decode* pass), still comfortably real-time | | `1.5` | ~0.625 | ~+20% wall (CFG doubles the *decode* pass), still comfortably real-time |
| `> 1.5` | worse | more compute AND worse sound — strictly dominated | | `> 1.5` | worse | more compute AND worse sound — strictly dominated |
So cfg above 1.5 costs more *and* sounds worse. Stay ≤ 1.4. ## 5. Empirically-measured per-voice emotion presets (PROVISIONAL, 2026-07-18)
## 5. Starting-point dial-sets (examples — copy and tweak, not presets) An empirical sweep (valence×arousal grid, scored by an emotion classifier +
speaker-identity retention) established that the four named emotions do **not**
respond uniformly to naive dial-twisting:
Tuned by ear on BritishFemale. Treat these as departure points, not a menu. - **happy / sad** — work via the named sliders (already usable).
- **angry** — the *named* `angry` slider is weak-to-broken (on BritishFemale it
misfires as **fear**). Driving the **valence/arousal axes** instead (negative
valence + high arousal) rescues it.
- **surprised** — **dead as an emotion class** on this engine (never activates,
any dial). The usable stand-in is **startled-happy** (high arousal + positive
valence), which reads as bright surprise.
The winning cell is **different per voice**, so a single global preset is unsafe.
The gateway therefore exposes emotion presets that **resolve against the voice**:
```
POST /v1/audio/speech { "input": "...", "voice": "BritishFemale", "preset": "angry" }
```
### Preset names
| preset | resolves to | aliases |
|---|---|---|
| `angry` | axes anger cell for the voice | — |
| `happy` | axes happy cell (better identity than the named happy slider) | — |
| `startled_happy` | the "surprised" product stand-in | `surprised`, `startled` |
| `sad` | named-slider preset (unchanged; not axes-tuned yet) | — |
| `neutral`/`warm`/`excited`/`intense`/`whisper` | voice-independent flavor presets (§6) | — |
### Calibrated cells (the 3 default voices)
All expressive (`accurate_mode:false`), `emotion_cfg_scale:1.5`, pure-axes (no
sliders). `emo` = emotion-classifier target prob, `id` = speaker-identity
retention (floor 0.65; neutral ~0.85).
| Voice | preset | valence | arousal | strength | emo | id |
|---|---|---|---|---|---|---|
| AmericanFemale | `angry` | 0.4 | +1.0 | 1.0 | 0.53 | 0.69 |
| AmericanFemale | `happy` / `startled_happy` | +0.6 | +0.8 | 1.0 | 1.0 | 0.80 |
| AmericanMale | `angry` (drama) | 0.6 | +0.8 | 1.2 | 1.0 | 0.62 |
| AmericanMale | `happy` / `startled_happy` | +0.3 | +1.0 | 1.0 | 0.98 | 0.76 |
| BritishFemale | `angry` | 0.4 | +0.8 | 1.0 | 0.99 | 0.73 |
| BritishFemale | `happy` / `startled_happy` | +0.6 | +1.0 | 1.0 | 1.0 | 0.74 |
**Uncalibrated voices** (Cora + cloned characters) fall back to a mid-region
default (`angry` = v0.5/a+0.9/str1.1; `happy`/`startled_happy` = v+0.5/a+0.9)
until they earn a measured row. The live table is in `GET /v1/dials`
`voice_emotion_presets`.
### Caveats (why "provisional")
1. **Non-monotonic surface** — the exact `(valence, arousal, strength)` is
pinned; do NOT interpolate a preset's neighborhood in a UI.
2. **Anger costs identity** — retention ~0.620.73 vs ~0.85 neutral. Acceptable
for a drama beat, not for identity-critical dialogue.
3. **Surprised-as-class is dead** — never expose a "surprised" *slider* promise;
map user-intent surprised/startled/shocked to the `startled_happy` preset.
4. **`emotion_strength` is not a smooth knob** (esp. AmericanMale: 1.0 = mild,
1.2 = full anger, higher flips to *disgust*). The `angry` preset pins the
voice's tuned strength; send an explicit `emotion_strength` to shift tier
(e.g. `1.0` on AmericanMale for a softer anger).
5. **Two-regime policy** — identity-critical lines: `accurate_mode:true`,
emotion off or soft. Tagged drama beats: these presets. Don't stack a high
named slider AND a high axes preset (single-condition wins).
Baked provisional pending confident ear-validation on emotion-congruent text
(neutral-sentence auditing was inconclusive — the dials clearly differ, but a
flat line understates landing). Re-tune when we revisit.
## 6. Starting-point flavor dial-sets (examples — copy and tweak)
Tuned by ear on BritishFemale. Departure points, not a menu. (Distinct from §5:
these are hand-tuned flavors, not measured emotion cells.)
```jsonc ```jsonc
// warm — friendly, unhurried // warm — friendly, unhurried
{ "emotion_enabled": true, "emotion_sliders": {"happy": 0.4}, { "emotion_enabled": true, "emotion_sliders": {"happy": 0.4},
"emotion_valence": 0.4, "emotion_strength": 0.8, "emotion_cfg_scale": 1.3 } "emotion_valence": 0.4, "emotion_strength": 0.8, "emotion_cfg_scale": 1.3 }
// excited — bright, brisk (NO 'surprised' — that chipmunks it)
{ "emotion_enabled": true, "emotion_sliders": {"happy": 0.55},
"emotion_valence": 0.5, "emotion_arousal": 0.35, "emotion_strength": 0.75,
"emotion_cfg_scale": 1.25, "speaking_rate_enabled": true, "speaking_rate_bucket": 4 }
// sad — slow, quiet // sad — slow, quiet
{ "emotion_enabled": true, "emotion_sliders": {"sad": 0.8}, { "emotion_enabled": true, "emotion_sliders": {"sad": 0.8},
"emotion_valence": -0.6, "emotion_arousal": -0.5, "emotion_strength": 1.0, "emotion_valence": -0.6, "emotion_arousal": -0.5, "emotion_strength": 1.0,
"emotion_cfg_scale": 1.4, "speaking_rate_enabled": true, "speaking_rate_bucket": 1 } "emotion_cfg_scale": 1.5, "speaking_rate_enabled": true, "speaking_rate_bucket": 1 }
// intense — emphatic, tense (angry only; no 'surprised')
{ "emotion_enabled": true, "emotion_sliders": {"angry": 0.55},
"emotion_valence": -0.25, "emotion_arousal": 0.45, "emotion_strength": 0.85,
"emotion_cfg_scale": 1.3 }
// whisper — hushed, breathy // whisper — hushed, breathy
{ "emotion_enabled": true, "emotion_sliders": {"sad": 0.2}, { "emotion_enabled": true, "emotion_sliders": {"sad": 0.2},
@@ -101,41 +170,41 @@ Tuned by ear on BritishFemale. Treat these as departure points, not a menu.
"speaking_rate_enabled": true, "speaking_rate_bucket": 2 } "speaking_rate_enabled": true, "speaking_rate_bucket": 2 }
``` ```
## 6. LLM client system-prompt snippet (the real deliverable) ## 7. LLM client system-prompt snippet
Paste into an LLM client that drives TTS, so it emits a dial-set per utterance: Paste into an LLM client that drives TTS, so it emits a dial-set per utterance:
> You control speech emotion by emitting a JSON dial-set alongside each line. > You control speech emotion per line. For **anger** or **surprise/startle**, use
> Set `emotion_enabled: true`, then shape the delivery: > a preset: `{"preset": "angry"}` or `{"preset": "startled_happy"}` (the named
> - `emotion_sliders`: weights (00.8) over `happy` / `sad` / `angry` / > `angry`/`surprised` sliders misfire — don't use them). For everything else,
> `surprised`. Blend them. **Avoid `surprised` unless you truly want a gasp — > emit a dial-set: set `emotion_enabled: true`, then shape delivery with
> it pitches the voice up unnaturally.** > `emotion_sliders` (happy/sad ~00.8), `emotion_valence`/`emotion_arousal`
> - `emotion_valence` (1..+1): unpleasant → pleasant. `emotion_arousal` > (1..+1), `emotion_strength` (~0.71.1), and `emotion_cfg_scale`
> (1..+1): calm → energized. Use these for nuance between the four emotions. > (**1.01.4, never above 1.5**). Add `speaking_rate_enabled: true` +
> - `emotion_strength` (~0.71.1): how hard to push. `emotion_cfg_scale` > `speaking_rate_bucket` (0 slow … 7 fast) for pacing. Subtle by default, bold
> (**1.01.4, never above 1.5**): amplification — 1.0 is already expressive; > only when the moment earns it.
> 1.251.4 lands a line harder; above 1.5 is distorted garbage, do not.
> - `speaking_rate_enabled: true` + `speaking_rate_bucket` (0 slow … 7 fast) for
> pacing: fast when excited/anxious, slow when sad/weary.
> Match the dial-set to the line's actual feeling — subtle by default, bold only
> when the moment earns it.
## 7. Voice cloning (reference audio) ## 8. Voice cloning (reference audio)
Cloning is text-independent (Qwen3 speaker embedding) — **no transcript needed.** Cloning is text-independent (Qwen3 speaker embedding) — **no transcript needed.**
Give it **~1524s of clean, single-speaker audio** (Zyphra's own default voices Give it **~1524s of clean, single-speaker audio** (Zyphra's own default voices
run 8.723.9s). Shorter clips (≤5s) under-condition the embedding and audibly run 8.723.9s). Shorter clips (≤5s) under-condition the embedding and audibly
degrade, worse under emotion steering. Register a reusable voice via degrade, worse under emotion steering. A cloned voice has **no calibrated
`POST /tts/speakers` (caches the embedding; reference by `speaker_embedding_id`). emotion row** yet — it uses the §5 fallback; add a measured row (or map it to the
nearest default) when its emotions matter.
## 8. Follow-ups (not done here) ## 9. Follow-ups (not done here)
- **Gateway code:** align `dials.py` `emotion_cfg_scale` help/`max` metadata with - **Ear-validation on congruent text** — the §5 cells were baked provisional off
§3 (carry the "deaf by 1.5" warning; no clamp added — no cap by ruling). Trim an inconclusive neutral-text audition; confirm on emotion-appropriate lines and
the in-code `PRESETS` to a couple of examples (or drop) once current `preset:` promote from provisional (or re-tune).
usage in asset-engine / gateway-chat is confirmed. - **Sad axes pass** — sad was not in the axes sweep; still the named-slider
- **Version control:** DONE (2026-07-17) — `vh/zonos-gateway` stood up on gitea preset. Owes a per-voice text/strength pass on the 3 defaults.
(private); this spec's canonical home is now `docs/EMOTION-DIALS-SPEC.md` in - **AmericanMale angry sweet-spot** — no single cell clears both the emo and
that repo (the eshpfi copy is a mirror). Remaining: git-connect / CI-wire the identity floors; it ships as drama (str 1.2) with a documented soft override
deployed irv-ml1 working tree (build + deploy on push, like the other sister (str 1.0). A finer strength ladder could find a middle tier.
services) with a deploy key. - **Clone-character rows** — Emmie/Penny/Natalie/Miranda use the fallback; run a
1-row confirm at the baked cell per voice, or map to the nearest default.
- **CI / deploy wiring** — the deployed irv-ml1 tree is a hand-updated build
context (not git / not CI-deployed). Git-connect + build-on-push with a deploy
key, like the other sister services.