Files
esh-pfi-infrastructure/persistent-memory.d/2026-07-18-zonos-emotion-canonical.md
T
vh 89611eb06b memory: snapshot — Zonos emotion-tuning + voice-cloning (8 voices, dial-in studio, emotion canonical) for /clear
Rewrote in-flight for the Zonos character-voice work: 4 cloned voices + host-managed
gateway voices, streaming dial-in studio (source saved to ~/development/zonos-tools/),
and the empirical emotion sweep canonical (single-emotion, two-regime accurate/expressive;
happy/sad usable, angry/surprised broken on named dirs -> axes sweep next). Captured #365
closed + WT#368 forensics + personal agent-memory scrub. Open loops: yt-voice-clipper
yields test, dvalin axes-sweep numbers, re-arm monitor + read mail.
2026-07-18 00:25:18 -07:00

4.6 KiB
Raw Blame History

  • [2026-07-18] Zonos2 emotion CANONICAL from an empirical sweep + the voice-cloning pipeline.

Voice-cloning pipeline (established this session). Source zips at /mnt/smithy/voice_clones/<name>.zip (irv-ml1 NFS from nh3-nas; remount post-reboot) — each = diarized single-speaker podcast clips + manifest.jsonl (per-clip WhisperX mean_score, word timestamps, text) + metadata.csv. ~/development/zonos-tools/assemble_voice.py <dir> ranks by mean_score and concatenates top clips to ~1524s (Zyphra's blessed clone-ref length; single clip if already ≥15s). Drop the assembled <Name>.wav into the gateway voices dir → voice:"name". 4 characters cloned: Emmie, Penny, Natalie, Miranda (+ Zyphra defaults AmericanFemale/Male/British/Cora) = 8 voices in zonos-gateway. Clone is inline speaker_audio_base64 (text-independent Qwen3 speaker embedding — NO transcript); /tts/speakers registration is session-scoped (needs X-TTS-Session-ID), so the gateway holds the ref wav and clones per-call.

Gateway voices are host-managed (bind-mount, added this session). Added ./voices:/app/voices:ro to /opt/docker/compose/zonos-gateway/compose.yaml (committed to vh/zonos-gateway + eshpfi mirror 438cd35). So adding a voice = drop the wav + docker compose restart zonos-gateway (registry rebuilds at boot; NO image rebuild). This also un-stranded the other voices (deploy build context had only Cora before). Voice wavs committed to the repo for backup.

Emotion mechanism (Zyphra canonical, from their README @194c0a3). Additive direction vectors: 4 named (happy/sad/angry/surprised) + valence/arousal axes. emotion_strength 1.0 = per-voice calibrated (calibration.json optimizes emotion2vec recognizability only, NOT identity). accurate_mode is THE trade-off: true = closer voice match (identity), false = expressive mode (emotion lands, identity drifts). Zyphra's strong recipe: accurate_mode:false + cfg~1.5. Single-emotion is blessed; mixing is unblessed (and degrades the clone — operator confirmed by ear). "deaf by 1.5" — cfg past 1.5 distorts + costs ~2× compute.

THE SWEEP (~/development/zonos-tools/emotion_sweep.py). 4 cloned voices × 4 named emotions × {accurate,expressive}×{cfg 1.0,1.3,1.5} @ strength 1.0, single-emotion, neutral sentence + a neutral baseline per voice (~100 clips). Scored on TWO axes: emotion-landing = emotion2vec iic/emotion2vec_plus_large target-emotion prob [0-1]; identity = resemblyzer speaker-embedding cosine vs the clone reference (neutral baseline ~0.85). Scoring env: uv run --with resemblyzer --with funasr --with "numpy<2" --with soundfile --with requests --with "setuptools<80" --with torchaudio (setuptools<80 for webrtcvad's pkg_resources; torchaudio for funasr).

RESULTS (mean across the 4 voices) — emotion, best setting, emo/id:

  • happy — exp cfg1.5 0.80/0.68 (soft: exp cfg1.0 0.76/0.69) → WORKS
  • sad — exp cfg1.5 0.53/0.57 (only working cell; id below the ~0.65 floor) → modest
  • angry — acc cfg1.3 / exp cfg1.5 tied at ~0.25 emo → WEAK (named ceiling ~0.25)
  • surprised — max ~0.015 across ALL settings → NON-FUNCTIONAL on the named direction Accurate + low cfg = identity/suppress regime (emo→0); expressive REQUIRED for emotion to land, at ~0.150.28 identity cost.

dvalin-smithy-dev synthesis (adopted, triaged genuine-adds; thread 01KXT12FN0AS5A3WMKEK06BVPS):

  1. Treat identity as a hard FLOOR (~0.65), not a free variable in emo×id.
  2. Two-regime policy — Regime A (default, identity-critical dialogue): accurate_mode:true, cfg 1.0, emotion off (text carries it) or soft-happy (exp cfg1.0). Regime B (tagged drama beats): accurate_mode:false, cfg 1.5, single emotion or axes. Line-type→regime heuristic (exposition→A, grief→B+sad, confrontation→B+axes-angry, shock→B+axes-arousal).
  3. Axes-first for the broken emotions — angry ≈ valence 0.6..0.8 / arousal +0.5..+0.8; surprised ≈ valence +0.2..+0.4 / arousal +0.7..+1.0 (exp cfg1.5); or "startled-happy" (happy + high arousal) as a surprised stand-in. These are PROVISIONAL — the sweep did NOT test axes.

NEXT (highest VoI, operator to green-light): an axes sweep for angry/surprised (valence×arousal grid) — the only path to rescue the two broken named emotions; then a strength ladder at the best cells + emotion-congruent text (neutral content understates landing) + per-voice tables + a 2nd emotion judge / human pairwise. Then bake the happy/sad canonical into gateway presets. I owe dvalin the axes-sweep numbers.

See reference_zonos_tts_stack; dials-first spec at vh/zonos-gateway docs/EMOTION-DIALS-SPEC.md.