Files
esh-pfi-infrastructure/voices/README.md
T
vh fca1a545f1 feat(voices): canonical voice corpus + dots.tts-optimized refs
Engine-agnostic voice corpus: canonical source clip + transcript per voice,
per-engine reference sets derived by derive.py from engines.yaml profiles.
First residents donut/glados/emmie/miranda optimized + verified clean for
dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference
audio into output otherwise). canonical/ + transcripts/ tracked; derived/
gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
2026-08-10 00:32:10 -07:00

3.2 KiB
Raw Blame History

Canonical voice corpus

Engine-agnostic source of truth for cloned voice identities. Each voice is stored once as a canonical source clip + an accurate transcript; per-engine reference sets (dots.tts, chatterbox, zonos, …) are derived from it on demand by derive.py. Adding a new TTS engine is "add a profile to engines.yaml and re-derive" — not "re-hunt every voice."

Why this exists

TTS engines disagree on what a reference clip must be:

Engine SR Transcript? Reference shape
dots.tts 48kHz required ~≤10s, sentence-bounded, accurate transcript
chatterbox-fast 24kHz no any length, audio-only
Zonos2 44.1kHz no any length, audio-only + emotion dials

Keeping one canonical source per voice + a derivation step means a voice cloned for Zonos a year ago can be re-optimized for whatever engine comes next without re-sourcing the audio.

The dots.tts sensitivity finding (load-bearing)

dots.tts conditions each generation on (reference audio + its transcript) and regurgitates reference content into the output when the transcript is wrong or ends mid-clause. Symptoms seen during the 2026-08 burn-in: a mismatched transcript collapsed output to 0.16s; an over-long reference with a repetitive transcript prefixed the output with reference lines; a transcript trimmed mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable recipe — encoded in derive.py for the dots profile — is trim to a clean ~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of exactly that clip.

Layout

voices/
  manifest.yaml         # voice registry: canonical path, transcript, SR, provenance
  engines.yaml          # per-engine reference requirements
  derive.py             # canonical -> derived/<engine>/<voice>.{wav,txt}
  canonical/<v>.wav     # source clip, best available SR (git-tracked, small + curated)
  transcripts/<v>.txt   # full accurate transcript of the canonical source
  derived/              # per-engine reference sets (GITIGNORED — regenerable)
    dots/<v>.{wav,txt}
    chatterbox/<v>.wav

Usage

# derive dots-ready references for every voice (needs faster-whisper for the trim):
python derive.py dots

# just two voices:
python derive.py dots donut glados

# a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml):
python derive.py chatterbox

derived/ is gitignored — treat it as a build output. Deploy a derived set to a live engine by copying derived/<engine>/ into that stack's refs dir (e.g. dots' voices mount, chatterbox /worktank/chatterbox/reference_audio/).

Adding a voice

  1. Drop the best available source clip in canonical/<name>.wav (highest SR, cleanest, ~10–30s is plenty).
  2. Add its row to manifest.yaml (SR, duration, provenance).
  3. python derive.py dots <name> — writes the transcript + dots reference and, if you wire it, a verify pass.

Provenance discipline

Record where each source came from in manifest.yaml. Unknown origin is fine to start (origin unrecorded) but should be filled in when known — a canonical corpus is only as trustworthy as its provenance.