Engine-agnostic voice corpus: canonical source clip + transcript per voice, per-engine reference sets derived by derive.py from engines.yaml profiles. First residents donut/glados/emmie/miranda optimized + verified clean for dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference audio into output otherwise). canonical/ + transcripts/ tracked; derived/ gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
3.2 KiB
Canonical voice corpus
Engine-agnostic source of truth for cloned voice identities. Each voice is
stored once as a canonical source clip + an accurate transcript; per-engine
reference sets (dots.tts, chatterbox, zonos, …) are derived from it on
demand by derive.py. Adding a new TTS engine is "add a profile
to engines.yaml and re-derive" — not "re-hunt every voice."
Why this exists
TTS engines disagree on what a reference clip must be:
| Engine | SR | Transcript? | Reference shape |
|---|---|---|---|
| dots.tts | 48kHz | required | ~≤10s, sentence-bounded, accurate transcript |
| chatterbox-fast | 24kHz | no | any length, audio-only |
| Zonos2 | 44.1kHz | no | any length, audio-only + emotion dials |
Keeping one canonical source per voice + a derivation step means a voice cloned for Zonos a year ago can be re-optimized for whatever engine comes next without re-sourcing the audio.
The dots.tts sensitivity finding (load-bearing)
dots.tts conditions each generation on (reference audio + its transcript) and
regurgitates reference content into the output when the transcript is wrong
or ends mid-clause. Symptoms seen during the 2026-08 burn-in: a mismatched
transcript collapsed output to 0.16s; an over-long reference with a repetitive
transcript prefixed the output with reference lines; a transcript trimmed
mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable
recipe — encoded in derive.py for the dots profile — is trim to a clean
~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of
exactly that clip.
Layout
voices/
manifest.yaml # voice registry: canonical path, transcript, SR, provenance
engines.yaml # per-engine reference requirements
derive.py # canonical -> derived/<engine>/<voice>.{wav,txt}
canonical/<v>.wav # source clip, best available SR (git-tracked, small + curated)
transcripts/<v>.txt # full accurate transcript of the canonical source
derived/ # per-engine reference sets (GITIGNORED — regenerable)
dots/<v>.{wav,txt}
chatterbox/<v>.wav
Usage
# derive dots-ready references for every voice (needs faster-whisper for the trim):
python derive.py dots
# just two voices:
python derive.py dots donut glados
# a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml):
python derive.py chatterbox
derived/ is gitignored — treat it as a build output. Deploy a derived set to a
live engine by copying derived/<engine>/ into that stack's refs dir
(e.g. dots' voices mount, chatterbox /worktank/chatterbox/reference_audio/).
Adding a voice
- Drop the best available source clip in
canonical/<name>.wav(highest SR, cleanest, ~10–30s is plenty). - Add its row to
manifest.yaml(SR, duration, provenance). python derive.py dots <name>— writes the transcript + dots reference and, if you wire it, a verify pass.
Provenance discipline
Record where each source came from in manifest.yaml. Unknown origin is fine to
start (origin unrecorded) but should be filled in when known — a canonical
corpus is only as trustworthy as its provenance.