# Canonical voice corpus Engine-agnostic source of truth for cloned voice identities. Each voice is stored **once** as a canonical source clip + an accurate transcript; per-engine reference sets (dots.tts, chatterbox, zonos, …) are **derived** from it on demand by [`derive.py`](derive.py). Adding a new TTS engine is "add a profile to [`engines.yaml`](engines.yaml) and re-derive" — not "re-hunt every voice." ## Why this exists TTS engines disagree on what a reference clip must be: | Engine | SR | Transcript? | Reference shape | |---|---|---|---| | **dots.tts** | 48kHz | **required** | ~≤10s, **sentence-bounded**, accurate transcript | | chatterbox-fast | 24kHz | no | any length, audio-only | | Zonos2 | 44.1kHz | no | any length, audio-only + emotion dials | Keeping one canonical source per voice + a derivation step means a voice cloned for Zonos a year ago can be re-optimized for whatever engine comes next without re-sourcing the audio. ## The dots.tts sensitivity finding (load-bearing) dots.tts conditions each generation on (reference audio **+ its transcript**) and **regurgitates reference content into the output** when the transcript is wrong **or ends mid-clause**. Symptoms seen during the 2026-08 burn-in: a mismatched transcript collapsed output to 0.16s; an over-long reference with a repetitive transcript prefixed the output with reference lines; a transcript trimmed mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable recipe — encoded in `derive.py` for the `dots` profile — is **trim to a clean ~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of exactly that clip.** ## Layout ``` voices/ manifest.yaml # voice registry: canonical path, transcript, SR, provenance engines.yaml # per-engine reference requirements derive.py # canonical -> derived//.{wav,txt} canonical/.wav # source clip, best available SR (git-tracked, small + curated) transcripts/.txt # full accurate transcript of the canonical source derived/ # per-engine reference sets (GITIGNORED — regenerable) dots/.{wav,txt} chatterbox/.wav ``` ## Usage ```bash # derive dots-ready references for every voice (needs faster-whisper for the trim): python derive.py dots # just two voices: python derive.py dots donut glados # a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml): python derive.py chatterbox ``` `derived/` is gitignored — treat it as a build output. Deploy a derived set to a live engine by copying `derived//` into that stack's refs dir (e.g. dots' voices mount, chatterbox `/worktank/chatterbox/reference_audio/`). ## Adding a voice 1. Drop the best available source clip in `canonical/.wav` (highest SR, cleanest, ~10–30s is plenty). 2. Add its row to `manifest.yaml` (SR, duration, provenance). 3. `python derive.py dots ` — writes the transcript + dots reference and, if you wire it, a verify pass. ## Provenance discipline Record where each source came from in `manifest.yaml`. Unknown origin is fine to start (`origin unrecorded`) but should be filled in when known — a canonical corpus is only as trustworthy as its provenance.