feat(voices): canonical voice corpus + dots.tts-optimized refs
Engine-agnostic voice corpus: canonical source clip + transcript per voice, per-engine reference sets derived by derive.py from engines.yaml profiles. First residents donut/glados/emmie/miranda optimized + verified clean for dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference audio into output otherwise). canonical/ + transcripts/ tracked; derived/ gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
This commit is contained in:
@@ -0,0 +1,78 @@
|
||||
# Canonical voice corpus
|
||||
|
||||
Engine-agnostic source of truth for cloned voice identities. Each voice is
|
||||
stored **once** as a canonical source clip + an accurate transcript; per-engine
|
||||
reference sets (dots.tts, chatterbox, zonos, …) are **derived** from it on
|
||||
demand by [`derive.py`](derive.py). Adding a new TTS engine is "add a profile
|
||||
to [`engines.yaml`](engines.yaml) and re-derive" — not "re-hunt every voice."
|
||||
|
||||
## Why this exists
|
||||
|
||||
TTS engines disagree on what a reference clip must be:
|
||||
|
||||
| Engine | SR | Transcript? | Reference shape |
|
||||
|---|---|---|---|
|
||||
| **dots.tts** | 48kHz | **required** | ~≤10s, **sentence-bounded**, accurate transcript |
|
||||
| chatterbox-fast | 24kHz | no | any length, audio-only |
|
||||
| Zonos2 | 44.1kHz | no | any length, audio-only + emotion dials |
|
||||
|
||||
Keeping one canonical source per voice + a derivation step means a voice cloned
|
||||
for Zonos a year ago can be re-optimized for whatever engine comes next without
|
||||
re-sourcing the audio.
|
||||
|
||||
## The dots.tts sensitivity finding (load-bearing)
|
||||
|
||||
dots.tts conditions each generation on (reference audio **+ its transcript**) and
|
||||
**regurgitates reference content into the output** when the transcript is wrong
|
||||
**or ends mid-clause**. Symptoms seen during the 2026-08 burn-in: a mismatched
|
||||
transcript collapsed output to 0.16s; an over-long reference with a repetitive
|
||||
transcript prefixed the output with reference lines; a transcript trimmed
|
||||
mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable
|
||||
recipe — encoded in `derive.py` for the `dots` profile — is **trim to a clean
|
||||
~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of
|
||||
exactly that clip.**
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
voices/
|
||||
manifest.yaml # voice registry: canonical path, transcript, SR, provenance
|
||||
engines.yaml # per-engine reference requirements
|
||||
derive.py # canonical -> derived/<engine>/<voice>.{wav,txt}
|
||||
canonical/<v>.wav # source clip, best available SR (git-tracked, small + curated)
|
||||
transcripts/<v>.txt # full accurate transcript of the canonical source
|
||||
derived/ # per-engine reference sets (GITIGNORED — regenerable)
|
||||
dots/<v>.{wav,txt}
|
||||
chatterbox/<v>.wav
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# derive dots-ready references for every voice (needs faster-whisper for the trim):
|
||||
python derive.py dots
|
||||
|
||||
# just two voices:
|
||||
python derive.py dots donut glados
|
||||
|
||||
# a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml):
|
||||
python derive.py chatterbox
|
||||
```
|
||||
|
||||
`derived/` is gitignored — treat it as a build output. Deploy a derived set to a
|
||||
live engine by copying `derived/<engine>/` into that stack's refs dir
|
||||
(e.g. dots' voices mount, chatterbox `/worktank/chatterbox/reference_audio/`).
|
||||
|
||||
## Adding a voice
|
||||
|
||||
1. Drop the best available source clip in `canonical/<name>.wav` (highest SR,
|
||||
cleanest, ~10–30s is plenty).
|
||||
2. Add its row to `manifest.yaml` (SR, duration, provenance).
|
||||
3. `python derive.py dots <name>` — writes the transcript + dots reference and,
|
||||
if you wire it, a verify pass.
|
||||
|
||||
## Provenance discipline
|
||||
|
||||
Record where each source came from in `manifest.yaml`. Unknown origin is fine to
|
||||
start (`origin unrecorded`) but should be filled in when known — a canonical
|
||||
corpus is only as trustworthy as its provenance.
|
||||
Reference in New Issue
Block a user