Files
esh-pfi-infrastructure/voices/README.md
T
vh fca1a545f1 feat(voices): canonical voice corpus + dots.tts-optimized refs
Engine-agnostic voice corpus: canonical source clip + transcript per voice,
per-engine reference sets derived by derive.py from engines.yaml profiles.
First residents donut/glados/emmie/miranda optimized + verified clean for
dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference
audio into output otherwise). canonical/ + transcripts/ tracked; derived/
gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
2026-08-10 00:32:10 -07:00

79 lines
3.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Canonical voice corpus
Engine-agnostic source of truth for cloned voice identities. Each voice is
stored **once** as a canonical source clip + an accurate transcript; per-engine
reference sets (dots.tts, chatterbox, zonos, …) are **derived** from it on
demand by [`derive.py`](derive.py). Adding a new TTS engine is "add a profile
to [`engines.yaml`](engines.yaml) and re-derive" — not "re-hunt every voice."
## Why this exists
TTS engines disagree on what a reference clip must be:
| Engine | SR | Transcript? | Reference shape |
|---|---|---|---|
| **dots.tts** | 48kHz | **required** | ~≤10s, **sentence-bounded**, accurate transcript |
| chatterbox-fast | 24kHz | no | any length, audio-only |
| Zonos2 | 44.1kHz | no | any length, audio-only + emotion dials |
Keeping one canonical source per voice + a derivation step means a voice cloned
for Zonos a year ago can be re-optimized for whatever engine comes next without
re-sourcing the audio.
## The dots.tts sensitivity finding (load-bearing)
dots.tts conditions each generation on (reference audio **+ its transcript**) and
**regurgitates reference content into the output** when the transcript is wrong
**or ends mid-clause**. Symptoms seen during the 2026-08 burn-in: a mismatched
transcript collapsed output to 0.16s; an over-long reference with a repetitive
transcript prefixed the output with reference lines; a transcript trimmed
mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable
recipe — encoded in `derive.py` for the `dots` profile — is **trim to a clean
~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of
exactly that clip.**
## Layout
```
voices/
manifest.yaml # voice registry: canonical path, transcript, SR, provenance
engines.yaml # per-engine reference requirements
derive.py # canonical -> derived/<engine>/<voice>.{wav,txt}
canonical/<v>.wav # source clip, best available SR (git-tracked, small + curated)
transcripts/<v>.txt # full accurate transcript of the canonical source
derived/ # per-engine reference sets (GITIGNORED — regenerable)
dots/<v>.{wav,txt}
chatterbox/<v>.wav
```
## Usage
```bash
# derive dots-ready references for every voice (needs faster-whisper for the trim):
python derive.py dots
# just two voices:
python derive.py dots donut glados
# a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml):
python derive.py chatterbox
```
`derived/` is gitignored — treat it as a build output. Deploy a derived set to a
live engine by copying `derived/<engine>/` into that stack's refs dir
(e.g. dots' voices mount, chatterbox `/worktank/chatterbox/reference_audio/`).
## Adding a voice
1. Drop the best available source clip in `canonical/<name>.wav` (highest SR,
cleanest, ~10–30s is plenty).
2. Add its row to `manifest.yaml` (SR, duration, provenance).
3. `python derive.py dots <name>` — writes the transcript + dots reference and,
if you wire it, a verify pass.
## Provenance discipline
Record where each source came from in `manifest.yaml`. Unknown origin is fine to
start (`origin unrecorded`) but should be filled in when known — a canonical
corpus is only as trustworthy as its provenance.