Engine-agnostic voice corpus: canonical source clip + transcript per voice, per-engine reference sets derived by derive.py from engines.yaml profiles. First residents donut/glados/emmie/miranda optimized + verified clean for dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference audio into output otherwise). canonical/ + transcripts/ tracked; derived/ gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
36 lines
1.2 KiB
YAML
36 lines
1.2 KiB
YAML
# Canonical voice corpus registry. One row per voice; the canonical clip + its
|
|
# full transcript are the source of truth, engine-agnostic. derive.py reads this
|
|
# together with engines.yaml to produce per-engine reference sets.
|
|
|
|
voices:
|
|
donut:
|
|
canonical: canonical/donut.wav
|
|
transcript: transcripts/donut.txt
|
|
source_sr: 44100
|
|
duration_s: 16.3
|
|
character: "sassy fairy-charm kid"
|
|
provenance: "cloned from the 65-frost Booth bundle (2026-08)"
|
|
|
|
glados:
|
|
canonical: canonical/glados.wav
|
|
transcript: transcripts/glados.txt
|
|
source_sr: 16000
|
|
duration_s: 25.0
|
|
character: "GLaDOS — flat, deliberate, menacing-cheerful"
|
|
provenance: "Portal GLaDOS lines"
|
|
warning: "LOW-SR source (16kHz) — upgrade the canonical clip if a cleaner GLaDOS source surfaces"
|
|
|
|
emmie:
|
|
canonical: canonical/emmie.wav
|
|
transcript: transcripts/emmie.txt
|
|
source_sr: 24000
|
|
duration_s: 19.3
|
|
provenance: "Zonos clone added 2026-07-17; origin unrecorded"
|
|
|
|
miranda:
|
|
canonical: canonical/miranda.wav
|
|
transcript: transcripts/miranda.txt
|
|
source_sr: 24000
|
|
duration_s: 16.3
|
|
provenance: "Zonos clone added 2026-07-17; origin unrecorded"
|