feat(voices): canonical voice corpus + dots.tts-optimized refs

Engine-agnostic voice corpus: canonical source clip + transcript per voice,
per-engine reference sets derived by derive.py from engines.yaml profiles.
First residents donut/glados/emmie/miranda optimized + verified clean for
dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference
audio into output otherwise). canonical/ + transcripts/ tracked; derived/
gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
This commit is contained in:
vh
2026-08-10 00:32:10 -07:00
parent 58b58d1401
commit fca1a545f1
14 changed files with 286 additions and 0 deletions
+40
View File
@@ -0,0 +1,40 @@
# Per-engine reference requirements. derive.py reads this to turn a canonical
# source + transcript into an engine-ready reference set under derived/<engine>/.
#
# Fields:
# sample_rate native SR the engine wants
# resample true = derive.py should resample to sample_rate
# (NOTE: resample is a follow-up — see the resample TODO
# in derive.py; dots resamples internally so it's false there)
# needs_transcript engine requires a per-reference transcript file
# ref_max_seconds cap on derived reference length (null = uncapped)
# ref_sentence_bounded transcript/clip must end on a sentence boundary (. ! ?)
# — set for engines that leak reference content otherwise
engines:
dots:
description: "dots.tts (rednote-hilab) — continuous-AR 48kHz zero-shot clone"
sample_rate: 48000
resample: false # runtime auto-resamples at load; keep source SR
needs_transcript: true # REQUIRED and must be accurate + sentence-bounded
ref_max_seconds: 10
ref_sentence_bounded: true
notes: >
Transcript accuracy AND sentence-boundary are load-bearing: a mismatched or
mid-clause transcript makes dots regurgitate reference audio into the output.
chatterbox:
description: "chatterbox-fast (Turbo) — streaming 24kHz clone"
sample_rate: 24000
resample: true
needs_transcript: false # audio-only clone; server globs its refs dir live
ref_max_seconds: null
ref_sentence_bounded: false
zonos:
description: "Zonos2 — expressive 44.1kHz clone + emotion dials"
sample_rate: 44100
resample: true
needs_transcript: false
ref_max_seconds: null
ref_sentence_bounded: false