Engine-agnostic voice corpus: canonical source clip + transcript per voice, per-engine reference sets derived by derive.py from engines.yaml profiles. First residents donut/glados/emmie/miranda optimized + verified clean for dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference audio into output otherwise). canonical/ + transcripts/ tracked; derived/ gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
41 lines
1.7 KiB
YAML
41 lines
1.7 KiB
YAML
# Per-engine reference requirements. derive.py reads this to turn a canonical
|
|
# source + transcript into an engine-ready reference set under derived/<engine>/.
|
|
#
|
|
# Fields:
|
|
# sample_rate native SR the engine wants
|
|
# resample true = derive.py should resample to sample_rate
|
|
# (NOTE: resample is a follow-up — see the resample TODO
|
|
# in derive.py; dots resamples internally so it's false there)
|
|
# needs_transcript engine requires a per-reference transcript file
|
|
# ref_max_seconds cap on derived reference length (null = uncapped)
|
|
# ref_sentence_bounded transcript/clip must end on a sentence boundary (. ! ?)
|
|
# — set for engines that leak reference content otherwise
|
|
|
|
engines:
|
|
dots:
|
|
description: "dots.tts (rednote-hilab) — continuous-AR 48kHz zero-shot clone"
|
|
sample_rate: 48000
|
|
resample: false # runtime auto-resamples at load; keep source SR
|
|
needs_transcript: true # REQUIRED and must be accurate + sentence-bounded
|
|
ref_max_seconds: 10
|
|
ref_sentence_bounded: true
|
|
notes: >
|
|
Transcript accuracy AND sentence-boundary are load-bearing: a mismatched or
|
|
mid-clause transcript makes dots regurgitate reference audio into the output.
|
|
|
|
chatterbox:
|
|
description: "chatterbox-fast (Turbo) — streaming 24kHz clone"
|
|
sample_rate: 24000
|
|
resample: true
|
|
needs_transcript: false # audio-only clone; server globs its refs dir live
|
|
ref_max_seconds: null
|
|
ref_sentence_bounded: false
|
|
|
|
zonos:
|
|
description: "Zonos2 — expressive 44.1kHz clone + emotion dials"
|
|
sample_rate: 44100
|
|
resample: true
|
|
needs_transcript: false
|
|
ref_max_seconds: null
|
|
ref_sentence_bounded: false
|