feat(voices): canonical voice corpus + dots.tts-optimized refs
Engine-agnostic voice corpus: canonical source clip + transcript per voice, per-engine reference sets derived by derive.py from engines.yaml profiles. First residents donut/glados/emmie/miranda optimized + verified clean for dots.tts (sentence-bounded ref + accurate transcript — dots leaks reference audio into output otherwise). canonical/ + transcripts/ tracked; derived/ gitignored (regenerable). Records the dots.tts burn-in in persistent-memory.
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
# Per-engine reference sets are build outputs — regenerate with derive.py.
|
||||
derived/
|
||||
@@ -0,0 +1,78 @@
|
||||
# Canonical voice corpus
|
||||
|
||||
Engine-agnostic source of truth for cloned voice identities. Each voice is
|
||||
stored **once** as a canonical source clip + an accurate transcript; per-engine
|
||||
reference sets (dots.tts, chatterbox, zonos, …) are **derived** from it on
|
||||
demand by [`derive.py`](derive.py). Adding a new TTS engine is "add a profile
|
||||
to [`engines.yaml`](engines.yaml) and re-derive" — not "re-hunt every voice."
|
||||
|
||||
## Why this exists
|
||||
|
||||
TTS engines disagree on what a reference clip must be:
|
||||
|
||||
| Engine | SR | Transcript? | Reference shape |
|
||||
|---|---|---|---|
|
||||
| **dots.tts** | 48kHz | **required** | ~≤10s, **sentence-bounded**, accurate transcript |
|
||||
| chatterbox-fast | 24kHz | no | any length, audio-only |
|
||||
| Zonos2 | 44.1kHz | no | any length, audio-only + emotion dials |
|
||||
|
||||
Keeping one canonical source per voice + a derivation step means a voice cloned
|
||||
for Zonos a year ago can be re-optimized for whatever engine comes next without
|
||||
re-sourcing the audio.
|
||||
|
||||
## The dots.tts sensitivity finding (load-bearing)
|
||||
|
||||
dots.tts conditions each generation on (reference audio **+ its transcript**) and
|
||||
**regurgitates reference content into the output** when the transcript is wrong
|
||||
**or ends mid-clause**. Symptoms seen during the 2026-08 burn-in: a mismatched
|
||||
transcript collapsed output to 0.16s; an over-long reference with a repetitive
|
||||
transcript prefixed the output with reference lines; a transcript trimmed
|
||||
mid-clause ("…we will") leaked a stray "we'll" into the output. The reliable
|
||||
recipe — encoded in `derive.py` for the `dots` profile — is **trim to a clean
|
||||
~≤10s clip ending on a sentence boundary (. ! ?) with an accurate transcript of
|
||||
exactly that clip.**
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
voices/
|
||||
manifest.yaml # voice registry: canonical path, transcript, SR, provenance
|
||||
engines.yaml # per-engine reference requirements
|
||||
derive.py # canonical -> derived/<engine>/<voice>.{wav,txt}
|
||||
canonical/<v>.wav # source clip, best available SR (git-tracked, small + curated)
|
||||
transcripts/<v>.txt # full accurate transcript of the canonical source
|
||||
derived/ # per-engine reference sets (GITIGNORED — regenerable)
|
||||
dots/<v>.{wav,txt}
|
||||
chatterbox/<v>.wav
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# derive dots-ready references for every voice (needs faster-whisper for the trim):
|
||||
python derive.py dots
|
||||
|
||||
# just two voices:
|
||||
python derive.py dots donut glados
|
||||
|
||||
# a no-transcript engine (copies canonical; resample = follow-up, see engines.yaml):
|
||||
python derive.py chatterbox
|
||||
```
|
||||
|
||||
`derived/` is gitignored — treat it as a build output. Deploy a derived set to a
|
||||
live engine by copying `derived/<engine>/` into that stack's refs dir
|
||||
(e.g. dots' voices mount, chatterbox `/worktank/chatterbox/reference_audio/`).
|
||||
|
||||
## Adding a voice
|
||||
|
||||
1. Drop the best available source clip in `canonical/<name>.wav` (highest SR,
|
||||
cleanest, ~10–30s is plenty).
|
||||
2. Add its row to `manifest.yaml` (SR, duration, provenance).
|
||||
3. `python derive.py dots <name>` — writes the transcript + dots reference and,
|
||||
if you wire it, a verify pass.
|
||||
|
||||
## Provenance discipline
|
||||
|
||||
Record where each source came from in `manifest.yaml`. Unknown origin is fine to
|
||||
start (`origin unrecorded`) but should be filled in when known — a canonical
|
||||
corpus is only as trustworthy as its provenance.
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,125 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Derive per-engine reference sets from the canonical voice corpus.
|
||||
|
||||
Reads manifest.yaml + engines.yaml and, for a chosen engine, writes
|
||||
derived/<engine>/<voice>.wav (plus <voice>.txt when the engine needs a
|
||||
transcript).
|
||||
|
||||
Usage:
|
||||
python derive.py <engine> [voice ...] # default: every voice in manifest
|
||||
|
||||
Deps: pyyaml, soundfile. faster-whisper is imported lazily, only when an engine
|
||||
sets ref_sentence_bounded (dots) — it picks a clean sentence-boundary trim and
|
||||
its exact transcript. Run under a venv that has these (on irv-ml1 the dots +
|
||||
whisper venvs already do).
|
||||
|
||||
Known follow-up: `resample: true` engines (chatterbox, zonos) currently COPY the
|
||||
canonical clip at its source SR rather than resampling — a proper resample step
|
||||
(soundfile + a resampler) is a TODO. dots sets resample:false (it resamples
|
||||
internally at load), so the dots path is complete.
|
||||
"""
|
||||
import sys
|
||||
import wave
|
||||
import pathlib
|
||||
import shutil
|
||||
import yaml
|
||||
|
||||
ROOT = pathlib.Path(__file__).parent
|
||||
|
||||
|
||||
def load():
|
||||
manifest = yaml.safe_load((ROOT / "manifest.yaml").read_text())["voices"]
|
||||
engines = yaml.safe_load((ROOT / "engines.yaml").read_text())["engines"]
|
||||
return manifest, engines
|
||||
|
||||
|
||||
DANGLING = {"and", "but", "so", "or", "the", "a", "an", "that", "to", "my",
|
||||
"because", "with", "of", "for", "as", "i", "we", "it", "is"}
|
||||
|
||||
|
||||
def sentence_bounded_trim(src, target_s, model, min_s=6.0):
|
||||
"""Return (end_seconds, transcript) for a clip ending on a real sentence
|
||||
boundary.
|
||||
|
||||
Accumulates whisper segments and takes the FIRST point past `min_s` where the
|
||||
running transcript ends in . ! ? — searching up to target_s+4 so a run-on
|
||||
conversational source (no boundary early) still lands on a real sentence end
|
||||
rather than a dangling clause. Only if the source has no boundary at all in
|
||||
that window does it fall back to a best-effort trim with the trailing dangling
|
||||
conjunction/article stripped — a partial-clause tail is exactly what dots.tts
|
||||
regurgitates into its output.
|
||||
"""
|
||||
target_s = float(target_s) if target_s else 10.0
|
||||
hard_max = target_s + 4.0
|
||||
segs = list(model.transcribe(src, beam_size=5)[0])
|
||||
acc, best_end, best_txt = [], None, None
|
||||
for s in segs:
|
||||
if s.end > hard_max:
|
||||
break
|
||||
acc.append(s)
|
||||
txt = " ".join(x.text.strip() for x in acc).strip()
|
||||
if txt.endswith((".", "!", "?")):
|
||||
best_end, best_txt = s.end, txt
|
||||
if s.end >= min_s:
|
||||
break
|
||||
if best_end is not None:
|
||||
return best_end, best_txt or ""
|
||||
# no sentence boundary in-window — best effort, strip the dangling tail
|
||||
end = acc[-1].end if acc else 0.0
|
||||
words = " ".join(x.text.strip() for x in acc).strip().rstrip(",").split()
|
||||
while words and words[-1].lower().strip(",.") in DANGLING:
|
||||
words.pop()
|
||||
return end, " ".join(words)
|
||||
|
||||
|
||||
def trim_wav(src, dst, end_s):
|
||||
w = wave.open(str(src))
|
||||
sr = w.getframerate()
|
||||
frames = w.readframes(int(end_s * sr))
|
||||
w.close()
|
||||
o = wave.open(str(dst), "w")
|
||||
o.setnchannels(1)
|
||||
o.setsampwidth(2)
|
||||
o.setframerate(sr)
|
||||
o.writeframes(frames)
|
||||
o.close()
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 2:
|
||||
sys.exit("usage: derive.py <engine> [voice ...]")
|
||||
engine = sys.argv[1]
|
||||
manifest, engines = load()
|
||||
if engine not in engines:
|
||||
sys.exit(f"unknown engine '{engine}'; have {list(engines)}")
|
||||
prof = engines[engine]
|
||||
names = sys.argv[2:] or list(manifest)
|
||||
|
||||
outdir = ROOT / "derived" / engine
|
||||
outdir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
model = None
|
||||
if prof.get("ref_sentence_bounded"):
|
||||
from faster_whisper import WhisperModel
|
||||
model = WhisperModel("base.en", device="cpu", compute_type="int8")
|
||||
|
||||
for v in names:
|
||||
vc = manifest[v]
|
||||
src = ROOT / vc["canonical"]
|
||||
dst_wav = outdir / f"{v}.wav"
|
||||
if prof.get("ref_sentence_bounded"):
|
||||
end, txt = sentence_bounded_trim(str(src), prof.get("ref_max_seconds") or 10, model)
|
||||
trim_wav(src, dst_wav, end)
|
||||
if prof.get("needs_transcript"):
|
||||
(outdir / f"{v}.txt").write_text(txt + "\n")
|
||||
print(f"{engine}/{v}: {end:.1f}s sentence-bounded | {txt}")
|
||||
else:
|
||||
# TODO: resample to prof['sample_rate'] when resample:true
|
||||
shutil.copy(src, dst_wav)
|
||||
if prof.get("needs_transcript"):
|
||||
(outdir / f"{v}.txt").write_text((ROOT / vc["transcript"]).read_text())
|
||||
print(f"{engine}/{v}: copied canonical ({vc.get('source_sr')}Hz) -> {dst_wav.name}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,40 @@
|
||||
# Per-engine reference requirements. derive.py reads this to turn a canonical
|
||||
# source + transcript into an engine-ready reference set under derived/<engine>/.
|
||||
#
|
||||
# Fields:
|
||||
# sample_rate native SR the engine wants
|
||||
# resample true = derive.py should resample to sample_rate
|
||||
# (NOTE: resample is a follow-up — see the resample TODO
|
||||
# in derive.py; dots resamples internally so it's false there)
|
||||
# needs_transcript engine requires a per-reference transcript file
|
||||
# ref_max_seconds cap on derived reference length (null = uncapped)
|
||||
# ref_sentence_bounded transcript/clip must end on a sentence boundary (. ! ?)
|
||||
# — set for engines that leak reference content otherwise
|
||||
|
||||
engines:
|
||||
dots:
|
||||
description: "dots.tts (rednote-hilab) — continuous-AR 48kHz zero-shot clone"
|
||||
sample_rate: 48000
|
||||
resample: false # runtime auto-resamples at load; keep source SR
|
||||
needs_transcript: true # REQUIRED and must be accurate + sentence-bounded
|
||||
ref_max_seconds: 10
|
||||
ref_sentence_bounded: true
|
||||
notes: >
|
||||
Transcript accuracy AND sentence-boundary are load-bearing: a mismatched or
|
||||
mid-clause transcript makes dots regurgitate reference audio into the output.
|
||||
|
||||
chatterbox:
|
||||
description: "chatterbox-fast (Turbo) — streaming 24kHz clone"
|
||||
sample_rate: 24000
|
||||
resample: true
|
||||
needs_transcript: false # audio-only clone; server globs its refs dir live
|
||||
ref_max_seconds: null
|
||||
ref_sentence_bounded: false
|
||||
|
||||
zonos:
|
||||
description: "Zonos2 — expressive 44.1kHz clone + emotion dials"
|
||||
sample_rate: 44100
|
||||
resample: true
|
||||
needs_transcript: false
|
||||
ref_max_seconds: null
|
||||
ref_sentence_bounded: false
|
||||
@@ -0,0 +1,35 @@
|
||||
# Canonical voice corpus registry. One row per voice; the canonical clip + its
|
||||
# full transcript are the source of truth, engine-agnostic. derive.py reads this
|
||||
# together with engines.yaml to produce per-engine reference sets.
|
||||
|
||||
voices:
|
||||
donut:
|
||||
canonical: canonical/donut.wav
|
||||
transcript: transcripts/donut.txt
|
||||
source_sr: 44100
|
||||
duration_s: 16.3
|
||||
character: "sassy fairy-charm kid"
|
||||
provenance: "cloned from the 65-frost Booth bundle (2026-08)"
|
||||
|
||||
glados:
|
||||
canonical: canonical/glados.wav
|
||||
transcript: transcripts/glados.txt
|
||||
source_sr: 16000
|
||||
duration_s: 25.0
|
||||
character: "GLaDOS — flat, deliberate, menacing-cheerful"
|
||||
provenance: "Portal GLaDOS lines"
|
||||
warning: "LOW-SR source (16kHz) — upgrade the canonical clip if a cleaner GLaDOS source surfaces"
|
||||
|
||||
emmie:
|
||||
canonical: canonical/emmie.wav
|
||||
transcript: transcripts/emmie.txt
|
||||
source_sr: 24000
|
||||
duration_s: 19.3
|
||||
provenance: "Zonos clone added 2026-07-17; origin unrecorded"
|
||||
|
||||
miranda:
|
||||
canonical: canonical/miranda.wav
|
||||
transcript: transcripts/miranda.txt
|
||||
source_sr: 24000
|
||||
duration_s: 16.3
|
||||
provenance: "Zonos clone added 2026-07-17; origin unrecorded"
|
||||
@@ -0,0 +1 @@
|
||||
This is just not acceptable, Carl. I like my butterfly charm. It makes it so fairies like me, and it is pretty. It's part of my fit. I don't want to take it off. I don't see why I can't just wear two charms at the same time. Stupid angel of the caucus spaniel had like four or five tags.
|
||||
@@ -0,0 +1 @@
|
||||
I think I mentioned but I read your book because my my dear friend Nupa told me that I should and every now and again I would see you come up. I don't know. I take my job seriously I guess and so interviews to me felt a lot like chess and it required so much energy.
|
||||
@@ -0,0 +1 @@
|
||||
Welcome to test chamber 4. You're doing quite well. Once again, excellent work. As part of a required test protocol, we will not monitor the next test chamber. You will be entirely on your own. Good luck! As part of a required test protocol, our previous statement suggesting that we would not monitor this chamber was an outright fabrication. Good job! As part of a required test protocol, we will not monitor the next test protocol.
|
||||
@@ -0,0 +1 @@
|
||||
It's great. I mean, it's definitely comforting to go back to Australia when I come from there. So, you know, I get to see my parents, I get to see my friends and hang out. And I know the city really well because this was my fourth movie that I did in...
|
||||
Reference in New Issue
Block a user