memory: correct Fish-cloning finding — Fish clones competently (ECAPA 0.79); 'not British' was the undefined bug, not Fish/transcript

This commit is contained in:
vh
2026-06-01 15:26:58 -07:00
parent 7f9dc2dcf1
commit 3b54519d60
+33 -19
View File
@@ -118,13 +118,16 @@ _As of 2026-06-01:_
`01KT2K2SY9N7AY69R9V0B4RXSW` → asset-engine-dev). **Workaround until fixed: `01KT2K2SY9N7AY69R9V0B4RXSW` → asset-engine-dev). **Workaround until fixed:
click the dropdown (fires change) or hit the API directly.** My half click the dropdown (fires change) or hit the API directly.** My half
(`blendable: false` catalog flag) is queued — see Recent decisions. (`blendable: false` catalog flag) is queued — see Recent decisions.
- **Fish cloning verdict UNRESOLVED — re-test parked.** My in-session "Fish is - **Fish cloning RESOLVED — it clones competently; "not British" was the
a weak cloner" conclusion was confounded (bogus reference transcript + maybe undefined bug.** ECAPA-TDNN re-test (2026-06-01): an Imogen-referenced Fish
`reference_id: "undefined"` from the form bug + weak resemblyzer encoder) — clone scores **~0.79 cosine vs the real `Imogen.wav`** (firmly same-speaker)
see Tried and abandoned. Literature says Fish/OpenAudio is SOTA cloning. and only **~0.10 vs Fish's no-reference default voice**. So Fish faithfully
Whisper has the real transcript ("If the red of a second bow…"). Clean clones Imogen WHEN it receives the reference. The operator's "Imogen doesn't
re-test (correct transcript + ECAPA scoring) ready when wanted. Future voice sound British at all" was the `"undefined"` select bug feeding Fish its
clones are operator-handled (pitch-shift deepening abandoned). DEFAULT voice (≈ a different speaker), NOT a Fish cloning failure. Earlier
resemblyzer-based "weak cloner" verdict RETRACTED. No Fish-side fix needed —
the blocker is entirely the undefined bug (asset-engine-dev). Future voice
clones operator-handled (pitch-shift deepening abandoned).
- **On-host consented voice library** — ~992 real-person clips cached in the - **On-host consented voice library** — ~992 real-person clips cached in the
kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK
(CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS + (CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS +
@@ -151,6 +154,15 @@ _As of 2026-06-01:_
## Recent decisions ## Recent decisions
- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting
the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79
cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default;
transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves
identity (affects pronunciation, not timbre). Root cause of "Imogen sounds
nothing like British" = the `"undefined"` select bug feeding Fish its default
voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no
Fish-side fix needed. (Methodology lessons → Tried and abandoned.)
- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from - `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from
catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host
(`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed (`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed
@@ -248,13 +260,14 @@ _25 older entries archived to archival-memory.md._
optional `<name>.txt`) or inline base64 `references`. The catalog now uses optional `<name>.txt`) or inline base64 `references`. The catalog now uses
`reference_id`. `reference_id`.
- `[2026-06-01]` **Fish reference transcript as a provenance note** (not the - `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY**
actual spoken words) — Fish/fish-speech uses the reference transcript to (disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt`
disambiguate phonemes, and a misaligned transcript hurts cloning badly. I for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778)
staged VCTK voices with provenance-note `.txt` sidecars, which CONFOUNDED my / empty (0.738) transcripts all clone Imogen about equally — the transcript
"Fish is a weak cloner" verdict (that verdict is retracted/unproven). Lesson: affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds
stage the REAL transcript (VCTK ships ground-truth; or ASR via Whisper) for like. The real culprit for "not British" was the `"undefined"` select bug, not
any clone reference. the transcript. (A correct transcript still marginally helps pronunciation —
cheap to stage, not load-bearing.)
- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen - `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen
to contralto/mezzo) — Fish ignores small reference shifts and overshoots to contralto/mezzo) — Fish ignores small reference shifts and overshoots
@@ -263,11 +276,12 @@ _25 older entries archived to archival-memory.md._
Abandoned at every depth; all variants deleted. Finer independent Abandoned at every depth; all variants deleted. Finer independent
pitch/formant control needs praat (not installed). Future clones = operator's. pitch/formant control needs praat (not installed). Future clones = operator's.
- `[2026-06-01]` **resemblyzer for cloning-fidelity scoring** — weak/dated - `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring** — its
encoder (2019 LSTM, English-biased) + a synthetic-vs-natural domain gap that dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen
depresses cosine regardless of true fidelity → conclusions were soft. Use clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish
ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) for rigorous is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the
speaker-verification scoring. same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to
default). Use ECAPA, not resemblyzer, for speaker-verification.
- `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream - `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream
packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660 packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660