memory: correct Fish-cloning finding — Fish clones competently (ECAPA 0.79); 'not British' was the undefined bug, not Fish/transcript

This commit is contained in:
2026-06-01 15:26:58 -07:00
parent 7f9dc2dcf1
commit 3b54519d60
+33 -19
View File
@@ -118,13 +118,16 @@ _As of 2026-06-01:_
`01KT2K2SY9N7AY69R9V0B4RXSW` → asset-engine-dev). **Workaround until fixed:
click the dropdown (fires change) or hit the API directly.** My half
(`blendable: false` catalog flag) is queued — see Recent decisions.
- **Fish cloning verdict UNRESOLVED — re-test parked.** My in-session "Fish is
a weak cloner" conclusion was confounded (bogus reference transcript + maybe
`reference_id: "undefined"` from the form bug + weak resemblyzer encoder)
see Tried and abandoned. Literature says Fish/OpenAudio is SOTA cloning.
Whisper has the real transcript ("If the red of a second bow…"). Clean
re-test (correct transcript + ECAPA scoring) ready when wanted. Future voice
clones are operator-handled (pitch-shift deepening abandoned).
- **Fish cloning RESOLVED — it clones competently; "not British" was the
undefined bug.** ECAPA-TDNN re-test (2026-06-01): an Imogen-referenced Fish
clone scores **~0.79 cosine vs the real `Imogen.wav`** (firmly same-speaker)
and only **~0.10 vs Fish's no-reference default voice**. So Fish faithfully
clones Imogen WHEN it receives the reference. The operator's "Imogen doesn't
sound British at all" was the `"undefined"` select bug feeding Fish its
DEFAULT voice (≈ a different speaker), NOT a Fish cloning failure. Earlier
resemblyzer-based "weak cloner" verdict RETRACTED. No Fish-side fix needed —
the blocker is entirely the undefined bug (asset-engine-dev). Future voice
clones operator-handled (pitch-shift deepening abandoned).
- **On-host consented voice library** — ~992 real-person clips cached in the
kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK
(CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS +
@@ -151,6 +154,15 @@ _As of 2026-06-01:_
## Recent decisions
- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting
the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79
cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default;
transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves
identity (affects pronunciation, not timbre). Root cause of "Imogen sounds
nothing like British" = the `"undefined"` select bug feeding Fish its default
voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no
Fish-side fix needed. (Methodology lessons → Tried and abandoned.)
- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from
catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host
(`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed
@@ -248,13 +260,14 @@ _25 older entries archived to archival-memory.md._
optional `<name>.txt`) or inline base64 `references`. The catalog now uses
`reference_id`.
- `[2026-06-01]` **Fish reference transcript as a provenance note** (not the
actual spoken words) — Fish/fish-speech uses the reference transcript to
disambiguate phonemes, and a misaligned transcript hurts cloning badly. I
staged VCTK voices with provenance-note `.txt` sidecars, which CONFOUNDED my
"Fish is a weak cloner" verdict (that verdict is retracted/unproven). Lesson:
stage the REAL transcript (VCTK ships ground-truth; or ASR via Whisper) for
any clone reference.
- `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY**
(disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt`
for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778)
/ empty (0.738) transcripts all clone Imogen about equally — the transcript
affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds
like. The real culprit for "not British" was the `"undefined"` select bug, not
the transcript. (A correct transcript still marginally helps pronunciation —
cheap to stage, not load-bearing.)
- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen
to contralto/mezzo) — Fish ignores small reference shifts and overshoots
@@ -263,11 +276,12 @@ _25 older entries archived to archival-memory.md._
Abandoned at every depth; all variants deleted. Finer independent
pitch/formant control needs praat (not installed). Future clones = operator's.
- `[2026-06-01]` **resemblyzer for cloning-fidelity scoring**weak/dated
encoder (2019 LSTM, English-biased) + a synthetic-vs-natural domain gap that
depresses cosine regardless of true fidelity → conclusions were soft. Use
ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) for rigorous
speaker-verification scoring.
- `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring**its
dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen
clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish
is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the
same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to
default). Use ECAPA, not resemblyzer, for speaker-verification.
- `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream
packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660