memory: correct Fish-cloning finding — Fish clones competently (ECAPA 0.79); 'not British' was the undefined bug, not Fish/transcript
This commit is contained in:
+33
-19
@@ -118,13 +118,16 @@ _As of 2026-06-01:_
|
|||||||
`01KT2K2SY9N7AY69R9V0B4RXSW` → asset-engine-dev). **Workaround until fixed:
|
`01KT2K2SY9N7AY69R9V0B4RXSW` → asset-engine-dev). **Workaround until fixed:
|
||||||
click the dropdown (fires change) or hit the API directly.** My half
|
click the dropdown (fires change) or hit the API directly.** My half
|
||||||
(`blendable: false` catalog flag) is queued — see Recent decisions.
|
(`blendable: false` catalog flag) is queued — see Recent decisions.
|
||||||
- **Fish cloning verdict UNRESOLVED — re-test parked.** My in-session "Fish is
|
- **Fish cloning RESOLVED — it clones competently; "not British" was the
|
||||||
a weak cloner" conclusion was confounded (bogus reference transcript + maybe
|
undefined bug.** ECAPA-TDNN re-test (2026-06-01): an Imogen-referenced Fish
|
||||||
`reference_id: "undefined"` from the form bug + weak resemblyzer encoder) —
|
clone scores **~0.79 cosine vs the real `Imogen.wav`** (firmly same-speaker)
|
||||||
see Tried and abandoned. Literature says Fish/OpenAudio is SOTA cloning.
|
and only **~0.10 vs Fish's no-reference default voice**. So Fish faithfully
|
||||||
Whisper has the real transcript ("If the red of a second bow…"). Clean
|
clones Imogen WHEN it receives the reference. The operator's "Imogen doesn't
|
||||||
re-test (correct transcript + ECAPA scoring) ready when wanted. Future voice
|
sound British at all" was the `"undefined"` select bug feeding Fish its
|
||||||
clones are operator-handled (pitch-shift deepening abandoned).
|
DEFAULT voice (≈ a different speaker), NOT a Fish cloning failure. Earlier
|
||||||
|
resemblyzer-based "weak cloner" verdict RETRACTED. No Fish-side fix needed —
|
||||||
|
the blocker is entirely the undefined bug (asset-engine-dev). Future voice
|
||||||
|
clones operator-handled (pitch-shift deepening abandoned).
|
||||||
- **On-host consented voice library** — ~992 real-person clips cached in the
|
- **On-host consented voice library** — ~992 real-person clips cached in the
|
||||||
kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK
|
kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK
|
||||||
(CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS +
|
(CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS +
|
||||||
@@ -151,6 +154,15 @@ _As of 2026-06-01:_
|
|||||||
|
|
||||||
## Recent decisions
|
## Recent decisions
|
||||||
|
|
||||||
|
- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting
|
||||||
|
the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79
|
||||||
|
cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default;
|
||||||
|
transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves
|
||||||
|
identity (affects pronunciation, not timbre). Root cause of "Imogen sounds
|
||||||
|
nothing like British" = the `"undefined"` select bug feeding Fish its default
|
||||||
|
voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no
|
||||||
|
Fish-side fix needed. (Methodology lessons → Tried and abandoned.)
|
||||||
|
|
||||||
- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from
|
- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from
|
||||||
catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host
|
catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host
|
||||||
(`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed
|
(`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed
|
||||||
@@ -248,13 +260,14 @@ _25 older entries archived to archival-memory.md._
|
|||||||
optional `<name>.txt`) or inline base64 `references`. The catalog now uses
|
optional `<name>.txt`) or inline base64 `references`. The catalog now uses
|
||||||
`reference_id`.
|
`reference_id`.
|
||||||
|
|
||||||
- `[2026-06-01]` **Fish reference transcript as a provenance note** (not the
|
- `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY**
|
||||||
actual spoken words) — Fish/fish-speech uses the reference transcript to
|
(disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt`
|
||||||
disambiguate phonemes, and a misaligned transcript hurts cloning badly. I
|
for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778)
|
||||||
staged VCTK voices with provenance-note `.txt` sidecars, which CONFOUNDED my
|
/ empty (0.738) transcripts all clone Imogen about equally — the transcript
|
||||||
"Fish is a weak cloner" verdict (that verdict is retracted/unproven). Lesson:
|
affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds
|
||||||
stage the REAL transcript (VCTK ships ground-truth; or ASR via Whisper) for
|
like. The real culprit for "not British" was the `"undefined"` select bug, not
|
||||||
any clone reference.
|
the transcript. (A correct transcript still marginally helps pronunciation —
|
||||||
|
cheap to stage, not load-bearing.)
|
||||||
|
|
||||||
- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen
|
- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen
|
||||||
to contralto/mezzo) — Fish ignores small reference shifts and overshoots
|
to contralto/mezzo) — Fish ignores small reference shifts and overshoots
|
||||||
@@ -263,11 +276,12 @@ _25 older entries archived to archival-memory.md._
|
|||||||
Abandoned at every depth; all variants deleted. Finer independent
|
Abandoned at every depth; all variants deleted. Finer independent
|
||||||
pitch/formant control needs praat (not installed). Future clones = operator's.
|
pitch/formant control needs praat (not installed). Future clones = operator's.
|
||||||
|
|
||||||
- `[2026-06-01]` **resemblyzer for cloning-fidelity scoring** — weak/dated
|
- `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring** — its
|
||||||
encoder (2019 LSTM, English-biased) + a synthetic-vs-natural domain gap that
|
dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen
|
||||||
depresses cosine regardless of true fidelity → conclusions were soft. Use
|
clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish
|
||||||
ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) for rigorous
|
is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the
|
||||||
speaker-verification scoring.
|
same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to
|
||||||
|
default). Use ECAPA, not resemblyzer, for speaker-verification.
|
||||||
|
|
||||||
- `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream
|
- `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream
|
||||||
packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660
|
packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660
|
||||||
|
|||||||
Reference in New Issue
Block a user