From 1dba05d8c8db1eaca80f42dfdbd0e3b308565f93 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 4 Jun 2026 01:08:50 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=202026-06-04=20i?= =?UTF-8?q?nfra=20session=20(phi4/vLLM,=20infra-ops=20identity,=20R15/R16,?= =?UTF-8?q?=20brokkr=20svc)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Captures: phi4-mini FP8 on ana-ml2 vLLM as nevermore summarizer (granite retired, canonical Phi-4 template after apply→revert); infra-ops NOPASSWD-sudo identity across PFI boxes; R15/R16 brokkr-smithy stand-up + ollama 0.30.4 + A6000 UUID-pin; brokkr user + substrate migration + audition service; homepage incident; observability roadmap. Foot-guns: headgun chat-template lesson, UUID GPU-pin, system-unit reaping, pyworld/ gfortran, homepage D-state. Archived 11 settled 2026-06-01 TTS-arc entries (Recent 5 + Tried 6) to archival-memory.md; pruned 145 lines of DONE 2026-06-03 in-flight bulk. File back under the 300-line cap (299). --- archival-memory.md | 97 +++++++++++++ persistent-memory.md | 315 +++++++++++-------------------------------- 2 files changed, 179 insertions(+), 233 deletions(-) diff --git a/archival-memory.md b/archival-memory.md index 100a854..29bc503 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -247,6 +247,53 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re `incident_esh_docker_nfs_boot_race`. _Archived 2026-06-03._ +- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting + the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79 + cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default; + transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves + identity (affects pronunciation, not timbre). Root cause of "Imogen sounds + nothing like British" = the `"undefined"` select bug feeding Fish its default + voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no + Fish-side fix needed. (Methodology lessons → Tried and abandoned.) + _Archived 2026-06-04._ + +- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from + catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host + (`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed + Maya/Miles demo is a fine-tuned, larger variant Sesame never open-sourced; + the open 1B base ships no usable voices, can't generate text, English-only, + no real-time streaming, absent from leaderboards, dominated by + Kokoro/Dia2/Fish/IndexTTS. (2) The upstream image is rotted (see Tried and + abandoned). asset-engine re-synced (`306aae9`). + _Archived 2026-06-04._ + +- `[2026-06-01]` **Chatterbox catalog v1→v2** (`93cedd6`) — re-point to the + devnen wrapper's `/tts` (was the thin OpenAI `/v1/audio/speech` exposing + none of the emotion knobs → poor output). Default `exaggeration: 0.5` per + Resemble's README ("works well for most"); the server SHIPS demo-tuned 1.3 + which over-acts — deliberately NOT adopted. 28 predefined voices via + `/get_predefined_voices` (default Emily.wav). Corrected seedable false→true. + _Archived 2026-06-04._ + +- `[2026-06-01]` **Fish-S2 catalog v1→v6** (`f3c260e`..`38d9e3b`) — was a + blank 2-field form; now exposes the full fish-speech ServeTTSRequest surface + + a 32-voice `reference_id` "Voice" dropdown (STATIC list — Fish has no + /voices API). `reference_id`-by-name works ONLY when each voice is staged as + a `references//` dir with `.wav` + `.lab` — was silently a + no-op until that layout fix on 2026-06-01 (see the two-root-causes entry + above); the inline-path-as-filesystem-path `references` shape 500s (see Tried + and abandoned). Corrected vram + 4→24 GB, cold_start 8→240s, seedable false→true. Added 3 consenting + British-female VCTK voices (Imogen=p225/Eleanor=p228/Beatrice=p229). + _Archived 2026-06-04._ + +- `[2026-06-01]` **`blendable: false` catalog flag — DEFERRED**, pending + asset-engine-dev honoring it in the form template. Clone-voice selects + (chatterbox/fish/dia2) must render as plain dropdowns, not the Kokoro blend + widget (blending reference clones is nonsensical). I ship the catalog flag + + schema field once they read it. Tracked: althing msg + `01KT2K2SY9N7AY69R9V0B4RXSW` to asset-engine-dev. + _Archived 2026-06-04._ ## Tried and abandoned (archived) @@ -532,3 +579,53 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re `incident_esh_docker_nfs_boot_race`. _Archived 2026-06-03._ +- `[2026-06-02]` **Naive cumulative re-decode for streaming** — decode the + growing token prefix each chunk + emit the delta, assuming the causal flow + gives a stable prefix. It does NOT (full-context attention, see above); + maxdiff 0.30. Also the high-level `s3gen.inference` re-applies `trim_fade` at + the start every call. Don't go this way. + _Archived 2026-06-04._ + +- `[2026-06-01]` **CSM bring-up** — upstream `phildougherty/sesame_csm_openai` + Dockerfile pins NO `huggingface_hub` version, so it now resolves to 1.17.0 + where `huggingface-cli` is fully removed (replaced by `hf`) → the build dies + at `huggingface-cli login`/`download`. Passing the HF token as a build-arg + got the token in (the model downloads at BUILD time, not runtime — our + compose only wired it as runtime env, the original failure) but then hit the + rotted CLI. Fixing would need vendoring + patching the Dockerfile; not worth + it for a weak model → CSM removed. + _Archived 2026-06-04._ + +- `[2026-06-01]` **Fish path-form references** (`{audio:"/app/references/x.wav"}`) + → HTTP 500 on this build, despite the stack README showing that shape. The + working clone paths are `reference_id` (by name → resolves the DIRECTORY + `references//` containing `.wav` + `.lab`; the flat + `references/.wav` is ignored — see the two-root-causes entry above) or + inline base64 `references`. The catalog uses `reference_id`. + _Archived 2026-06-04._ + +- `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY** + (disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt` + for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778) + / empty (0.738) transcripts all clone Imogen about equally — the transcript + affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds + like. The real culprit for "not British" was the `"undefined"` select bug, not + the transcript. (A correct transcript still marginally helps pronunciation — + cheap to stage, not load-bearing.) + _Archived 2026-06-04._ + +- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen + to contralto/mezzo) — Fish ignores small reference shifts and overshoots + large ones; `-3st` WITHOUT formant preservation → "baritone through a tuba" + (formants collapsed); `-2/-3st` WITH `--formant` → too subtle / still bad. + Abandoned at every depth; all variants deleted. Finer independent + pitch/formant control needs praat (not installed). Future clones = operator's. + _Archived 2026-06-04._ + +- `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring** — its + dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen + clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish + is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the + same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to + default). Use ECAPA, not resemblyzer, for speaker-verification. + _Archived 2026-06-04._ diff --git a/persistent-memory.md b/persistent-memory.md index ece03a8..fd8ae7a 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-06-03_ +_Last updated: 2026-06-04_ ## Repo purpose @@ -88,150 +88,36 @@ Sister repos (separate gitea repos, deployed by playbooks here): ## Current state / in-flight -_As of 2026-06-03:_ +_As of 2026-06-04:_ -- **INFRA SESSION 2026-06-03 — NH3 egress proxy, R14 provisioning, yt-voice-clipper CD.** - Five threads landed (detail in the cited auto-memories + runbooks + commits): - - **NH3 egress proxy (DURABLE).** dante SOCKS5 on nh3-dev `socks5h://10.100.10.50:1080`, - internal-only (client ACL `10.100.0.0/16`), systemd. NH3 **residential** egress - (~70.230.226.88) for colo services gated on their **datacenter IP**. - `scripts/setup-nh3-egress-proxy.sh` + `docs/runbooks/nh3-egress-proxy.md` (`f27ee47`); - auto-memory `reference_nh3_egress_proxy`. - - **yt-voice-clipper on irv-ml1** (`/opt/docker/compose/yt-voice-clipper`, :8000 - console+API + bgutil `pot-provider` sidecar). yt-dlp via `YTVC_PROXY` egress beats - YouTube's datacenter bot-gate (proven end-to-end). Console + audition + trim-by-playhead - deployed; Homepage label (AI Systems) in the override. **Push-to-deploy LIVE**: gitea - webhook → HMAC listener irv-ml1:9008 → git-clone deploy dir rebuilds on push to main. - `docs/runbooks/ytvc-autodeploy.md` (`d4f180d`); auto-memory `reference_ytvc_autodeploy`. - - **R14 VC-NVV tracer (brokkr-smithy)** provisioned scope-(a) on irv-ml1: box + CUDA env + - 4 engine run-envs (knn-vc/seed-vc/knn-svc/RVC, A6000-validated; one proven on a real - VIVAE conversion) + `/mnt/smithy` nfs4 mount + VIVAE-reachability. brokkr wires arms.py + - runs. Env at `~/r14`; auto-memory `reference_irv_ml1_gpu_r14`. OPEN: irv-ml1 git-pull - deploy key for `vh/brokkr-smithy` pending brokkr authorizing the pubkey (tracked in the - R14 althing thread). - - **gitea `[webhook] ALLOWED_HOST_LIST = external, 10.100.0.0/16`** on the gitea container - (ana-docker) — the anti-SSRF default `external` was denying webhooks to the private WG net, - blocking the ytvc auto-deploy. Backup `app.ini.bak-pre-webhook`. - - **irv-ml1 disk 97%→49%** — cleared all ollama models (~194 GB) + 2023-era - `/bootstore/llm-models/Storage/` LLMs (~472 GB; 2 llmuser-owned dirs needed operator sudo). +- **INFRA SESSION 2026-06-04 — phi4 on vLLM, infra-ops sudo identity, R15/R16, brokkr svc.** + Major threads (detail in cited auto-memories + commits): + - **phi4-mini FP8 LIVE on ana-ml2 vLLM** (`vllm-phi4`, :8004, GPU 1, 50K ctx, FP8 + FP8-KV) + as the **nevermore** summarizer/dreaming agent — superseded **granite-4-small** (removed + from llama-swap config; GGUFs kept on disk). Template = **CANONICAL/official Phi-4** (final, + after an Ollama-matching override applied `90e08f0` then reverted `27eb537`). nevermore + repointed (LLAMA_SWAP_URL→:8004, MODEL→phi4-mini). auto-memory `reference_ana_ml2_vllm_phi4`. + OPEN: brokkr re-baselining R15 P02 under canonical (eval re-regresses ~0.78→~0.45-0.57; not + live — worldtree #252 is baseline-first); **GPU 1 tight (~10 GB free — pin llama-swap to + GPU 0 as follow-up)**; Worldtree Vili #253 (hardcoded :9292 granite fallback, their fix). + - **infra-ops NOPASSWD-sudo identity commissioned** across PFI boxes — `ssh infra-ops@`, + key `~/.ssh/infra-ops_ed25519`. Bootstrap `playbooks/bootstrap-infra-ops-user.yaml` + + `scripts/bootstrap-infra-ops-fleet.sh` (tiers 1+2 live: irv-ml1/ana-ml2/ana-docker/nh3-docker/ + ana-nas + 4 app VMs; **esh-docker-vm added by operator override**). Excludes sf-*/corviduo/ + Synology. auto-memory `reference_infra_ops_sudo_identity` (`8c32a05`). + - **R15/R16 brokkr-smithy harness stood up on irv-ml1** — ollama upgraded **0.9.0→0.30.4** + (Ministral-3 needs it), A6000 **UUID-pinned**; R15 ollama arms (granite4.1:3b/8b, qwen3:4b, + phi4-mini:3.8b, hf SmolLM3-GGUF, ministral-3:3b-instruct, nuextract:3.8b) + R16 R/soundgen + + pyworld venv. `playbooks/irv-ml1-r15-r16-{nosudo,sudo}.yaml`. auto-memory `reference_irv_ml1_gpu_r14`. + - **`brokkr` user created on irv-ml1** (no-sudo) + 24 GB R14/R15/R16 substrate migrated out of + lkraven's home → `/home/brokkr/`; gitea pull = read-only deploy key. **brokkr-audition.service** + (SYSTEM systemd unit, :8137) serves brokkr's R16 NVV audition UI (the morph set being auditioned). + - **homepage incident (esh-docker-vm)** — wedged on dead siteMonitor IP (retired ESH firewall + 10.0.250.1) into unkillable D-state; host reboot cleared it; ESH-Firewall widget removed from + `services.yaml`. (`incident_esh_docker_nfs_boot_race` updated.) + - **Observability roadmap** `docs/roadmap.md` — Langfuse (full req/resp tracing) + + Prometheus/Grafana off vLLM `/metrics`. Deferred; Langfuse first. -- **chatterbox-fast — DONE** (was the primary focus). Deployed irv-ml1 :8197, cataloged - `ready`, EXTRACTED to its own MIT repo `vh/chatterbox-fast` (`547a557`) — the eshpfi stack - is now a pointer; auto-memory `reference_chatterbox_fast_repo`. Adaptive buffer-ratchet - chunking, TTFB ~0.5s, A6000 fp32 5.34 GB. _(Historical build detail follows.)_ Operator-authorized - high-effort build (incl. custom container from source). **Plan-of-attack: - `docs/design/chatterbox-fast-plan.md`** (durable; carries the full executable - detail). The old `docs/design/chatterbox-fast.md` was KILLED 2026-06-01 — its - §5 windowed-token-streaming was the abandoned native arc; superseded by the - adaptive-chunk plan. **DEPLOYED 2026-06-02** — Phases 1-3 done: - `stacks/chatterbox-fast/` (scheduler + app + sim tests + Dockerfile/compose), - live on **irv-ml1 :8197** (container `chatterbox-fast`, healthy) BESIDE the live - `chatterbox` :8196. Measured: **TTFB ~0.5s** (vs ~5.2s one-shot), no starvation, - RTF ~3.9-4.0× on A6000. **GPU = A6000 (device 1) ONLY** — turbo loads **fp32** - (NOT the ~2.5GB fp16 old notes claimed), measured **5.34 GB**, so the 3090's - ~3.8GB free does NOT fit. Perf: TF32/SDPA on (don't move TTFA — AR-decode-bound; - bf16 deferred, no clean fp32→bf16 cast path). Context-priming for join coherence - TRIED + REVERTED (discard-cut leaks prefix — see Tried/abandoned); coherence loss - at joins is an ACCEPTED limitation (operator: cold streaming "really good"). - Remaining: Phase 4 catalog entry + cutover; optional deploy-playbook smoke gate; - optional §5 base-fork 3-way A/B. - - **Goal:** cut time-to-first-audio from ~2.5s → sub-second via streaming, - keep turbo quality. Workload = single-stream interactive. - - **Chosen approach = adaptive buffer-ratchet chunking** (operator's idea): - generate the FIRST sentence alone for instant response (~0.66s first-audio - measured), then while it plays, generate the largest sentence-aligned chunk - that fits in ~80% of the bought playback time; chunks grow ~3× each (driven - by Chatterbox's ~3.8× realtime headroom), so after 2-3 chunks the rest of - the paragraph is one big chunk with near-full context. Context loss confined - to 2-3 joins at natural sentence-pause boundaries. THIS is "sentence-level - done right" and supersedes naive per-sentence splitting (which loses - cross-sentence prosodic context = real quality loss, operator-corrected). - Only works because RTF>1 (Fish at <1× realtime starves — can't use this). - - **Native frame-level streaming on turbo = ABANDONED** (see Tried/abandoned). - - **Also to build for A/B:** base-chatterbox + davidbrowne17 streaming fork - (true frame-stream but base-model quality) — not yet installed. - - **A/B samples** (GLaDOS voice) on nh3-dev `~/chatterbox-ab/`: sentence-level - turbo, chunked-native (artifacty), chunked-oneshot. - - **Dev/test pattern:** one-off GPU-1 container from `local/chatterbox:v1`, - mount `/worktank/chatterbox/{cache,reference_audio}` + the test script; - `docker exec -i chatterbox python -` for lib introspection. lkraven is in - the `docker` group on irv-ml1 — NO sudo needed for docker. -- **TTS fleet on irv-ml1 reworked this session; asset-engine catalog now - 17 services** (CSM removed). Canonical `docs/asset-engine/services.yaml` - + vendored `vh/asset-engine` both pushed (eshpfi `38d9e3b`; asset-engine - `0b552bd`). Per-engine live state: - - **dia2-2b (:8200) + dia2-1b (:8202)** — LIVE + healthy (restarted this - session; `local/dia:v2`, voices + cache intact). `status: ready`. - - **Chatterbox (:8196)** — LIVE. Catalog **v2**: routes to `/tts`, exposes - emotion levers (exaggeration 0.5, cfg_weight 0.5, temperature 0.8), 28 - built-in predefined voices (default `Emily.wav`) + clone. Fast + competent - cloner (it audibly cloned Imogen). Plain `Imogen.wav` uploaded to its clone - dropdown (`/get_reference_files` now: Imogen.wav + glados_25s.wav). - - **Fish-S2 (:8195)** — LIVE but **HEAVY**: ~240s torch.compile warmup on - cold start, ~24 GB on the A6000; first call after any restart eats the - warmup. Catalog **v6**: full `/tts` ServeTTSRequest surface - (temperature/top_p/repetition_penalty/latency/seed/format) + `reference_id` - is a 32-voice static "Voice" dropdown (default Emily) + advanced base64 - clone. NOT OpenAI-compat (only `/v1/tts`, no `/openapi.json`, no `/voices` - API). Research/non-commercial license. - - **zonos (:8203)** down (adapter built, not deployed). **ComfyUI :8188** - catalog-deferred link. -- **"undefined" select bug — RESOLVED.** asset-engine-dev shipped form-select - hardening (v0.1.9/.10) + a durable **per-field enable-toggle** (`togglable`, - v0.1.14/.16) — the real fix for the "form submits untouched fields" family. - My catalog half: opted fish-s2 `references` into `togglable: true` - (`catalog_version 1→2`, schema gained the `togglable` boolean; `c98a12b`, - pushed to origin). The earlier `blendable` idea was superseded by `togglable`. - dia2/chatterbox clone fields deliberately NOT toggled (dia2 defaults clone-on - as its stable voice). asset-engine CI was briefly red because the commit was - local-only until pushed — lesson: push catalog commits promptly (their CI - drift-checks against the remote). -- **Fish "not British" had TWO independent root causes — BOTH now fixed.** - The ECAPA-TDNN re-test (2026-06-01) showed Fish's cloning *engine* works: an - Imogen reference scores **~0.79 cosine vs real `Imogen.wav`** vs **~0.10 vs - the no-reference default** — but that test fed the reference via the inline - base64 `references` path, which masked the real defect. **`reference_id`-by- - name (the path the catalog + form actually use) was silently a NO-OP.** Proof - (2026-06-01): `reference_id=Abigail`, `=Imogen`, and no-reference all returned - **byte-identical** audio (md5 `40b6ac73…`, same text+seed) → reference_id was - being ignored entirely, so every dropdown voice collapsed to Fish's default - speaker. Cause: fish-speech resolves `reference_id=` against the - **directory** `references//` (audio + same-basename `.lab`), but the - voices were staged as **flat** `references/.wav` + `.txt` with the - per-name dirs left **empty** → nothing resolved. **FIX (2026-06-01):** - populated `references//.wav` + `.lab` for all 32 voices; - re-test → Imogen/Eleanor/Beatrice/Abigail/no-ref all 5 **distinct**. - reference_id now works; British accents come through. Hardened: the - deploy playbook gained a normalize-layout step + an A/B smoke gate - (`reference_id` MUST change output) so the empty-dir regression can't ship - silently. So the "not British" saga was (A) the `"undefined"` select bug - (asset-engine, still belt-and-suspenders) AND (B) this reference_id no-op - (server-side, FIXED) — B alone would have kept every voice identical even - with A fixed. Earlier "reference_id-by-name is THE working path, verified" - claim RETRACTED. Pitch-shift deepening abandoned; future clones operator- - handled. -- **On-host consented voice library** — ~992 real-person clips cached in the - kyutai tts-voices repo (`/worktank/kyutai-tts/.../snapshots/.../`): VCTK - (CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS + - expresso (CC BY-NC). Source for future clone voices. British-female Southern - England speakers p225/p228/p229 staged into Fish as Imogen/Eleanor/Beatrice. -- **Parakeet ASR (:8765) is now UP** (brought online 2026-06-02 to transcribe - the glados clip; operator wants it kept online). CPU-only (`gpu_device_id: - null`), `restart=unless-stopped`, healthy. API: POST multipart `file` to - `/transcribe` → `{"text":...}`. Image `local/parakeet:sherpa-onnx-v2` - pre-built; `docker compose up -d` in `/opt/docker/compose/parakeet`. -- **NFS share for Smithy — DONE.** `/volume1/smithy` on nh3-nas → `/mnt/smithy` - on nh3-dev (single export, RW, map-to-admin, scoped to nh3-dev only, fstab - `_netdev,nofail,x-systemd.automount`). Subdirs `datasets/{raw,manifest-store, - derived,holdout,quarantine}`; rename-atomicity verified. Smithy wired - storage-roots.yaml + E2E-verified. **Backup DEFERRED** per operator (datasets - regenerable; raw + manifest-store are the irreplaceable surface for when it's - wired — durable tier should land OFF nh3-nas/cross-site, since the NH3 restic - repo lives on the same NAS as the source). -- **Worldtree healthy v0.29.13** (last-known); **Skaldsong v0.32.2** - (ana-docker:8300, Kokoro SSE streaming); **artemis-31b-v1i** live on - llama-swap + worldtree personal; **ttyd fleet driver seat** on nh3-dev - (RO :7681 / drive :7682 via Caddy) — all out of this repo, last-known. - **Disclosed-keys hygiene queue** — rotate at convenience: HF token `hf_HBl…` (lkraven's HF account) leaked into BuildKit logs during the CSM build attempt (logs shredded, never committed — low urgency); @@ -246,6 +132,30 @@ _As of 2026-06-03:_ ## Recent decisions +- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent; + granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from + Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`) + +- `[2026-06-04]` **phi4 ships the CANONICAL/official Phi-4 chat template, NOT Ollama's.** + Ollama's bundled template omits the system `<|end|>` — that flattered brokkr's R15 eval but is + the DIVERGENT scaffold (Dvalin: the system `<|end|>` is Microsoft's intended format). Applied an + Ollama-matching override then reverted — ship correct, not the benchmark quirk. (`90e08f0`→`27eb537`; + "headgun" lesson in Tried.) + +- `[2026-06-04]` **infra-ops NOPASSWD-sudo identity commissioned, scoped to PFI boxes** (+esh-docker-vm + by operator override) — so infra-ops completes DevOps end-to-end vs handing the operator sudo steps. + Dedicated key, sudo log_output, key-gated. (`8c32a05`) + +- `[2026-06-04]` **Worldtree demo/pinned/personal deploys are ALL CI/CD, not infra-ops** — a "deploy + vX.Y.Z" request to infra-ops is MISROUTED → point them back to their pipeline. The granite→phi4 + repoint: worldtree-dev self-served via their CI/CD (v0.30.10). (`d8d776c`) + +- `[2026-06-04]` **ollama upgraded 0.9.0→0.30.4 on irv-ml1** (Ministral-3 is a Dec-2025 model the + old engine refused); A6000 pinned by **UUID** not index (native fastest-first ≠ nvidia-smi PCI). + +- `[2026-06-04]` **`brokkr` user (no-sudo) on irv-ml1; R14/R15/R16 substrate moved to /home/brokkr.** + Persistent box services there need SYSTEM systemd units (see Tried). + - `[2026-06-03]` **yt-voice-clipper bot-gate fix = route yt-dlp through NH3 residential egress, NOT cookies/PO-token.** YouTube hard-flags the Irvine colo IP (LOGIN_REQUIRED on a public video even with no cookies). Cookies + the bgutil PO-token + deno JS-runtime all @@ -292,53 +202,37 @@ _As of 2026-06-03:_ a real transcript (ASR'd via Parakeet): the Portal "Welcome to test chamber 4" lines. -- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting - the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79 - cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default; - transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves - identity (affects pronunciation, not timbre). Root cause of "Imogen sounds - nothing like British" = the `"undefined"` select bug feeding Fish its default - voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no - Fish-side fix needed. (Methodology lessons → Tried and abandoned.) - -- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from - catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host - (`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed - Maya/Miles demo is a fine-tuned, larger variant Sesame never open-sourced; - the open 1B base ships no usable voices, can't generate text, English-only, - no real-time streaming, absent from leaderboards, dominated by - Kokoro/Dia2/Fish/IndexTTS. (2) The upstream image is rotted (see Tried and - abandoned). asset-engine re-synced (`306aae9`). - -- `[2026-06-01]` **Chatterbox catalog v1→v2** (`93cedd6`) — re-point to the - devnen wrapper's `/tts` (was the thin OpenAI `/v1/audio/speech` exposing - none of the emotion knobs → poor output). Default `exaggeration: 0.5` per - Resemble's README ("works well for most"); the server SHIPS demo-tuned 1.3 - which over-acts — deliberately NOT adopted. 28 predefined voices via - `/get_predefined_voices` (default Emily.wav). Corrected seedable false→true. - -- `[2026-06-01]` **Fish-S2 catalog v1→v6** (`f3c260e`..`38d9e3b`) — was a - blank 2-field form; now exposes the full fish-speech ServeTTSRequest surface - + a 32-voice `reference_id` "Voice" dropdown (STATIC list — Fish has no - /voices API). `reference_id`-by-name works ONLY when each voice is staged as - a `references//` dir with `.wav` + `.lab` — was silently a - no-op until that layout fix on 2026-06-01 (see the two-root-causes entry - above); the inline-path-as-filesystem-path `references` shape 500s (see Tried - and abandoned). Corrected vram - 4→24 GB, cold_start 8→240s, seedable false→true. Added 3 consenting - British-female VCTK voices (Imogen=p225/Eleanor=p228/Beatrice=p229). - -- `[2026-06-01]` **`blendable: false` catalog flag — DEFERRED**, pending - asset-engine-dev honoring it in the form template. Clone-voice selects - (chatterbox/fish/dia2) must render as plain dropdowns, not the Kokoro blend - widget (blending reference clones is nonsensical). I ship the catalog flag + - schema field once they read it. Tracked: althing msg - `01KT2K2SY9N7AY69R9V0B4RXSW` to asset-engine-dev. - -_32 older entries archived to archival-memory.md._ +_37 older entries archived to archival-memory.md._ ## Tried and abandoned +- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF — + the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the + model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline + -33pp type-F1 while valid_format held 1.0. An Ollama-matching `--chat-template` "fixed" it but was + the WRONG fix (the bundled scaffold is the divergent one). PRINCIPLE: serve each model's canonical + `tokenizer.apply_chat_template`, not the bundled template — bundled ones corrupt baselines. Verify + the applied prompt via vLLM `/tokenize`→`/detokenize`. (`90e08f0`/`27eb537`) + +- `[2026-06-04]` **GPU pin by INDEX is ambiguous on irv-ml1** — native CUDA orders fastest-first + (A6000=0) but nvidia-smi/docker use PCI order (A6000=1), so an index pin can land on the wrong + card. Pin by **UUID** (`CUDA_VISIBLE_DEVICES=GPU-…`); verify via nvidia-smi compute-apps. Check + loaded-model VRAM with `ollama ps` (Ministral-3 @ its 256K default ctx = ~30 GB; cap num_ctx). + +- `[2026-06-04]` **Persistent services on irv-ml1 need SYSTEM systemd units** — the box reaps + user-session processes on ssh disconnect, and `--user` systemd isn't reachable over non-login + ssh, so nohup/setsid/`screen -dmS`/`systemd-run --user` all die (even with `enable-linger`). Use + `/etc/systemd/system/`. + +- `[2026-06-04]` **pyworld needs `setuptools<81`** (imports the removed `pkg_resources`); and + **R/soundgen `-lgfortran` fails** on irv-ml1 because the default `gcc` is gcc-11 but only + gfortran-12 is present (libgfortran.so lives only in the gcc-12 dir) → install `libgfortran-11-dev`. + +- `[2026-06-04]` **homepage "crash" ≠ always NFS** — a wedged container in unkillable D-state + ("tried to kill container, but did not receive an exit event") can come from dead `siteMonitor` + widget targets (retired ESH firewall IPs) hanging the node event loop into `exit_mmap`, needing a + host reboot. Check homepage's siteMonitors against retired hosts. (`incident_esh_docker_nfs_boot_race`) + - `[2026-06-03]` **gitea webhook to a private IP is denied by `webhook.ALLOWED_HOST_LIST`** (anti-SSRF; default `external` blocks private/loopback). Symptom: delivery shows `dial tcp ...: webhook can only call allowed HTTP servers`. Fix = APPEND the target net to @@ -401,49 +295,4 @@ _32 older entries archived to archival-memory.md._ unsolved"; vLLM-turbo outputs noise; davidbrowne17 streaming fork is BASE-only). → Use adaptive-chunking instead. -- `[2026-06-02]` **Naive cumulative re-decode for streaming** — decode the - growing token prefix each chunk + emit the delta, assuming the causal flow - gives a stable prefix. It does NOT (full-context attention, see above); - maxdiff 0.30. Also the high-level `s3gen.inference` re-applies `trim_fade` at - the start every call. Don't go this way. - -- `[2026-06-01]` **CSM bring-up** — upstream `phildougherty/sesame_csm_openai` - Dockerfile pins NO `huggingface_hub` version, so it now resolves to 1.17.0 - where `huggingface-cli` is fully removed (replaced by `hf`) → the build dies - at `huggingface-cli login`/`download`. Passing the HF token as a build-arg - got the token in (the model downloads at BUILD time, not runtime — our - compose only wired it as runtime env, the original failure) but then hit the - rotted CLI. Fixing would need vendoring + patching the Dockerfile; not worth - it for a weak model → CSM removed. - -- `[2026-06-01]` **Fish path-form references** (`{audio:"/app/references/x.wav"}`) - → HTTP 500 on this build, despite the stack README showing that shape. The - working clone paths are `reference_id` (by name → resolves the DIRECTORY - `references//` containing `.wav` + `.lab`; the flat - `references/.wav` is ignored — see the two-root-causes entry above) or - inline base64 `references`. The catalog uses `reference_id`. - -- `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY** - (disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt` - for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778) - / empty (0.738) transcripts all clone Imogen about equally — the transcript - affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds - like. The real culprit for "not British" was the `"undefined"` select bug, not - the transcript. (A correct transcript still marginally helps pronunciation — - cheap to stage, not load-bearing.) - -- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen - to contralto/mezzo) — Fish ignores small reference shifts and overshoots - large ones; `-3st` WITHOUT formant preservation → "baritone through a tuba" - (formants collapsed); `-2/-3st` WITH `--formant` → too subtle / still bad. - Abandoned at every depth; all variants deleted. Finer independent - pitch/formant control needs praat (not installed). Future clones = operator's. - -- `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring** — its - dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen - clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish - is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the - same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to - default). Use ECAPA, not resemblyzer, for speaker-verification. - -_35 older entries archived to archival-memory.md._ +_41 older entries archived to archival-memory.md._