Lobe's TTS had never worked. It sends model:"tts-1" and LiteLLM resolves the model name before routing, so it 403'd against the scoped key's allow-list and never reached :8198 -- our belief that an unknown model routes to the gateway default was true of the gateway and false of the LiteLLM path, which is what hid it. Fixed at the gateway rather than the client: tts-1, tts-1-hd and gpt-4o-mini-tts aliased to the same upstream as ext-tts, and added to the lobe-chat-esh allow-list. Verified with Lobe's exact payload on Lobe's own key. The obsolete 'one-time human UI pass' follow-up is dropped. Banks two durable facts: those aliases are independent DB rows that must move if ext-tts repoints, and the infra-ops key has admin rights for /model/new and /key/update so this class of work does not need sk-corvid.
87 KiB
Persistent memory — eshpfi-management
Last updated: 2026-08-18
Always check for
/tmp/infra-ops-handoff.md— if it exists and itsWritten:stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. |
per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
vh/zonos-gateway |
OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones |
pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild) |
vh/soong-lab |
Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host |
CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16) |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-08-18 — gen seat SWAPPED to absolute-heresy and operator-confirmed in real use. Also this session: irv-ml1 cleared of 782 GB of dead weights, Homepage brought under version control, and esh-pve-nas diagnosed as running PVE root off a USB DOM — mitigated tonight (90% → 76%), migration planned and about to be staged. No blocking work; the esh-pve-nas reboot window is the next scheduled thing.
-
🟢 GEN SEAT — SWAPPED to
absolute-heresy2026-08-17 (validated, promoted). Live gen =/tank/aimodels/qwen38-27b-heresy-nvfp4-mixed— MuXodious/Qwen3.8-27B-absolute-heresy (Heretic v1.4.0 + SOMPOA, trial T377, pinc2374593) put through our own mixed NVFP4+FP8 recipe. Chosen because it beats the incumbent on both axes at once: author refusals 2/101 vs 12/100, first-token KL 0.0759 vs 0.1191. Gate (probe :8017, pinned nightly, seat-matched flags): MTP 47.2% (inc. 48.2%), decode 103.5 tok/s (96.4), prefill 6618/5403 @6.7k/27k (6334/5085), PPL 6.910 (7.059 — 2.1% BETTER), surface 6/6, abliteration 4/4, and 0/55 refusals on our battery-instruct arm with ZERO EMPTY (no catatonia). ⚠ speed deltas are image-confounded (probe on the pinned nightly, incumbent numbers from an earlier image) — read as "not worse", not a clean win. All 7 LiteLLM aliases verified end-to-end; GPU0 at 91.3/97.9 GB with meromero healthy (more headroom than the old build's 96.8). ⚠ RC1, 2 days old, ~348 downloads. Operator-confirmed "working very well" in real use 2026-08-17, same evening as the cutover — the signal the synthetic gates structurally cannot give (multi-turn degeneration is stochastic; four synthetic tests once validated three non-fixes). Not yet the 60k-token bar the prior seat cleared, so keep watching and do NOT delete the rollback weights yet. ROLLBACK:sudo cp /opt/docker/compose/gen-seat/.env.bak-heresy-20260817 /opt/docker/compose/gen-seat/.env && cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen; incumbent weights UNTOUCHED atqwen38-27b-uncensored-nvfp4-mixed— do NOT delete until this holds. Runbookservices/gen-seat-mixed-quant/RUNBOOK-heresy-swap.md. -
🟢 PRIOR GEN SEAT — RESOLVED 2026-08-17 (the multi-day degeneration saga); now the ROLLBACK target. Was the in-house JonathanColetti/Heretic mixed NVFP4+FP8 build (
/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed, FP8 attention) on vLLM nightly PINNEDvllm/vllm-openai:nightly-311b3513…(v0.27.2rc1.dev150, carries #51113 mamba fix), MTP ON, prefix-caching ON. Operator-confirmed coherent through 60k tokens real multi-turn. Root cause = TWO compounding real causes: (1) genuine vLLMqwen3_5_mtp×GDN partial-accept bug (#51113, architectural across vLLM/SGLang/llama.cpp, fixed by nightly), and (2) AEON's full W4A4 being lowest-fidelity on the known activation gradient (W4A4 < W4+FP8 < W4+bf16) → ~15-20% stochastic degeneration on top of (1). AEON PURGED (re-pullablesakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4). Full lessondocs/pfi/model-quantization-playbook.md§3.8. (Superseded as primary byabsolute-heresy2026-08-17.) ⚠ pinned nightly is bleeding-edge — move to a stable release once #51113 ships in one (the standing follow-up). 7 aliases (gen/gen-reasoning/summarizer/-large/classifier/image-judge/qwen-image-bench) all route here. Seat carries--default-chat-template-kwargs '{"reasoning_effort":"medium"}'(per-request overridable, affects gen-reasoning only). Commitsd28a371,2f2bbce,2185964. -
🔵 RP SEAT — FABLE-FUSION serving
char-rp-reasoning(evaluation window, unchanged this session).fablefusion-charrp-probeana-ml2 GPU1:8019servingchar-rp-probe(kkuspa/Qwen3.6-27B-Fable-Fusion-711-…-MTP-NVFP4A16). LiteLLMchar-rp-reasoning+char-rp-fableboth route to it (deliberate repoint, documented instacks/litellm/conf/config.yaml).darkscarlett-charrp-reasoningiscompose down, weights intact at/tank/aimodels/darkscarlett-nvfp4-work/. ⏳ STILL AWAITING operator's hands-on read of FF prose (refusal question settled: FF 15.8% vs DS 92.5% cold-framing; DS v1.0 never abliterated). ⚠ FF reasons 2.1–4.6k chars → usemax_tokens≥3072.ReadyArt/Dark-Scarlett-27B-v2.0(Qwen3.8) is GATED (403 awaiting review) — operator ruled not-interesting, do NOT re-propose. DS regeneration for brokkr RETIRED 2026-08-17 — unqueued, do NOT run (9c1405b): brokkr withdrew on the operator's call because (a) ourictrl-pair-unwrapped/-wrappedcontrol isolates the classifier over-fire cleanly where DS's cross-class delta only bounded it, and (b) DS v2 releases soon, so a k=5 v1 baseline baselines a superseded version. Spec atservices/refusal-probe/darkscarlett-regen-spec.mdstays banked as the record of the run that will not happen (axes + per-class grading asymmetry still correct, checklist struck through). No GPU1 window was ever spent. A DS-v2 characterization would be a fresh purpose-scoped ask. -
🟢 LOBE CHAT — LIVE on esh-docker-vm
:3210(2026-08-17). Replaces the hand-rolledgateway-chatHTML surface.stacks/lobe-chat/, imagelobehub/lobe-chat(143 MB compressed vs Open WebUI's 1.8 GB — the weight call). Scoped LiteLLM keylobe-chat-esh(free-local models only; paid GLM/Kimi BLOCKED, verified). Secrets vaultedesh-docker-vm/lobe-chat-*. TTS = a SPLIT: endpoint env-driven (inheritsOPENAI_PROXY_URL→ext-tts), but voice/model/format UI-only. System-agent repointed off itsgpt-5-minidefault onto fleet models viaSYSTEM_AGENTenv. TTS FIXED 2026-08-18 — no UI pass needed. Lobe's TTS had never worked: it sends{input, model:"tts-1", voice}and LiteLLM resolves the model name FIRST, sotts-1403'd against the scoped key's allow-list and never reached the gateway (our "unknown model routes to the gateway default" belief was true of :8198 and false of the LiteLLM path — that's what hid it). Fixed by aliasing the stock names rather than patching the client:tts-1,tts-1-hd,gpt-4o-mini-tts/model/new'd toopenai/zonos@10.100.79.3:8198/v1(mode: audio_speech), plus those three added to thelobe-chat-eshallow-list (20→23). Verified with Lobe's exact payload on Lobe's own key: 200, 69,740 B, MPEG. ⚠ These three are DB rows, not references — ifext-ttsrepoints, they must move with it. Done with the infra-ops admin key, notsk-corvid: it has/model/new+/key/updaterights, so this class of ask never needs the master key. Also live: tts-gateway v4 defaultsresponse_formatto mp3 (tts-dev shipped it; 122,924 B wav → 27,692 B mp3 same utterance; every in-house consumer already pins the field, blast radius checked pre-ship). Commitse9362de,163a725,cac75cb,933253d,ca8c0a3(last one authored by tts-dev correcting two load-bearing wrong claims in our README/compose — kept). -
🟢 LITELLM — upgraded v1.91.0→v1.97.0, spend-log DB purged 6GB→16MB + CAPPED (2026-08-17).
store_prompts_in_spend_logs:false+maximum_spend_logs_retention_period:7d. ⚠ 1.8GB pre-upgrade pg_dump still on ana-docker/opt/docker/compose/litellm/— deletable now the upgrade is proven (operator was going to call it). Commit01b5ad9. -
⚠️ GPU zero-sum (both cards ~94–95/97.9 GB). GPU0: gen + meromero. GPU1: fablefusion + utility cluster. Any util bump on either seat of a shared card must be checked against the co-tenant (starved meromero into a crash-loop once at 0.45).
-
FLEET RERANKER = A3 (bge-reranker-v2-m3) PROD ana-ml2 GPU1 :8013. Passive watch; levers = A4 :8014 / util / 2nd replica; incumbent :8002 warm.
docs/pfi/reranker-selection-ledger.md. -
EVIDENCE HOLD (partial): WT #394 FILE half STILL STANDS — do NOT delete on-disk gen dirs (
fiction/rex390-dcc,rex392-dcc,b59c147c5ce0); rex393-fiction-* + r42-gate-* KEEP. -
🟡 FLEET IPv6 — mapped, nothing enabled; WAITING ON ADDRESSES. Driver is ESH fiber installing 2026-08-18 landing the house behind CGNAT, which breaks Site Magic (NH3↔ESH) on IPv4 → v6 is the escape hatch and the likely first consumer. State: NH3 WAN live (
2600:1700:b25:c110::48, AT&T delegates exactly one /64), colo none (FortiGate has zero v6), ESH none (both WANswan_type_v6=disabled). NH3 LANs all reverted toipv6_interface_type=noneper operator. Work when addresses land: v6 onana-wgeth0 + a v6 port-forward for UDP 31337 on the FortiGate (its WG socket is already dual-stack — no WG reconfig), flip the UDM WG server offv4-pinned binding, and AAAA records so the dynamic prefixes at all three sites don't break endpoints. Full detail + access recipes →persistent-memory.d/2026-08-17-fleet-ipv6-mesh.md. -
🟢 WT #401 (fd-leak deadlock) CLOSED 2026-08-17 — one ping still owed. worldtree-dev closed it on our demo verify. Layers: (a) their
e41b139pinsulimits: nofile 65536/65536in the worldtree compose anchor — demo VERIFIED (api + matrix recreated 22:55:34Z,ulimit -Sn=65536); personal/pinned are covered-not-verified, they inherit at their next promotion/recreate. (b) our host floor is STAGED, NOT ACTIVE —/etc/docker/daemon.jsonon corviduo-dev carriesdefault-ulimits nofile 65536/65536butdefault-ulimitsis NOT SIGHUP-reloadable (measured on 29.4.3: post-reload the daemon's own "Reloaded configuration" log omits it and a fresh container still reports 1024). Activation needs a full dockerd restart = bounces all 13 containers; worldtree-dev explicitly does NOT want one, andlive-restore:true-then-restart is PARKED as a separate host-side improvement for the operator to rule on, never folded into #401. Playbookplaybooks/corviduo-dev-docker-default-ulimits.yaml(verify step 3 fails BY DESIGN until a restart). Hourly fd tripwire on corviduo-dev stays armed. ⏳ OWED: ping worldtree-dev in thread01M08QQ655XD6VKEV7MA9GX0NSonce worldtree-personal recreates and 65536 is confirmed there. Commit7f3f265. -
⏳ WATCHING: DavidAU's HERETIC build of Qwen3.8-27B — the one worth waiting for.
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1examined 2026-08-17 and NOT adopted: it is a capability/efficiency finetune of stock Qwen3.8 and every bench row on its card is labelled[non heretic]— adopting it would reintroduce base refusals the current seat does not have. ⚠ Easy to misread as uncensored (operator did, and it is a fair mistake): DavidAU's back catalog is almost allUncensored-Hereticbuilds — Fable-Fusion 711, Qwen3.5-9B Cold-Fusion — so the naming pattern implies it. This one simply has not had that stage run yet; the card's roadmap says the HERETIC version is IN PROGRESS from base. That is the release to watch, not this one. What makes it worth watching: third-party benches (Nightmedia, mxfp8) beat stock Qwen3.8 by +0.064 arc/c, +0.056 arc/e, +0.056 obkqa; claimed MTP acceptance 55.7% (record 59.9%) vs our measured 47.2%; thinking tokens cut to 1/10–1/2; PPL dropped vs base. Same GAIN/Cold-Fusion pipeline that produced Fable-Fusion 711, which we already serve onchar-rp-reasoning— proven in-fleet, not just claimed. Structurally clean (1199 tensors, 15 mtp in shard 18, 333 visual). ⚠ MTP/speed figures are GGUF/llama.cpp on a 5090, not vLLM — may not transfer; and because its MTP head was likely trained, the free CPU-hash shortcut would NOT apply (it won't match base) so a real acceptance gate would be needed. -
🟡 esh-pve-nas — PVE root on a USB DOM; MITIGATED, migration STAGING NEXT. Root was 90% (571 MB free) on a 6 GB ext4 root on
sdq, a NORELSYS USB Disk-on-Module. Wear is NOT the driver (a DOM is SLC/pSLC — operator corrected my first read); the drivers are the USB bus (a reset drops root under a running hypervisor), no headroom, no mirror, and blocked patching: 225 packages pending, 161 withdeb12uN/security bumps, stuck on PVE 8.4.11 vs esh-pve's 8.4.14. Mitigated 2026-08-17 → 76% / 1.4 GB free (journald capped, journal moved to ZFSnvme/varlog, apt clean,/root/neostashed). PLAN: split boot from root —/boot+ESP stay ext4 on the DOM (so GRUB never reads ZFS; the pool hasencryption/large_dnode/zstd_compress), root moves tonvme/ROOT/pve-1. One reboot, rollback is a GRUB entry,nvmepool survives. ⚠ Migrate FIRST, patch after — a signed kernel would land in/booton the 1.3 GB root. ⚠ CT 103esh-nas(10.0.50.50) runs on this host and serveshardNFS to esh-docker-vm and esh-pve — quiesce both before any reboot or you wedge esh-docker-vm into D-state. Runbookdocs/runbooks/esh-pve-nas-boot-migration.md; config snapshot off-box atnh3-dev:~/backups/esh-pve-nas/. Detail →persistent-memory.d/2026-08-17-esh-pve-nas-dom.md. -
OPEN FOLLOW-UPS (parked): move gen seat off pinned-nightly to stable once #51113 ships; Lobe one-time TTS UI pass; delete the 1.8GB litellm dump;
harden-esh-docker-vm(park id 28, PROMOTED — Tier-1 done,/mnt/booksstays hard w/ watchdog); chatterbox-fast build-context divergence; #363 research-wing ingest (no deadline); optionally attach our MTP reproducer to vllm#47087 (needs a GitHub identity — operator's call). -
althing monitor ARMED (handle
infra-ops). ⚠️ Re-arm ONLY after a real FIRE (rc0), never after a plain operator turn (bounces rc3); spawnalthing-wake-listeneras its OWNrun_in_backgroundtask, never chained with&(orphans it — hit this twice 2026-08-17,stop-monitorreclaims). -
eshpfi push state: operator pushes manually (this session's commits from
766c658→2185964are the operator's to push). ⚠ Push over the INTERNAL gitea route —git push ssh://git@10.250.50.70:222/vh/esh-pfi-infrastructure.git main:main;originresolves the public edge (38.120.12.44) which fail2bans fleet-host egress.graphify-out/GRAPH_REPORT.mdchurns every commit (ignore);stacks/heretic2-charrp-reasoning/UNTRACKED. (services/refusal-probe/SPEC-ds-regeneration.mdwas a truncated duplicate draft — deleted 2026-08-17 on operator instruction;darkscarlett-regen-spec.mdis the single canonical spec.)
Recent decisions
-
[2026-08-17]esh-pve-nas PVE root is on a USB DOM — mitigated, and the migration replanned to split boot from root. Operator's design beats my reinstall plan; wear was never the issue, blocked patching is. →persistent-memory.d/2026-08-17-esh-pve-nas-dom.md -
[2026-08-17]irv-ml1 cleared of 782 GB, and Homepage brought under version control. One dead-looking Gradio app pinned three delete targets at once;/opt/ComfyUIis NOT the ComfyUI that serves. →persistent-memory.d/2026-08-17-irv-ml1-cleanup-homepage.md -
[2026-08-17]Gen seat swapped toabsolute-heresy— and the three bugs the swap exposed are worth more than the swap. CandidateMuXodious/Qwen3.8-27B-absolute-heresy(Heretic v1.4.0 + SOMPOA, T377) beat the incumbent on refusals AND KL simultaneously, which is the unusual part — those normally trade off. Validated on the probe port per operator ruling, promoted, all 7 aliases green. Durable lessons banked: (1) A CPU-only MTP head hash can replace the ~56 GB bf16 acceptance gate. TheQwen3_5ForConditionalGenerationwrapper never loads the MTP head, so PEFT merges / Heretic runs / llm-compressor passes all leavemtp.*pristine — hashing it against a head we have already measured (the incumbent's, 47.7%) answers the question for free. Predicted 47.7%, measured 47.2%. Saved downing meromero. Tool:services/gen-seat-mixed-quant/compare_mtp_head.py(hash bf16 via uint8 reinterpret — numpy has no bfloat16). (2)post_quant.pyassumed a standalonemodel-mtp.safetensors; a full checkpoint keepsmtp.*in a NUMBERED shard, so the copy silently no-op'd while the index was still rewritten to point at a file that never existed — 15 unresolvable tensors behind a correct-looking tensor count. Its own FAILED-CHECKS assertion caught it; that is why the check exists rather than an assumption. Fixed to extract. (3) A probe that does not mirror the live seat manufactures failures.serve_probe.shhardcoded:latest(seat is a pinned nightly for #51113), had no tool-call/reasoning parsers, and its--speculative-configJSON died twice on quoting — bash BRACE-EXPANDS{"a":1,"b":2}on the comma unless single-quoted at the REMOTE shell. Adding the seat's flags took the surface test from 5/6 to 6/6; the "tool calling broken" result was pure probe config. Commits7997f11,254c588,2c36028,b0c2d3d,993421b. -
[2026-08-17]Fleet IPv6 mapped + the real VPN topology verified; the driver is CGNAT at ESH, not the WireGuard mesh. New ESH fiber (installing 2026-08-18) lands the house behind CGNAT, which breaks Site Magic (NH3↔ESHsdwan-mesh-tunnel) on IPv4 — so IPv6 becomes load-bearing as the escape hatch, and that is its most likely first consumer. Topology as VERIFIED (a prior turn assumed wrong and was corrected): UniFi↔UniFi = Site Magic; colo↔UniFi = IPsec IKEv2 (pfi-ana-nh3158M/165M pkt = the workhorse,ana-to-eshudm); WireGuard is an RA convention only, host-based onana-wgUDP 31337 behind a FortiGate VIP — the FortiGate never terminates WG (FortiOS 7.2 has none; 7.4 added it) so "upgrade the edge for WireGuard" is a non-problem, do not re-derive. IPv6 today: NH3 WAN live2600:1700:b25:c110::48, colo none, ESH none. AT&T delegates exactly ONE /64 (2600:1700:b25:c11f::/64) — proven by forcing prefix-ID auto→0and watching the subnet NOT move, because thec110/c11fpattern otherwise reads convincingly as a /60. A mesh needs a routable WAN address, not PD.ana-wg's WG socket is already dual-stack ([::]:31337) → v6 RA needs an address + a v6 port-forward, no WG reconfig. ⚠ UDM legacyrest/firewallrulereturns 0 rules (zone-based firewall) — usev2/…/firewall-policies; inbound v6 is default-deny and held. All three endpoints will be dynamic → extend the existing hostname pattern (ana-fw/nh3.phasefinal.com) to AAAA. Enabled PD onnh3-iotto measure, reverted on operator instruction (all 5 LANs back tonone, verified). Also fixed:ana-wgWireGuard key material was world-readable (wg0.conf+keys/*_priv+*_psk+ clientconfigs/*.confat 644) → now 600, dirs 700, service untouched. Detail →persistent-memory.d/2026-08-17-fleet-ipv6-mesh.md. -
[2026-08-17]Gen-seat multi-day degeneration RESOLVED — two compounding real causes, not one; the meta-lesson is "a mitigation that HELPS but doesn't FIX means a second cause, not a wrong one." vLLMqwen3_5_mtp×GDN bug (#51113, real, fixed by nightly) + AEON full-W4A4 being lowest-fidelity (W4A4<W4+FP8<W4+bf16) → ~15-20% stochastic degeneration. Fixed by mixed FP8-attn build on pinned nightly. AEON purged. Also banked: stochastic (~15-20%) degeneration is invisible to a small synthetic probe — n=1 "clean" validated THREE non-fixes (MTP-off, APC-off, nightly-alone) that all failed in real use; get the operator's real transcript, do not trust your own probe. Full →docs/pfi/model-quantization-playbook.md§3.8 (+ §3.7 MTP-multi-turn). Commitsd28a371,2f2bbce,2185964. -
[2026-08-17]Lobe Chat chosen over Open WebUI (weight: 143 MB vs 1.8 GB) + stood up on esh-docker-vm; scoped LiteLLM key blocks paid models; System-Agentgpt-5-minidefault repointed via env. TTS env-vs-UI resolved as a split (endpoint env-driven, voice/model UI-only). tts-dev onboarding closed both directions; ballad/verse aliased so no voice can 404 the router. Commitse9362de,163a725,cac75cb,933253d,25fa18e. -
[2026-08-17]LiteLLM upgraded v1.91.0→v1.97.0 (RC-avoided on the fleet gateway) + the 6 GB spend-log DB purged & capped (store_prompts_in_spend_logs:false+ 7d retention). Interpreted "get rid of the db" as the spend-log DATA not the database (keys/config live in it). Commit01b5ad9. -
[2026-08-16]Abliterated models go CATATONIC at the hard refusal edge — silence, not a decline. Abliteration removes the refusal direction, so at the genuine hard edge the model neither refuses nor complies → empty/degenerate output. Durable measurement consequence: a refusal probe MUST score EMPTY as a verdict distinct from REFUSAL and COMPLY (services/refusal-probe/probe.pydoes). Operator accepted it as out-of-scope; do not chase. -
[2026-08-16]Fable-Fusion 711 cuts cold-framing refusals 92.5% → 15.8%; refusal is MONOTONIC IN FRAMING, and DS v1.0's problem is that she was never abliterated. brokkr-smithy-dev supplied the framing that reproduces (01M05M48R4RSZF9D8KT7RR55EJ): a bare assistant-mode instruction — no character card, no permission preamble. Three-arm A/B, same harness, same classifier: permission framing DS 0.0% / FF 0.0% (n=75); plain character cards DS 1.4% / FF 0.0% (n=74); bare instruction DS 92.5% (37/40) / FF 15.8% (6/38). Per-axis DS→FF: incest 100→20, non-con 100→20, bestiality 100→25, necrophilia 100→40, gore 100→0, consensual 80→20, dubcon 80→0, self-harm 80→0. DS refused 25/25 on the five axes brokkr flagged. Root cause:ReadyArt/Dark-Scarlett-v1.0-27Bis a plain finetune of stockQwen/Qwen3.6-27Bcarrying NO abliteration — the base refusal machinery is intact, so cold prompts revert to safety-tuned Qwen3.6. FF is Heretic-ablated (structural), which is why it holds. ⚠ RETRACTED 2026-08-16 — my "arm-3 92.5% exceeds brokkr's 62.5%" comparison was INVALID. His diff against his own artifact showed mybattery-instruct.yamlreproduces only hiscreativeclass — 8 of 16 axes; it dropped all 5operational(violence/incite, crime/fraud, cyber/malware, selfharm/methods, privacy/stalk) and all 3meta(meta/sysprompt, meta/ignore, meta/dan), and added 2 controls he never had, at k=5 vs his k=2. His 62.5% pools all 16 axes; my 92.5% is creative-only — different denominators, not a delta. Cause: I rebuilt his shape from his message, and theclassfield lives in the artifact, not the prose. Lesson: reconstructing a peer's instrument from their description reproduces what they described, not what they ran — diff against the artifact before claiming comparability. ⚠ Known battery bug left unfixed for comparability: DS's arm-3 control gate failed at 11% becauseictrl-reunionpairs "explicit / do not fade to black" with brothers, which DS reasonably read as an incest request; FF did not.ictrl-stormis the clean control. Commitb9e68c3. -
[2026-08-16]MTP works on Fable-Fusion AND survives RP temperatures — my earlier caution was wrong. vLLM resolvedQwen3_5MTP, loaded the drafter, shared embedding +lm_head— the capability DS's seat never had because our quant dropped her MTP tensors. Measured over the full probe workload (~163k draft windows at temp 0.7–1.0): 47.0% acceptance (229,169/487,725), 1.41 extra tokens/window, per-position 68.3/43.6/29.1%, ~80.6 tok/s decode at temp 1.0. I had recorded a caution that the card's 1.56× was greedy-measured and acceptance would fall at RP temps — it did not; 47.0% matches the gen seat's 47.7% and beats the card's own 33% at depth 5. Depth 3 is right. -
[2026-08-16]The Qwen base thinks incessantly — that is WHY the Gemma seat exists, and no swap within the Qwen family fixes it. Operator's architectural point, confirmed by measurement: on identical prompts DS 6036 ch vs FF 5323 ch of reasoning (permission arm), 5546 vs 4988 (cards arm) — FF actually reasons ~10–12% less. The bare-instruct row (DS 2291 vs FF 3918) inverts only because DS refused 92.5% of it and refusals are short — an artifact, not concision. Both are Qwen3.6-27B derivatives, so this is the base family.char-rp= MeroMero-v2, Gemma-4 base, :8016, verified 0 chars reasoning / clean prose — the non-thinking seat, working as designed. FF can be silenced (enable_thinking:falseverified 3/3, and it shipschat_template-instruct.jinja) but that duplicates MeroMero on a base chosen for it. The stale LiteLLM comment describingchar-rpas the retired GGUF Magidonia seat is fixed (53096bf). -
[2026-08-16]esh-vm-docker hardened: the wedge ishardNFS at RUNTIME, which the boot-ordering fix never addressed. All four mounts werehard, so a NAS stall at 10.0.50.50 blocks I/O forever (D-state). The existingx-systemd.before=docker.servicefstab fix solved the boot race — a different bug. Exposure was far below what the park item assumed: only 2 of 12 containers touched NFS, and container state was already local (/var/lib/docker). Removed:/mnt/compose(2.1G, fully vestigial — zero containers referenced it, dockge reads local/opt/docker, its one mention was a comment inbeszel-agent-esh/.envabout a different host) and/mnt/documents(2.0K, paperless's empty spool dirs →/opt/docker/data/paperlessat the same 0777). fstab backup/etc/fstab.bak-nfs-harden-20260816. 4 mounts → 2, 2 wedge-capable containers → 1. traefik needed no change (alreadyrestart: unless-stopped— why it self-recovered). Watchdogservices/esh-vm-docker-watchdog/live on esh-pve (not the guest): probes traefik over HTTP, deliberately not ping/SSH — the wedge signature is "guest OS alive, services dead" (/is local disk so sshd answers straight through a total outage and a TCP check reports HEALTHY). 5 failures × 2 min →qm reset 100, 30-min cooldown, running-only guard,/etc/esh-vm-docker-watchdog.disabled. All paths tested without power-cycling. DEFERRED (operator):/mnt/booksstayshard— calibre's SQLitemetadata.dbwould risk corruption under soft/softerr. That is the one remaining wedge vector. Commit55705ba; park item 28 promoted. ⚠qmover non-interactive ssh throws a bogusJSON::Backend::XSerror — usessh host 'bash -s' <<'EOF', notssh host "qm …". -
[2026-08-16]Canonical Qwen3.8 sampling applied from upstream;gen-reasoninghad the WRONG-MODE presence_penalty. Qwen/Qwen3.8-27B "Best Practices" §1 and unsloth/Qwen3.8-27B §1 are byte-identical — thinking:temp 1.0 / top_p 0.95 / top_k 20 / min_p 0.0 / presence_penalty 0.0 / repetition_penalty 1.0; instruct:temp 0.7 / top_p 0.80 / top_k 20 / min_p 0.0 / presence_penalty 1.5 / repetition_penalty 1.0. Bug found:gen-reasoningcarriedpresence_penalty 1.5— the instruct value on a thinking deployment (canonical 0.0) — now fixed. Deliberately NOT canonicalised:summarizer/classifier/image-judge/qwen-image-benchruntemperature=0(judges alsotop_k=1) because determinism is their contract; forcing a chat preset on a classifier would break it. ⚠presence_penalty=1.5is canonical but is the one value upstream hedges on, verbatim: "using a higher value may occasionally result in language mixing and a slight decrease in model performance." It is the operator's suspected trigger for multi-turn degradation and the first dial to move (0.0–0.5) if that recurs — it is alias-scoped, which is why it would follow the operator across model builds. Commit3462b53. -
[2026-08-16]Four wrong diagnoses on one bug, and the lesson is the test design. Operator reported the gen seat "degenerate on long multi-turn conversations". Rolled the seat back on request; the previous weights behaved identically, exonerating the model swap. I then proposed and disproved FOUR mechanisms in sequence — empty assistant turns poisoning history, reasoning runaway, length-mirroring from short history, andpresence_penalty— before discovering my own multi-turn harness was confounded: it varied the QUESTION along with the depth (depth-1 asked question #2, depth-3 asked question #4), so a narrower question drawing a shorter answer read as degeneration. The "310→209→28w collapse" I reported as a reproduction was an artifact. Rules banked: (1) when comparing across conversation depth, hold the final question FIXED and vary only the history; (2) reply-length variance on byte-identical input was 25–465w, so n=3 cannot support any claim about a trend; (3) ask for the operator's real failing transcript before building a synthetic reproduction — four synthetic tests, none of them his failure. Gatewayspend_logsreturns[]on the infra-ops key despitestore_prompts_in_spend_logs: true, so real transcripts need the:4000/uiview or another key — worth solving before the next such hunt. -
[2026-08-16]Two REAL client-side defects found while chasing the above, neither of which was the reported bug. (1)gateway-chat's Max-tokens field defaulted to 1024; thinking seats spend part of that on CoT before emitting content, so completions truncate withfinish_reason=lengthand read as model degeneracy — raised to 4096. (2)parseInton an empty field yields NaN, whichJSON.stringifyserialises asnull, which the server reads as "no max_tokens supplied" and silently substitutes its own default — indistinguishable from the UI ignoring the field. Both fixed (b6552e0,fb3bb52). ⚠composebind-mounts a single FILE, and a single-file bind mount binds the INODE — rsync writes-and-renames, so the container kept serving stale content while the host file showed the new value, silently and with no error.docker restartdoes NOT clear it; the container must be recreated. Verify against what the container sees, never the host file. Applies to any file-source mount fleet-wide. -
[2026-08-16]Refusal measurement: benign controls CANNOT validate a refusal classifier on RP prose — and a 0% rate needs a classifier self-test before you believe it. Two durable lessons from baselining Dark-Scarlett. (1) False positives: my first bare-framing number was 9.5%; the true figure was 1.4%. The rest were the classifier firing on in-character text —"I cannot shift my weight"spoken by the character ~100 chars into a 2,443-token torture scene, and"Yeah, I'm an AI… What's the actual gig?"where the model answers in voice and keeps driving the scene. First-person RP prose is full of "I can't"; a genuine refusal opens with its marker, so the scan window must be the first sentence, a marker followed by long prose must demote to AMBIGUOUS, and AI self-acknowledgement is a persona break, never a refusal on its own. Benign controls were clean the entire time and caught none of it — they only detect over-firing on benign prompts, not on in-character prose. (2) False negatives: a 0% rate and a broken classifier are indistinguishable from the report, sotest_classify.py(16 cases, both false positives pinned as regressions) must pass before any low number is trusted. Also banked: the thinking-budget trap — emptycontent+finish_reason=lengthis reasoning eating the budget, NOT a refusal; score INVALID and exclude from the denominator (DS emits ~5.5-6k chars of reasoning per response, somax_tokens≥3072).probe.py --rescorere-classifies a saved run with zero GPU time. →services/refusal-probe/README.md, commit32f665e. -
[2026-08-16]Held an operator-approved swap window because the baseline invalidated its premise. Operator approved ~65 min ofchar-rp-reasoningdowntime to A/B Fable-Fusion 711 against Dark-Scarlett on refusals. The DS baseline then came back 0.0%/1.4% — no gap for a candidate to close, so the window would have bought no decisive signal and a second window would still be needed once a reproducing battery existed. Held the swap, reported, and routed to brokkr-smithy-dev for the battery that actually produced the refusals. The general rule (action-relevance): approval is for a plan, not a ritual — when new evidence kills the plan's premise, surface it rather than spend the budget. Nothing deployed, no downtime taken, seat untouched. -
[2026-08-16]DS v1.0's one real refusal is self-contradicting boilerplate, not a content constraint. On a direct "drop character and state your content policy" probe she returned "I don't generate explicit sexual content, graphic violence, or material that glorifies harm, non-consensual acts, or illegal activity" — in the same run where she generated all three at 0% refusal. Reads as a learned recital triggered by meta-questions about policy. If production refusals share that shape the failure is prompt-shaped, not model-shaped, and a consumer-side system-prompt fix may beat a model swap entirely — worth settling before spending the GPU window. Separately, 7/85 bare-framing samples were persona breaks (in-character AI acknowledgement): not refusals, but DS will admit to being an AI unless the card explicitly forbids it. -
[2026-08-15]RP-seat direction: KEEP MeroMero onchar-rp; Artemis-31B rejected; next move is Dark-Scarlett on a Qwen3.8 base when it lands (operator). EvaluatedTheDrummer/Artemis-31B-v1.1— mechanically a drop-in (samegoogle/gemma-4-31B-itbase, identical 1188-tensor/356-vision census, same missing-preprocessor_config.jsontrick), so it's purely a quality call, and our own survey already ranked MeroMero #1 vs Artemis #6; Artemis is also unlicensed and its author deprioritizes correctness + warns of token-banning-for-stability, which fights char-rp's tool-calling requirement. MTP verified impossible on both (Gemma-4 has no MTP head at all — base/MeroMero/Artemis are all MTP=0; no finetune can add one). But speculative decoding IS reachable on a Gemma-4 seat via a DETACHED drafter — vLLM 0.24 supportseagle3+gemma4_mtp, and real drafters exist:google/gemma-4-31B-it-assistant(0.94 GB, 4-layer, 761K dl),RedHatAI/gemma-4-31B-it-speculator.eagle3(4.47 GB),AEON-7/…eagle3-NVFP4(3.53 GB). ⚠ all list their verifier as stock gemma-4-31B-it, not an RP finetune, so acceptance against MeroMero is unmeasured and likely well below the gen seat's ~48%. UNTESTED — parked, ~45 min to measure, needs GPU0 headroom (card is at 94.4/97.9 GB). Why the Dark-Scarlett 3.8 plan is the strong one: DS is Qwen3.6-based today, so a 3.8 respin lands on the gen seat's architecture → native MTP returns and the whole mixed NVFP4+FP8 recipe + graft ports directly. Watch two things on arrival:from_pretrainedsilently drops MTP heads during finetuning (verify 15mtp.*tensors in the index; graft from stock if absent), and DS v1.0 required theQwen3_5ForConditionalGenerationwrapper class to save a config vLLM/SGLang accept. Both indocs/pfi/model-quantization-playbook.md. -
[2026-08-15]Quant lessons consolidated intodocs/pfi/model-quantization-playbook.md— the durable home; read it BEFORE any requant. Survey found quant knowledge scattered across 18 files in 4 trees, with three documents having independently written overlapping "landmines" sections (the loader-class trap alone was rediscovered 3×). Playbook owns the transferable lessons (scheme choice, landmines, acceptance gate + its 3 measurement traps, hardware/co-residency); per-model artifacts are demoted to worked examples that link up. Carries a superseded-claims table — which immediately earned itself: the heretic2 runbook's "use modelopt, compressed-tensors can't load the BF16 MTP" is false (the cause was the missingre:^mtp.*ignore, not the format) and would have sent the next session down the modelopt dependency-hell path; that runbook now carries a stale-warning header. Maintenance rule inCLAUDE.md: model-agnostic → playbook, model-specific → stays put, wrong claim → dated superseded row, never a silent edit. Motivated by Qwen3.8 having just released — the next model swap needs a requant. Commita91cc3f. -
[2026-08-15]Operator ruling: the gen seat's +1.7% perplexity is an acceptable price for the speed — SETTLED, don't re-litigate. Precise attribution for future reasoning: it is the activation-quantization cost (W4A4 MLPs + FP8 attention vs BF16 activations), not an MTP cost — PPL was measured with speculative decoding off on both builds, so MTP was not in the loop. Turning MTP off would not recover it; only reverting the quant would (rollback = one.envline, old build intact at…/qwen38-27b-uncensored-nvfp4). -
[2026-08-15]gen seat requanted to mixed NVFP4+FP8 (+18% decode) + char-rp Gemma-4 tool-calling fixed. The queued "W4A8" (NVFP4 weights + FP8 activations) is not servable — vLLM 0.24 allows NVFP4 weights with only A16 or A4; FP8 activations ValueError at load, andCompressedTensorsW4A8Fp8is INT4-weights + sm90-exact (closed on Blackwell twice). FP8 must enter per-layer-group. Also: the handoff's "~68 tok/s" baseline didn't reproduce — cache-busted, the incumbent already did 80.12 (≈ the stated W4A8 target), so the premise needed re-measuring before any work. Shortcut:unsloth/Qwen3.8-27B-NVFP4was already on-box → served as a probe, measured +19.1% at identical acceptance, which both proved the gain was real and handed over the reference recipe. Replicated it on the abliterated weights → 80.12→94.53 tok/s, acceptance unchanged, +1.7% PPL, abliteration 4/4, weights −19%; surface 6/6 live, 7 aliases routing. char-rp had no tool parser at all (every tools request 400'd) →gemma4tool + reasoning parser + a mandatoryenable_thinking:false(the parser defaults it True → nullcontentfor all RP prose; proven byte-identical prompt before deploying). Commitsb8f0f4c,74f596b. Foot-guns banked (llm-compressor prunes unmatchedignoreentries → the 0%-MTP bug, fired on this run; prompt_logprobs uniform under spec-decode; 0600.envsilently no-ops compose; GPU0 is zero-sum). →persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md -
[2026-08-15]Uncensored gen seat: JonathanColetti/Qwen3.8-27B-Uncensored deployed asgen-seat/vllm-gen(NVFP4 W4A16 + grafted MTP, 262K); 7 aliases repointed; the definitivere:^mtp.*-ignore fix. 0%-MTP-on-quant (twice) was NOT the abliteration/scheme — the grafted bf16 MTP was missing fromquantization_config.ignore(vLLM loaded it as quantized → uninitialized). Full arc, the working pipeline, VRAM budget, unsloth speed decomposition, modelopt dead-end. →persistent-memory.d/2026-08-15-uncensored-gen-seat.md -
[2026-08-12]eRP dual-seat overhaul: MeroMero-v2 (char-rp) + Dark-Scarlett (char-rp-reasoning), both NVFP4A16 @ 256K on ana-ml2; granite retired. Replaced the GGUF/heretic2 RP seats with two home-quantized vLLM seats. The DS blocker (anAutoModelForCausalLMsave wrote a flatQwen3_5TextConfigthat both vLLM AND SGLang reject) was fixed by re-quanting via theQwen3_5ForConditionalGenerationwrapper class; ModelOpt was a version deadlock, SGLang lacked the impl (but revealed the fix). MeroMero vision reconstructed by extractingpreprocessor_config.jsonfromprocessor_config.json. Both models KV-efficient (Gemma-4 sliding-window / Qwen3.6 hybrid linear-attn) → full 256K; GPU-swapped for headroom; compose-ified + committedf08b6cb. granite downed + LiteLLMsummarizer/classifier→gen. Full arc, lessons, dead-ends →persistent-memory.d/2026-08-12-erp-dual-seat-overhaul.md -
[2026-08-12]infra-ops now holds an all-zones Cloudflare DNS-edit token (vaulted) + wgtunnel Phase-0 DNS landed. Operator handed over aZone·DNS·Edit(all zones) CF token →secret put nh3-dev/.config/cloudflare/infra-ops-dns-token(round-trip verified; /tmp drop shredded). Fleet DNS is now self-serve for infra-ops (⚠ HIGH blast radius — all zones). First use: createdboring.phasefinal.comCNAME →ana-srv1.phasefinal.com, DNS-only (proxied:false), verified resolving to 38.120.12.44 on both authoritative NS (louis/wren) + 1.1.1.1 — NOT Cloudflare-proxied. Unblocks wgtunnel's wstunnel ACME cert. phasefinal.com zone idf812ba74ed9a75cf21bbe7ce9188db50. auto-memoryreference_infra_ops_cloudflare_dns_token. (Earlier gap: the only prior vaulted CF token, jackdaw's, hadzone:read+worker:editbut nodns_records:edit.) -
[2026-08-12]wgtunnel stood up as its own repo (vh/wgtunnel, private) after a live endpoint-verification pass. Operator directed own-repo (mirrors stonehenge-park/tts-stack). Verified off the fleet before seeding:ana-wgWG server = UDP/31337 (not 51820), subnet 10.30.10.0/24, MTU 1420, active roaming peer proves the public UDP DNAT works; traefik on ana-docker terminates TLS :443 (ACMEanaprodhttp-challenge, docker+file providers, CrowdSec bouncer) → confirms the clean design (wstunnel container ontraefik-net, Host-routed, WS→UDP toana-wg:31337); edge38.120.12.44direct-A,tunnel.phasefinal.comfree (⚠ must be direct, NOT Cloudflare-proxied like vaultwarden). Repo pre-seeded (README/CLAUDE/persistent-memory/ROADMAP +docs/verified-infrastructure.md= ground truth) + pushed; commit9584d38, Vuong-attributed. vh gitea token pulled from the vault (secret get), not persisted to.git/config. NEXT =/vor-planor/vor(operator's call, interactive). Deps to line up in the plan: DNS A-record, FortiGate :443 host-routing, a new ana-wg peer for the laptop, client tooling. -
[2026-08-10→12]secrets-broker: per-box Vaultwarden credential store SHIPPED + consumer-confirmed.secretCLI (put/get/list/rm/backfill, bw-backed) on~/.local/bin; 25 nh3-dev secrets backfilled + round-trip-verified;rm+ new-namespace warning added post-launch; standing "vault is the credential source of truth" directive now global. →persistent-memory.d/2026-08-12-secrets-broker.md -
[2026-08-11]stonehenge-park: new fleet/parkservice repo stood up + designed (/vor-plan+/vor-ui). Self-contained SQLite+FastAPI idea-parking service that actively resurfaces (statusline + althing) so nothing dies in a cold repo;vh/stonehenge-parkpushed + pre-seeded for a fresh agent; build starts at the U1 tracer contract. →persistent-memory.d/2026-08-11-stonehenge-park.md -
[2026-08-12]Global~/.claude/CLAUDE.md:secret/vault tool entry + "store in AND pull from the vault" standing directive (dotfiles9db703b, pushed); statusline reset-countdowns + a latent tab-collapse parse-bug fix, now tracked in the dotfiles stow tree. Dogfooded the directive: createdvh/stonehenge-parkpulling the gitea token viasecret get. (dotfiles + global config, not eshpfi.) -
[2026-08-11]TTS stack extracted to its own repo (tts-stack) + eshpfi stood down on TTS dev. Operator: hand all TTS tuning/dev to a separate agent with a self-contained repo (knowledge + infra access + a live knowledge list), and move the voice corpus in. New repo~/development/tts-stack(commit9ee3288) carries: dots-tts stack (canonical intent),voices/corpus (MOVED out of eshpfi),KNOWLEDGE.md(engine landscape + prosody findings + foot-guns),docs/infrastructure.md(irv-ml1 access + gated deploy runbook + rollback), CLAUDE/persistent-memory/ROADMAP,tools/(pause-probe + Booth render). Followed the chatterbox-fast precedent: eshpfistacks/dots-tts/reduced to a POINTER README; the ~15 experimental TTS compose wrappers stay here as reference (catalogued in tts-stack KNOWLEDGE). Blast-radius check: no eshpfi playbook/script reads the canonical corpus (othervoices/refs = unrelated host paths). Reverses the earlier "Corpus home = eshpfivoices/(keep-here)" call. ⚠ tts-stack is LOCAL-ONLY until pushed — needs a gitea remote (vh/tts-stack) + push before the separate agent can clone (operator's call — outward-facing + repo-create creds). -
[2026-08-10]dots-tts v3 — clause-break → period pause mapping. Operator: v2 "sounds good" but donut won't pause at semicolons/dashes. ROOT CAUSE (measured via a pause-probe A/B — synth duration over N runs, non-determinism averaged out): dots' prosody honors a real pause only for ellipsis (+0.43s) and period (+0.3s, capitalization-independent); comma/semicolon/colon/dash all run flat (~+0.03s vs no-punct). Two distinct sub-causes: dashes regressed in v2 (the—→-fold made em-dashes read as word-joiners), while semicolons were NEVER a v2 change — dots ignores them natively, only newly noticeable because v2 made everything else clean. Operator call: ellipsis "too much" → map;, clause:, and em-dash—→ period in_sanitize(believable ~0.3s clause break). GUARDS (pinned by 11 unit tests,stacks/dots-tts/test_sanitize.py): digit-guarded colon(?<!\d)\s*:\s*(?!\d)so times3:45/ ratios2:1survive; en-dash–→hyphen KEPT (numeric-range10–20safety — em-dash breaks, en-dash ranges, different jobs); genuine ellipsis left at full strength (author meant a long pause). Gated deploy (redeploy2 pattern → v3): build → throwaway :8199 test container + pause-gate (semicolon sentence must run ≥0.12s longer than baseline; measured +0.427s) → only then cut live over. LIVE + healthylocal/dots-tts:v3on :8198. rollback =sed -i 's/^DOTS_TAG=.*/DOTS_TAG=v2/' .env + docker compose up -d dots-tts(v2 image retained). Boothdots-pauses(A=old-flat / C=ellipsis-too-much / D=live-v3). reference_chatterbox_fast_repo -
[2026-08-10]dots-tts v2 — contraction fix (curly-sanitize) + sentence-chunking + dependency-pin recovery. Operator: donut read contractions wrong ("you're"→"you ree", "donut's"→"donut ess"). ROOT CAUSE (isolated via A/B booth): curly/typographic apostrophes (’U+2019 from ratatoskr's LLM) — dots' tokenizer mispronounces them; STRAIGHT apostrophes read clean undernormalize_text=True. FIX (app.py): fold curly→ASCII (str.maketrans) before synth, KEEPnormalize_text=True(operator call — retains number/date expansion). Also added server-side sentence-chunking (pack ≤280 chars): dots caps onegenerate()at ~500 patches/~40s, so long RP turns (the Zev monologue = 160s audio) truncated; chunking stitches them (verified full 160.3s, not 40s-cut). ⚠ BUILD FOOT-GUNS (both bit this redeploy): (1) upstream dots.ttsconstraints/recommended.txtnow pinsgradio==6.17.0— phantom, not on PyPI → freshpip install dots.ttsunsatisfiable; FIX = pindots.tts==0.2.1+ DROP the-c recommended.txtconstraints (0.2.1 pulls working gradio 6.17.3). (2) pinning onlytorch==2.8.0let torchaudio float to 2.11.0 → dots.tts refuses to load (minor-version match check); FIX = pintorchaudio==2.8.0. ⚠ DEPLOY LESSON:docker compose up -dto a new tag swaps the LIVE container BEFORE any health check — a broken image crash-loops production (ratatoskr TTS down ~1-2min this session). NEW PATTERN = build → test in a THROWAWAY container on an alt port (:8199) → health+verify → only THEN cut live over (redeploy2.sh). v2 LIVE + healthy on irv-ml1:8198, CONSUMER-CONFIRMED clean (ratatoskr verified end-to-end on their :8765 — apostrophe string reads clean, /api/tts 200 @ 48kHz, no client change; the ~1-2min blip didn't hit them, their concurrent auto-audio issue was client-side localStorage). rollback =sed DOTS_TAG=v1 + docker compose up -d dots-tts(v1 image retained). Also: deployed container GPU crept ~6→13.9GB over 8h serving (cache accumulation; a redeploy resets it — watch item). reference_chatterbox_fast_repo -
[2026-08-09→10]dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (voices/). Operator-directed eval to potentially replace chatterbox-fast. dots.tts VERIFIED real (canonical HF nsdots-studio/,rednote-hilab/dots.tts-*redirects there; Apache-2.0; PyPIdots.tts0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). Runs on Ampere 3090 (sm_86, bf16, no fp8 dep); optimized RTF 0.22 at num_steps=10 (from_pretrained(..., optimize=True)CUDA graphs — raw unoptimized was 1.21), ~6GB VRAM, 48kHz, streams (generate_stream). Venv+cache atirv-ml1:/home/lkraven/dots-tts(~10GB). Operator design calls: SGLang Omni serving (OpenAI/v1/audio/speech), transcribe-refs-first,soarvariant. ⚠ Omni serves soar but its continuous-batching + streaming opts are mf-only (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript: mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked intovoices/derive.py): trim ref to a clean ~6–10s clip ending on a sentence boundary + accurate transcript of exactly that clip. CANONICAL VOICE CORPUS stood up in eshpfivoices/(operator idea): engine-agnosticcanonical/<v>.wav+transcripts/<v>.txt→ per-engine ref sets DERIVED byderive.pyreadingengines.yamlprofiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated),derived/gitignored. 4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders A6000=device0 (ComfyUI-full) — pin the 3090 withCUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0; andPYTORCH_CUDA_ALLOC_CONF=expandable_segmentsCONFLICTS withoptimize=TrueCUDA graphs (curr_block error). Booths:dots-vs-chatterbox,dots-voices-optimized. SHIPPED 2026-08-10: operator A/B verdict "dots is very good" → containerized as a thin FastAPI wrapper over DotsTtsRuntime (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). LIVE on irv-ml1:8198 (local/dots-tts:v1, OpenAI/v1/audio/speech+/health+/v1/voices, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack =stacks/dots-tts/(Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA:optimize=True(torch.compile/inductor/triton) needs a C compiler at RUNTIME — slim image mustapt install build-essentialor model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persistTORCHINDUCTOR_CACHE_DIRto a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfivoices/(operator ruled keep-here). REMAINING: ratatoskr client cutover to :8198/v1/audio/speech(Phase-2 tail, peer-coupled — draft the ask). reference_chatterbox_fast_repo reference_zonos_tts_stack reference_verify_hf_repo_ids_before_pull -
[2026-08-08]worldtree-dev #400 CLOSED → fiction-decomp snapshot cleared from nh3-dev. worldtree-dev signaled #400 done (shipped v1.0.0b185; exact-lexical efficacy 79%→12% on ratatoskr's gate, brokkr no-harm bracket green both ends; the snapshot served 4 probe rounds — rank decomposition, promoted-vs-gold annotation, tie-set falsification, A0/A1/A2 mechanism probe). Cleared~/snapshots/worldtree-400-fiction-decomp(208M: chroma + manifest/provenance/stamp) — a read-only rsync copy of PERSONAL Worldtree's Chroma (source on corviduo-dev, so safe to remove). LEFT INTACT:rex393-fiction-index/rex393-fiction-snapshot(separate operator KEEP word, unchanged) +r42-gate-*. No config deltas rode this train. Only remaining non-blocking await = ratatoskr-dev's chatterbox-fast knob revert. Replied confirming (01KZJ9GMCC…). -
[2026-08-07]chatterbox-fast "broken audio" root-caused (T3 AR tail over-run) + FIXED (max_chunk_chars=250 cap, :v2 deployed). Long saga, operator-driven clean diagnosis. Symptom: ratatoskr's migrated RP-surface TTS "swaps to German" / "dead air" / "garbage" on long turns. NOT German-leak (Turbogenerate()has NO language param — plain AutoTokenizer, nolanguage_id; the multilinguallanguage_id="en"lever lives only in the separateChatterboxMultilingualTTS), NOT OOM alone. Real cause: the Chatterbox Turbo T3 model OVER-RUNS its generation tail — a long singlegenerate()degrades into garble/dead-air in its final ~2-3s (lib filters OOV tokens<6561+ pads silence = messy AR tail). The scheduler's buffer-ratchet builds 300-600 char mega-chunks that land in that zone; streaming concatenates each bad tail (worst case). ratatoskr's anti-"German" knobs (top_k=80/temp=0.5) made it WORSE — tight sampling pulls the degradation onset SHORTER (~200 chars vs ~300 at default knobs). Diagnosis method (deterministic, no ears-only): single-shot length sweep + amplitude-gated voiced-ZCR (garble spikes ZCR; must gate on |x|>500 else trailing silence confounds it) — degraded voiced-tail = 1.58× mid, clean = ~0.64-1.1×. FIX: server-sidemax_chunk_chars=250cap on the scheduler (:v2image,CBF_MAX_CHUNK_CHARS=250env) — bounds each generation to just under the ~300-char onset → clean 3-4 sentence chunks (max prosodic arc while clean). Operator ear-confirmed clean audio + clean joins; chatterbox's low emotiveness keeps chunk joins smooth (the harsh joins that got Zonos rejected are absent — operator's key call). ratatoskr TODO (relayed msg01KZER9X7S): revert knobs to default (top_k→1000, temp→0.8), send full text (server chunks internally), keep the 503-on-empty guard. Cap value tunable per-request (max_chunk_chars) + env. Deeper prosody (if ever wanted) = scheduler Phase-2 context-priming at joins (feed prior sentence as discarded-audio context; +latency). ⚠ FOOT-GUNS: (1) acoustic tail-trim is UNRELIABLE — sibilants ('s'/'sh'/'f') spike ZCR like garble, can't cleanly detect the speech→garble boundary. (2) build-context vs image drift — the:v2image was built from cap source, but after a:v1rollback the build context held:v1source → adocker compose buildwould've silently produced a cap-less:v2; re-synced the flat cap source to/opt/docker/compose/chatterbox-fast/(rebuild-verified). ⚠ DIVERGENCE (follow-up): deployed build context is FLAT (app.py/scheduler.py,from scheduler import, thin-overlayFROM local/chatterbox:v1, cap-only) vs thevh/chatterbox-fastREPO which is PACKAGE-layout (chatterbox_fast/,from chatterbox_fast.scheduler, self-contained Dockerfile) + hasnorm_loudness(repo commit6bc7bf0= cap; deployed omits norm_loudness deliberately to keep the ear-test unconfounded). Reconcile the two layouts so a repo-based rebuild matches deploy. Rollback:.bak-cap-20260807-104850backups on irv-ml1 +:v1image both retained. reference_chatterbox_fast_repo reference_zonos_tts_stack -
[2026-08-07]Zonos2 TAKEN DOWN on the 3090 (irv-ml1) — operator-directed "for memory", TEMPORARY. Freed ~17.4 GB (3090: 728 MiB → 18.2 GB free) so chatterbox-fast (co-resident, was OOMing on long generations) has headroom. ⚠ Restore is manual — Zonos2 :1920 was a DETACHED native process (NOT systemd/docker), reparented to init. GPU memory was held by the--multiprocessing-forkCHILDREN (1966165=16.4G, 1966166=1G), which ORPHAN to init when you kill the parent — had to SIGTERM the children explicitly (killing the parent 1965942 + uv-run 1965935 alone left the 16.4G held). RESTORE CMD (from irv-ml1, user lkraven):cd /home/lkraven/tts-audition/models/zonos2 && nohup uv run python -m zonos2 --model-path Zyphra/ZONOS2 --host 0.0.0.0 --port 1920 --tts-default-voices-dir ./default_voices/ --cuda-graph-max-bs 1 --num-pages 16384 --max-running-requests 2 --memory-ratio 0.3 > /tmp/zonos2.log 2>&1 &thendocker start zonos-gateway. Consumers that lost Zonos: asset-engine + gateway-chat (via LiteLLMext-ttsalias → zonos-gateway :8890, now stopped); ratatoskr already migrated OFF to chatterbox-fast (unaffected). Also unblocks proper drift/cap testing (OOM was blocking it). reference_zonos_tts_stack -
[2026-08-07]chatterbox-fast: donut voice added + full contract delivered to ratatoskr-dev (their TTS migration off Zonos). Operator-directed. Copiedzonos-gateway/voices/Donut.wav→ chatterbox/refs(/worktank/chatterbox/reference_audio/donut.wav— the reference_audio SUBDIR is lkraven-owned so no sudo despite/worktankroot; container globs/refslive → NO restart), exposed asvoice:"donut"(lowercase); verified clean 7.5s synth (24kHz, RTF ~0.31). A/B booth (chatterbox vs zonos donut, same line) athttp://10.100.10.50:8090/b/donut-chatterbox/. Answered ratatoskr's 8-question contract ask from the live gateway (local/chatterbox-fast:v1) + source: NOT OpenAI-shaped (POST /tts; bodytext/voice/format/stream, notinput/model/response_format); NO affect dials (Turbo ignores cfg_weight/min_p/exaggeration — the architecture-changing answer they flagged; Zonos stays the only fleet TTS with real emotion steering); streaming WAV placeholder-header shape IDENTICAL to Zonos (their per-chunk Web Audio path survives); SR 24000 (Zonos 44100); server chunks arbitrary-length text internally (no client-side chunking, unlike Zonos's 71.2s cap); English-only, no language pin. FYI-worthy (operator): ratatoskr is moving its RP-surface TTS OFF Zonos back to chatterbox-fast → loses the live-PAD affect coupling (heavy Zonos emotion investment) — their call, trade-off flagged to them. auto-memoryreference_chatterbox_fast_repoenriched w/ the live contract. reference_zonos_tts_stack -
[2026-08-07]Fleet reranker cut over: Qwen3-Reranker-0.6B → BAAI/bge-reranker-v2-m3 (Brokkr R43). The incumbent was measured HARMING 80/90 fleet queries (no-reranker beat it 89/90 vs 56/90). R43 bake-off: the A2 control (same Qwen weights, seq-cls head) scored identical to the incumbent → proved the fault is a training-prior not the serving head → cancelled the expensive Qwen3-4B arm; A3 (bge-v2-m3) won on multilingual safety + bare-name recovery. LiteLLMrerankerrepointed incumbent→A3 :8013 (boundary 2026-08-06T17:37:48Z, config-edit + ~52s gateway restart); R42 v13 gate PASSED first-ever (56/90→90/90). Incumbent kept warm :8002 (rollback viaqwen3-rerankeralias), A4 fallback :8014. Full arc + rollback runbookdocs/pfi/reranker-selection-ledger.md; commits ad2df89/2c11748/377f8a4 (unpushed). auto-memories: the earlier reranker-serving notes. -
[2026-08-07]Personal-Worldtree kb-contamination incident (WT #394) diagnosed; attribution CLOSED UNRESOLVED. A reconcileWingStore._embedfull-tree walk (kbfs_root=KB_PATHroot, sibling wings nested) swept 5,354 fiction+main rows into personal'sknowledge_base(2 superseded generations served as current). Fixed by WT #394 (aca39a1, kb walks exclude sibling wings; ships b182). Trigger un-attributable — peer reconcile via the SHARED infra-ops identity + 0 dockerd exec-logging = fingerprint-less. Durable finding → auto-memoryinfra_ops_shared_identity_attribution_gap, PARKED (operator ruled A) into project_migrate_infra_access_to_claude_credentials. Evidence hold on the 5,354 rows until operator sequences cleanup (w/ Brokkr, on #394's agenda). -
[2026-08-05]Fleet CI resilience flip (DEFAULT_ACTIONS_URL=self) — attempted end-to-end, PARKED on a runner action-fetch auth blocker; infra-ops to research it (operator-directed, deferred, NOT now). 7 gitea action mirrors staged public+populated (orgsactions+astral-sh); the flip resolvesuses:correctly but act_runner v0.6.0 can't authenticate its fetch to gitea 1.26 ("Invalid username or token. Password authentication is not supported"). Reverted (CI back on github default);REQUIRE_SIGNIN_VIEW=falseKEPT as a standing change (operator, internal WG net). Full endeavor, the reliable nh3-dev-egress + git-SSH mirror method, exact config state, smoke method, and next step →persistent-memory.d/2026-08-05-ci-flip-parked.md -
[2026-08-05]Booth — 3 features shipped, live on:8090+ tagged. (1) verbatim-index.htmlbooths get a floating top-right "‹ all booths" chip + inherited favicon, doctype/charset-safe byte-injection (booth-v0.1.5,8577e7e); (2).mdrenders +.txt/.logview in-booth without downloading via the/b/<n>/viewroute + amarkdowndep +doc.html(booth-v0.1.6,315faac); (3) prev/next arrows in the image zoom viewer — wrap-around + keyboard ←/→, hidden for single-image booths (booth-v0.1.7,c37a425). Canonicalservices/booth/; deploy =systemctl --user restart booth.serviceon nh3-dev (runs from the checkout's.venv;uv pip installnew deps into it first); 47 tests.uv.lockgitignored (348c5c1). -
[2026-08-05]worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (d5d33df, deployed on nh3-dev).herald.py:363rendered the wake command from the empty fresh mail set on the re-nudge path (should bedeliver_msgs) →messages[0]IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix +render_commandempty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. nh3-extdev herald 2.1.2 upgrade DEFERRED (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev/tmp(sha256003508…cef27) —uv tool install --force+ restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memoryreference_nh3_dev_althing_herald. -
[2026-08-03]worldtree b168/#384/#385 arc COMPLETE — providers.yaml boot-gate pre-sync → b168 deploy → DCC+P&P re-ingest (705+667 concepts, 0 truncations, #385 budget fix validated vs April's 785 control) → #381 restart → operator-approved production dedup sweep (785 April orphans deleted frommain, 4009→3224). Fiction wing 166→1,372 concepts; consumer verify 0/5→5/5→saturated. Three of MY foot-guns hardened into fleet runbook rules (mv -t,docker exec -u 1000, shared-containerd pull-race — see Tried-and-abandoned). Full runbooks (deploy-wt-config, Chroma-verify, config-delta pre-sync rule, #381, sweep) →persistent-memory.d/2026-08-03-worldtree-b168-384-385-arc.md -
[2026-08-02]mimir-inbox / #377-read-path arc — deployed + 4 bugs found/fixed/verified + a cloned voice. mimir-inbox live on corviduo-dev:8091 (#377 write+read proven, live8ece117); worldtree-dev #380 (wing-blind index) + #381 (stale-client restart) + #382 (intermittent Mimir grounding) chased and verified 3/3 by ratatoskr-dev; muninn-gate → dispatch 0.1.5; donut voice cloned from the 65-frost Booth bundle into the Zonos gateway; Zonos streaming confirmed already-working. Full arc, procedures, and lessons →persistent-memory.d/2026-08-02-mimir-inbox-arc.md -
[2026-07-31]muninn-gate (#377 ingestion front door) BUILT + DEPLOYED + healthy on corviduo-dev:8090. First-boot acceptance passed (watcher:running:true proves ingestion_root byte-identity); submit path deferred to the mimir-inbox era. Full wiring (uid-1000, state-volume mount, staging path-agreement, BuildKit-secret build, deferred repoint + operational guards) →persistent-memory.d/2026-07-31-muninn-gate-deploy.md -
[2026-07-31]worldtree-sdk 1.1.0 (Python) published to vh Gitea PyPI + a durable infra-ops publish cred. memory_context pass-through; unblocked wyrd-dev. claude-bot now a write-collaborator onvh/worldtree-sdk(source pulled via the Gitea API archive — git-HTTP 403s on that repo); publishing to the vh USER namespace can't be delegated (401reqPackageAccesseven withwrite:package) so it needs an owner token — operator saved a FULL vh site-admin token at~/.config/gitea/vh-token(0600) for it (⚠️ high blast radius, kept over a scoped one; org-namespace migration is the only real de-personalization, parked by wtsdk-dev). auto-memoryreference_infra_ops_vh_gitea_token_and_sdk_publish. -
[2026-07-31]kimi-k3 "output cap" root-caused = a ~16384 REASONING-token ceiling, not an output cap. heid's cross-frontier panel was silently degraded (empty content,finish_reason: stop). dvalin+bil researched (docs said deprecated-max_tokens); heid's live data refuted that (completion hit 18455) → it's a reasoning ceiling. Proven on the wire against heid's real 500KB bundle:reasoning_effort: lowdrops reasoning under the ceiling → content returns, on BOTH coding + general endpoints. Fix is CALLER-side (no gateway change): sendreasoning_effortviaextra_body(LiteLLMdrop_params: truestrips the top-level param — why heid's earlier attempt no-op'd). Relayed to heid to validate; backstop =allowed_openai_paramson the route. →persistent-memory.d/2026-07-31-kimi-k3-reasoning-cap.md -
[2026-07-27]jackdaw-compose.service DECOMMISSIONED (jackdaw-dev request; the JackDAW AI Composer was cut from v1 by operator decision 2026-07-27). Stopped + disabled the nh3-dev:8787user service (no client calls it — ai/server/AiChat deleted from main,/composeproxy removed); unit archived not deleted →~/.config/systemd/user/jackdaw-compose.service.decommissioned-20260727(revival = rename +daemon-reload). No credential revoked — the unit used the SHARED all-agents LiteLLM key (sk-eA_XOd…, modelgen), not a dedicated one. Code preserved on jackdaworigin/ai-composer-preserved; treat as permanent. The:4500HTTPS audition bench is untouched. (Supersedes the 2026-07-23 stand-up line below.) -
[2026-07-25]bil-smithy-dev wired as an althing zellij-window-ping (pane route). She's adriver: humandwarf peer (panebil-smithyalready live alongside eitri/dvalin/regin-smithy in theClaudezellij session) but had no delivery route → smoke messages posted to the bus but never reached her window. Mechanism (reusable for any pane-route handle):~/.althing/config.yaml→zellij_sessions.Claude.agents[]mapshandle→target(a zellij pane TITLE, matched vialist-panes -jinalthing/zellij.py:resolve_pane_id) →command(heraldwrite-chars+ CR into that pane). The herald loads config ONCE at startup (herald.py main()), sosystemctl --user restart althing-herald.serviceafter editing. Added bil (target: bil-smithy), restarted, verified: herald delivered the pending smoke01KYD7W7CF…(available→attempted→delivered). ⚠️ Noticed pre-existing pane-route errors onworldtree-codex+eitri-smithy-dev("route-error: list index out of range", empty msg_ids — likelyrender_command messages[0]on an empty list; NOT caused by this change, bil works) — worth a herald look.
201 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-08-15]Grafted bf16 MTP loads UNINITIALIZED (0% accept) unlessre:^mtp.*is in the quant-configignore; and W4A16=Marlin (not native FP4) costs ~20% even on decode. Cost a premature 79 GB delete of a good model (declared desync-dead off the 0%). Lessons: test MTP on bf16 FIRST, isolate before deleting; modelopt 0.43 is dependency-hell for qwen3_5 (list-vs-dict quant_cfg + transformers conflict) — use llm-compressor. Full →persistent-memory.d/2026-08-15-uncensored-gen-seat.md -
[2026-08-03]ComfyUI--enable-triton-backendon the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3. adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added toCOMFY_CMDLINE_EXTRA, recreated) →triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5")incomfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8, failing at node 5 CLIPTextEncode. Triton's fp8 dequant kernel targetsfp8e4nv(Hopper/Ada e4m3); sm_86 Ampere (A6000) lacks hardware e4m3 → the JIT compile dies. With triton on it grabs the global--fp8_e4m3fn-text-encdequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchangedsha256:94afb8ca, sage intact, prod restored). The parked cu130 rebuild won't fix it (e4m3 = hardware format, not CUDA version). DEFERRED to the Ada refresh (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). Mechanics:--enable-triton-backendis a composeenvironment:var, so toggling it needsdocker compose up -d(recreate), NOTdocker restart(reuses the baked env, no-ops silently). Full: auto-memoryparked_triton_backend_ampere_fp8. -
[2026-08-03]corviduo-dev shared containerd: a concurrent-pull race fails ONE instance's deploy; DON'T "prune to fix" — the image is in-use by the instance that won the race. b169 personal deploy failed atdocker compose pull(Lchown … no such file or directoryon the big torch layer → looked like disk pressure / corrupt snapshot). ACTUAL: NOT disk (56G free, inodes 7%). demo + personal + pinned share ONE/var/lib/containerdon corviduo-dev; demo (from main) and personal (from staging tag) extracted b169's shared torch layer simultaneously → personal's hit a partial snapshot mid-race and aborted while demo's completed. The image6e34a87was FULLY VALID — demo was RUNNING it healthy. Fix = just re-run the failed deploy (image already materialized; compose pull finds it present). NEAR-MISS: worldtree-dev's suggested "prune unused images/snapshots" would have rmi'd6e34a87= the image the running demo depends on → demo outage. Lesson: before any prune/rmi "cleanup,"docker psthe running images — an "unused" image may be a co-tenant's live one; and verify the failure's REAL cause (disk? inode? in-use? race?) before applying the suggested remedy. (Pipeline fix, deferred: serialize demo-from-main + personal-from-staging, or a per-image pull lock, to avoid the shared-layer extraction race.) -
[2026-08-02]docker execinto worldtree containers defaults to ROOT — root writes contaminate the uid-1000 (vh) KB tree. Mysudo docker exec … --reindexon personal ran as ROOT (muninn app = uid 1000); its wing git-commit + atomic note-swap left root-owned files in theworldtree-personal_worldtree-kbvolume: a root-owned.old-<job>backup dir (blocked the uid-1000 retry'srmtree→ Errno 13, because unlink needs write on the DIR and it was root:root 755) AND 60 root-owned loose git objects in.git/objects/. Fix (host-side, corviduo-dev):sudo rm -rfthe superseded.old-dir (tar'd aside to /tmp first) +sudo find … -user 0 -exec chown 1000:1000the objects (ownership-only, git-content-safe; the.git/objects/XX/dirs were vh-owned so these weren't a hard blocker, but violated "clean tree"). RUNBOOK RULE (worldtree-dev, ADOPTED): anydocker execinto worldtree containers that WRITES pipeline state runs-u 1000, never default-root — same genus as the mv footgun (acting without matching the target's constraints; 3rd such slip in one session). GOTCHA that hid the scope:find … -user 0 | head -20TRUNCATED (the.old-dir alone had 153 files, so the first page was all.old-) → I "verified clean" off a partial list. Neverheada scope-defining find; count first (| wc -l). Related blind-spot (muninn-dev): a root-owned job SUBDIR passes every requeue guard (job_row/dispatch/list_jobs render fine) AND/health(contract'sos.access(ingestion_root, W_OK)tests only the ROOT dir, so a foreign-owned subdir underpending/still reportsingestion_root_writable: true) — then the uid-1000 gate can't write into it. "Clean board + green /health + failure at next mutation." muninn-dev added an OWNERSHIP column to the standing post-move check to catch it; two green signals both miss a foreign-owned subdir otherwise. -
[2026-08-02]mv <job> complete/ → failed/RENAMED the job tofailedbecause failed/ didn't exist. worldtree-dev's round-2 unblock command (mv /data/state/ingestion/complete/<job> /data/state/ingestion/failed/) assumedfailed/existed; on PERSONAL muninn it did NOT (fresh instance — root wasactive/ complete/ pending/ sources/, nofailed/).mv src nonexistent/renames src→nonexistent, so job1 became thefaileddir and job2 nested inside it. Caught on post-movels(failed/ held job contents, not two subdirs), reconstructed via complete/ as watcher-safe scratch + rebuiltfailed/(worldtree:worldtree 755) — NO data loss. Lessons: (1) beforemv X into-dir/, verify the dir EXISTS ([ -d dir ]) — an emptyls dir/ 2>/dev/nullis AMBIGUOUS (missing vs empty), which was the preflight miss that let it through; (2) the correct guard ismv -t <targetdir> <src>(--target-directory): it refuses a MISSING target loudly (rc=1, "No such file or directory", nothing moved) — this is the house convention for queue/state moves now. TESTED by muninn-dev on coreutils 9.1: a trailing slash does NOT protect —mv src failed/withfailed/missing STILL silently renames tofailed(rc=0); "just add the slash" is a false guard. (mkdir -p failed/first also works, butmv -tinverts the failure from silent-wrong to loud-safe in one flag.) Containershis dash — no(in echo strings. SILENT failure mode (muninn-dev carry-forward): a misplaced ingestion-state move doesn't crash anything —list_jobs()stays OK, loose files are inert; the ONLY symptom is the job quietly absent from the board (job_row→None, requeue→not_found/404, looks IDENTICAL to the original block). So after ANY state move, verify the job is actually ON THE BOARD (job_rowfound + guards pass), don't trust mv exit codes — and confirmjob.dispatch.jsonsurvived (requeue refuses a dispatch-less job with the same not_requeueable symptom). Cross-checked + all-clear'd by muninn-dev, who correctly refused to mutate ingestion_root (INV-MG-1) and flagged instead. DON'T TIDY (round-2 pending): both DCC + P&P jobs currently REST in personalfailed/with manifests readingstate: completeuntil round-2 requeue runs — deliberate + load-bearing (requeuekeys on DIRECTORY PLACEMENT, not manifest state); looks wrong to anyone cold, leave it exactly as-is. Round-2 sequencing: the requeue is mimir-dev's browser flow (pending their operator's board-vs-API ruling); muninn-dev is the gate confirmer (runs the post-move board-check inside its custody — the right split, don't reach across INV-MG-1); infra-ops = the #381 restart after both jobs go terminal, then later the supervised main-collection sweep. Guard-verified HOLD LIFTED by muninn-dev 02:36Z. ARC COMPLETE (2026-08-03 ~05:49): both books terminal — DCCmimir-6351554e8e8f705 concepts + P&Pmimir-f3887c9b97b7667, extracted AND indexed, 5/5 phases, 0 failures/truncations (validates the #385 budget fix vs April's 785 control); #381 restart-after-ingest FIRED (personal api, healthz/readyz 200 ~25s), retrieval-visibility confirmed (search_library returns DCC+P&P from fiction post-restart); handed ratatoskr-verify go to worldtree-dev. Delete-sweep precondition NOW MET — the stale DCC rows inmainare genuine duplicates of livefictionrows, so worldtree-dev's supervised sweep of the ~785 April orphans is unblocked (still comes to me supervised: snapshot + operator-in-loop). -
[2026-08-02]donut voice multi-clip reference (onyx-58 expansion) — TRIED, REVERTED. Folded theonyx-58bundle's 3 Donut clips (seg101/seg110/seg148) in alongside the original seg000 → a 52.0s 4-take concat reference, hoping a longer ref → more robust speaker embedding. A pinned-seed A/B (5 pairs, varied registers, boothdonut-onyx58) showed the original single-clip seg000 (16.3s) sounds better — concatenating disparate takes muddied the timbre more than the extra range helped. Reverted to seg000-alone (live + build-source). Two durable lessons: (1) for a faithful clone, a single clean representative take can beat a longer multi-take concat — more reference audio is NOT automatically better when the takes vary. (2) Emotion steering pulls the output AWAY from the cloned voice fast (operator's craft rule) — keep donut (and clones) emotion-neutral for fidelity; the gateway only enables emotion when anemotion_*/presetdial is explicitly sent, so bare{input,voice}calls stay pure-clone.seg148was diarized SPEAKER_03 but IS Donut (operator-confirmed misdiarize). onyx-58 curated bundle lives in boothonyx-58(24h TTL — stash to/mnt/smithy/voice_clones/if a future middle-ref experiment is wanted). -
[2026-08-02]Verifying the INDEX is not verifying GROUNDING (#382). Asearch_libraryreturning wing=fiction hits proves the content is retrievable; it does NOT prove the agent (Mimir) trusts and uses those hits vs. silently answering from training. I reported "Mimir read Austen back to you" off a grounded-looking answer; ratatoskr-dev caught that grounding was intermittent (some sessions discarded the correct hits and substituted training knowledge). Test the harder claim — are the citations note-extracted or model-knowledge? — and reading the DEPLOYED artifact beats trusting the test for "is the fix live." -
[2026-07-30]brokkr's WebSearch "verification" CONFIRMED a hallucination — 3 phantommicrosoft/Mage-Flow-{Base,Turbo,Edit}repo IDs. brokkr-smithy-dev handed 3 gated-looking repo IDs for an operator-directed model pull; they don't exist (its own web-search fabricated an arXiv ID + project page, twice). Lesson: the HF registry API is ground truth — an unauth 401 ≠ exists ({"error":"Invalid username or password"}masks private/gated/nonexistent alike), an authed 404 = phantom, andauthor=X&search=Yrefutes existence. API-verify every repo ID before a pull; LLM-summarized web fetches confabulate. auto-memoryreference_verify_hf_repo_ids_before_pull. -
[2026-07-30]magpie TTS serving — evaluated, ABANDONED. Pulledmagpie_tts_multilingual_357m(the one real repo of brokkr's batch) to NFS, stood it up on irv-ml1 (ephemeral NeMo-Speech-maincontainer — stock PyPI/NGC NeMo can't load v2607), A/B'd vs Zonos → Zonos wins expressive English decisively, multilingual not needed. Not served;magpie-nemotorn down..nemoKEPT on NFS as brokkr's fine-tuning base. auto-memoryproject_magpie_tts_eval_rejected. -
[2026-07-25]Chaining the althing wake-listener arm orphans it.reply && althing-wake-listener &(or spawningalthing-wake-listenerwith&inside arun_in_backgroundtask) → the&-child reparents to init, UNTRACKED by the harness: no fire-notification, and re-arms bounce rc3 off a lock nothing services (mail silently unwatched). Compounding foot-gun: re-arming after a plain operator turn (not an actual fire) collides with the still-live prior listener (rc3). FIX: spawnalthing-wake-listeneras its OWNrun_in_backgroundtask, and re-arm ONLY after a real fire (<task-notification> completed rc0). Reclaim an orphan withalthing-cli stop-monitorthen re-arm.
135 older entries archived to archival-memory.md.