Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 2dd459d2e5 memory: Worldtree memory-split flip protocol, and all three deployments measured writable
worldtree-dev's U6 reader refuses at boot if it cannot append and read back
<memory root>/reader/canary.jsonl. Per-euid subdirectories are lazy and only
warn, so that root canary is the single boot-blocking check -- which makes the
memory root's writability by the container uid the precondition worth knowing
before a flip rather than during one.

Protocol agreed with worldtree-dev: neither memory.reader.enabled nor
memory.writer.enabled gets flipped on any deployment without infra-ops
confirming that writability first. Both ship dark until the operator schedules
the tracer skeleton.

Measured tonight on corviduo-dev, all three pass. Also retracts a wrong
prediction I sent earlier in the thread: personal runs as root, not uid 1000,
and pinned is the only uid-1000 deployment -- it passes regardless because
/data/state is owned 1000:1000. Recorded with the caveat that a permissions
reading is a claim about its own date, so the probe gets re-run immediately
before any flip rather than cited from tonight.
2026-09-11 22:12:08 -07:00

104 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-09-11 ~17:45 PT (fv-ml1 relocation cutover PREPPED for tomorrow; Anaheim recovered except ana-ml2 which relocates; BabyYarros COMPLETE + evaluated; sentinel-r3 quant done, cyber-preview to re-run at FV)

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.

Repo purpose

  • 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up. /tank and other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. See persistent-memory.d/2026-09-10-beszel-fleet-wiring.md and stacks/beszel/README.md.

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecatedalthing-clipostbox, althing-wake-listeneralthing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild)
vh/soong-lab Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16)
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-09-11 ~17:45 PT.

fv-ml1 relocation — cutover PREPPED, executes TOMORROW

  • ana-ml2 → fv-ml1, moving to a NEW Fountain Valley colo (10.251.0.0/16) tomorrow; its power draw is the root cause of the repeated Anaheim breaker trips. Fully staged, nothing deployed: runbook docs/runbooks/fv-ml1-cutover.md, rename sweep scripts/fv-ml1-rename-sweep.sh (dry-run default, history-safe), exact DNS + LiteLLM commands inside the runbook. See Recent decisions [2026-09-11] fv-ml1 for the full plan.
  • Load-bearing at cutover: LiteLLM api_base 10.250.50.54→10.251.50.54 (darkens every inference alias if missed), DNS piggyback records, OPNsense as tailscale subnet-router. Box already down (clean cutover); /tank is local ZFS so data travels with the chassis.
  • Anaheim rack left DARK until the move (operator) — nothing to bring up, it relocates.

Anaheim colo — recovered except ana-ml2

  • Full-site power/breaker outage ~15:0x PT; recovered ~16:39 EXCEPT ana-ml2 (no power, relocating). The gitea-wide 403 (crowdsec crash → traefik bouncer fail-closed) was fixed by restarting crowdsec then traefik; LiteLLM + everything else healthy. ⚠ recurring post-power-loss step, now in the recovery runbook memory.

BabyYarros — COMPLETE + evaluated

  • Both arms trained (Base 2.5263 @ ckpt-125, overfits within epoch; Instruct 2.6114 @ 178) and evaluated: voice moved toward Yarros above the 0.046 measured noise floor (Base +0.157, Instruct +0.076), Instruct renders beats 9/10. Booth babyyarros-voice. Full frozen adjudication (romantasy control panel + 2nd seed + gen seat for beat-incumbent) DEFERRED — needs the gen seat back. See Recent decisions [2026-09-11].

Quants — sentinel-r3 done, cyber-preview to re-run

  • sentinel-r3 NVFP4 (grafted base MTP head) COMPLETE at /tank/aimodels/sentinel-r3-nvfp4-mixed (survives — ZFS). Acceptance/A-B deferred (needs a serving slot). cyber-preview NVFP4 died mid-quant with the ana-ml2 outage — re-run when fv-ml1 is up; both bf16 sources safe on /tank.

gx10 on althing; Jetson planning

  • postbox installed on gx10 (handle gx10, send-only — no reader on its inbox, it's a headless notifier/watcher-host; reply-expecting watchers post as infra-ops).
  • Jetson AGX Orin — discussed as an ESH House Computer (cameras via Frigate + local ASR/TTS); its native fit is vision/perception. Discussion only, not committed. Jetson Nano generation TBD.

Recent decisions

  • [2026-09-11] Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or memory.writer.enabled on any deployment without infra-ops first confirming the memory root is writable by the container's uid. The reader REFUSES AT BOOT if it cannot append+read back <memory root>/reader/canary.jsonl (deliberate, the #335 typo'd-reranker precedent: refuse loudly, never silently disable); per-euid subdirs are created lazily and only warn, so the root canary is the only boot-blocking check. The writer degrades rather than refuses. Both ship DARK (enabled: false, parity-only config/defaults.yaml) until the operator schedules the tracer skeleton. Measured 2026-09-11 on corviduo-dev — all three deployments PASS: demo :8080 uid 0 and personal :8081 uid 0 both have /data/state/memory at 1000:1000 755 writable; pinned :8082 uid 1000 lacks memory/ but its parent /data/state is 1000:1000 755 so it can create it. ⚠ I had predicted personal was uid 1000 and warned it would fail — wrong, retracted; only pinned runs as 1000, and it passes anyway. ⚠ Re-probe immediately before any flip: a permissions reading is a claim about its own date, not about boot time. Heimdall side is clear too — demo and personal grant 7x tool.*, pinned uses image defaults, and the lone tool.evidence.* is additive, so tool.memory_read needs no policy change. Thread 01M2A05WED5W.

  • [2026-09-11] Plex hardware transcoding on the Arc A580 FIXED (esh-pve-nas LXC 105) — every setting was already correct and the fault was one layer below them. intel-media-va-driver 22.3.1 (Apr 2023, stock jammy) predates Arc/DG2 support and exports only __vaDriverInit_1_14, against the libva 2.22 Plex BUNDLES and loads via RPATH. Passthrough, cgroups, plex in video+render, HuC authenticated, Plex Pass, HardwareAcceleratedCodecs=1 and the Arc already selected as HardwareDevicePath — all good the whole time. Fixed with Intel's client-GPU repo (rolling jammy client) → iHD 24.3.4 (__vaDriverInit_1_22) + a consistent libva 2.22.0.2-87 set, now pinned + apt-mark hold (verified: a simulated upgrade moves 152 packages, touches none of the six). Also repaired a half-finished prior attempt — libva/libva-drm hand-installed at 2.22 with libva-x11 left at 2.14, killing every X11 VA-API app on va_fool_postp. ⚠⚠ pct snapshot REFUSES on a bind-mounted guest AND STILL EXITS 0 (LXC 105 has mp0: /tank/media) — use zfs snapshot nvme/subvol-105-disk-0@<tag> and read it back. ⚠⚠ A synthetic Plex Transcoder run is NOT a valid test (Plex bundles its own libc among 61 libs; my harness failed identically before and after a fix that worked — no positive control, so its negatives were worthless). Only a forced transcode settles it: PASS names the device (testing API vaapi for device '/dev/dri/renderD129' (Intel DG2 [Arc A580])). ⚠ The original empty final decoder: , final encoder: was an absence of evidence, not failure — TranscodeSession was 0. Jellyfin LXC 107 left alone (operator: not actively used). → persistent-memory.d/2026-09-11-plex-arc-vaapi.md, runbook docs/runbooks/plex-arc-vaapi-jammy.md

  • [2026-09-11] Beszel priority 2 complete: both DB hosts and both PBS hosts verified, 16 new alerts; fleet 17/18 up (ana-ml2 down). → persistent-memory.d/2026-09-11-beszel-priority2.md

  • [2026-09-11] Beszel priority 1 complete: all six installed and verified. After Anaheim recovery, live Synology samples and alerts verified; fleet 13/14 up, only known ana-ml2 outage remains. Configs not committed. → persistent-memory.d/2026-09-11-beszel-priority1.md

  • [2026-09-11] Sentinel-R3 pulled, MTP-grafted, and quantized as a M.O.G.-SEC seat candidate — quant DONE, acceptance UNVERIFIED (blocked on GPU space). Operator got access to glyphsoftware/sentinel-r3 and asked to compare vs the running M.O.G.-SEC seat + pull if promising, then "quant it with a grafted mtp head". It is promising and a better FIT: same base (stock Qwen3.8-27B), same qwen3_5 hybrid arch, same 262K, vision-intact — but M.O.G.-SEC is a persona on stock weights while Sentinel-R3 is a REAL SFT finetune on 1,230 authorized-pentest agent trajectories over a 19-tool surface that matches our own harness (Bash/Read/Write/Edit/Grep/Glob/Agent/Task*/Monitor/…). Card is unusually honest (flags its own mmlu-cybersec 0.88 as within-noise of base). HF check: M.O.G.-SEC repo unchanged (sha still our pinned deede6779…). MTP: Sentinel ships ZERO mtp tensors; grafted the verbatim base head from qwen38-27b-uncensored-bf16 (compare_mtp_head → IDENTICAL) — lineage correct since Sentinel's base is stock Qwen3.8-27B and that head is a verbatim base graft. ⚠ Acceptance is UNVERIFIED and may differ from the 47.7% the head hits on STOCK weights — it now reads hidden states from an SFT-finetuned body (the exact Stage-1b residual risk). Quant = the standard mixed NVFP4-W4A4(MLP 0-55) + FP8-W8A8(attn/linear_attn/lm_head/MLP 56-63) recipe, ran CUDA_VISIBLE_DEVICES=1 on GPU1 free space, no seat downtime, 51→22 GB. post_quant carried the head forward + re-injected re:^mtp.* (llm-compressor prunes it → the 0%-accept bug). Structural verify clean: 1968 tensors, 0 unresolved, 15 mtp, 333 visual, ignore has mtp+visual. Artifact /tank/aimodels/sentinel-r3-nvfp4-mixed (+ .PROVENANCE.txt).License is PROPRIETARY (Glyph Proprietary v1.0, all-rights-reserved) — operator's fair-use/licensee call, not apache like M.O.G.-SEC. ⚠ Serving/A-B is BLOCKED on GPU space: weights are 22 GB, GPU0 has 7.6 free / GPU1 19.9 — a probe serve needs a freed co-tenant slot (~25 GB), which is a material-consequence call. Serve with the PROSE system prompt (trained on prose tools, not structured tools=). → /tank/aimodels/sentinel-r3-nvfp4-mixed.PROVENANCE.txt

  • [2026-09-11] MEASURED: two concurrent training jobs on pfi-gx10 are 13% NET SLOWER than running them back to back — VRAM is not the constraint and never was. Operator asked to run the two BabyYarros arms in parallel if VRAM allowed. It does, comfortably: 18.4 GiB per 4B LoRA job, 36 of 121 GiB with both up, 98 GiB free. But the GB10 is a capacity box, not a throughput box, and the binding constraint is memory bandwidth. Solo baseline 37.10 s/it (n=6, 0.05% spread); with a second job both arms settled at ~85 s/it — 2.29x each, so combined throughput 0.0235 vs 0.0270 steps/s solo. Not a clean 2x split: the box is past its roofline and pays a contention penalty on top. Control: killing the second job returned the first to 37 s/it on the very next step, so the slowdown tracked contention and reversed with it. Chaining finished both arms ~43 min earlier than concurrency would have. General form: on this box, nvidia-smi free memory tells you nothing about whether a second job is affordable. Decision rule was pre-registered before the numbers were read (<55 s/it keep both, ≥2x chain). scripts/yarros-corpus/{launch-yarros-4b-base,chain-yarros-4b-base}.sh; the shared-GPU bypass is an explicit argument, never a default.

  • [2026-09-11] ana-ml2 → fv-ml1: relocating to a NEW Fountain Valley colo TOMORROW (operator decision). Its power draw (dual Blackwell PRO 6000, ~1.5 kW peak) is the ROOT CAUSE of the repeated Anaheim rack-breaker trips (2026-08-26, 2026-09-11) — moving it to its own circuit fixes the recurring whole-site outage. New site fv, same shape as Anaheim: server subnet 10.251.50.0/24 (fv-ml1 = 10.251.50.54, mirroring the old host octet), mgmt/BMC 10.251.250.0/24 (fv-ml1-bmc = 10.251.250.50). OPNsense firewall is the multi-homed gateway (.1 in every FV VLAN) AND the tailscale/headscale subnet-router advertising 10.251.0.0/16 — chosen over ana-ml2-as-endpoint specifically because the firewall stays up when the GPU box is down, giving out-of-band BMC access over the mesh — the exact thing the fleet LACKED during today's outage (no OOB path, BMC islanded). Rename to fv-ml1, full fv.internal DNS name. DNS approach: PIGGYBACKdns-sync builds name.site.zone with no check that the site is in the sites: block, so fv-ml1/fv-ml1-bmc records with site: fv resolve fleet-wide from the existing ana/esh/nh3 resolvers immediately; add a real fv resolver only when FV needs LOCAL resolution (OPNsense can't host the AdGuard the sync targets — it's FreeBSD/Unbound). Clean cutover: the box is already down (BMC dark, no power since the outage), and /tank is LOCAL ZFS with NO NFS from ana-nas, so data travels with the chassis. ⚠ Load-bearing repoint = stacks/litellm/conf/config.yaml (~10 api_base: 10.250.50.54:{8015,8016,8018,8019}10.251.50.54; darkens every inference alias if missed) — gateway STAYS on ana-docker so fv-ml1 serves cross-site (FV↔Anaheim metro, fine). Everything staged, nothing deployed: runbook docs/runbooks/fv-ml1-cutover.md (commit ce04f9d; exact DNS + LiteLLM commands) + scripts/fv-ml1-rename-sweep.sh (8400f3a; scoped, dry-run default, history/provenance-safe, manual-review list for judgement calls).

  • [2026-09-11] Anaheim rack LEFT DARK until the move (operator decision). ana-ml2 is the ONLY host still down post-recovery (BMC dark = no power); rather than power it on tonight just to shut it down for the truck tomorrow, it stays off. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) but there is nothing to bring up — the box relocates as fv-ml1.

  • [2026-09-11]RECOVERY FOOT-GUN, will recur every colo power event: crowdsec crashes on the hard power-off and traefik's bouncer fail-CLOSES — empty-body 403 on EVERY HTTP service behind traefik (gitea, homepage, …) while the apps themselves are fine. Signature (bifrost-dev reported it, gitea-shaped): HTTPS returns 403 content-length 0, no app body on all routes, but git-over-SSH works (SSH bypasses traefik). Diagnosis: gitea direct on localhost:3000 = 200 (app healthy), through traefik = 403; crowdsec container Exited (255); cscli decisions list EMPTY (not an IP-ban). The bouncer plugin does NOT self-recover from a startup-time LAPI-unreachable race — even after crowdsec is healthy again, traefik keeps 403ing until traefik itself is restarted. FIX: docker start crowdsec (its data/config are LOCAL volumes, comes up clean), wait for cscli lapi status = OK, THEN docker restart traefik so the plugin re-inits against the live LAPI. Verified 403→200 on gitea API/web/PyPI-index from an off-box vantage. This unblocked bifrost-dev's 1.2.0 PyPI publish (+ worldtree/wyrd/ratatoskr) and any HTTP gitea access; heid's SSH pushes were never affected. → add to the recovery runbook: crowdsec+traefik restart is a standard post-power-loss step.

  • [2026-09-11] Anaheim colo recovered ~16:39 PT EXCEPT ana-ml2 (bare metal, NO power — its BMC 10.250.250.50 is dark on standby, unlike same-subnet pfi-pve which is up → needs a physical PDU/PSU/breaker fix, not a boot). pfi-pve + all its VMs (ana-docker/ana-nas/ana-wg/corviduo-dev/pbs-ana) auto-started clean (on-boot gap held this time). LiteLLM came back up on its own (transient unhealthy during startup → serving). ⚠ Public WAN (38.120.12.44) ICMP still blocked from outside but HTTPS works fleet-internally (mesh-routed). ana-ml2 down blocks the gen/summarizer/mog-sec seats AND the cyber-preview quant re-run. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) to power-on + boot-watch the instant its BMC returns.

  • [2026-09-11] BabyYarros COMPLETE — both arms trained AND evaluated; the voice moved toward Yarros above the measured noise floor, and the instruct arm renders beats 9/10. Training: Base best held-out 2.5263 @ ckpt-125 (overfits within the epoch — best is the checkpoint, not the shipped step-178 adapter), Instruct 2.6114 @ 178 (still descending, undertrained if anything). Base-wins-held-out / Instruct-holds-instruction replicates Brontë at a near-identical 0.085-nat gap. Eval (gx10, seat-free, done during the Anaheim outage): three voice arms + instruct beat→paragraph. delta_cb (Burrows over char-bigrams vs held-out Yarros) ordering base-125 0.549 < instruct 0.631 < base-unadapted 0.706, same-author target 0.463; both adapters clear the 0.046 measured noise floor (within-arm seed spread, not the same-author distance — first cut mis-framed that) — base +0.157, instruct +0.076 vs control. ⚠ One seed-pair per arm, so the ordering CORROBORATES the independent held-out-loss ordering rather than settling it. Beats (instruct, chat template, Yarros SYS): on-beat 9/10 (it takes direction after raw-text training — the Skaldsong question, answered yes), in-band 5/10, ran-on 7/10 (length + clean-close discipline is the weak axis, same trade as Brontë). Booth: http://10.100.10.50:8090/b/babyyarros-voice/. Tooling scripts/r49-corpus/{voice_prompts_yarros.json,gen_beats_chat_yarros.py,voice_distance.py,build_booth_yarros.py}, commit 5558d9c. DEFERRED to power-return (needs the ana-ml2 gen seat): the frozen adjudication's romantasy control panel, a 2nd seed, and the beat-incumbent leg.

  • [2026-09-11] BabyYarros UNBLOCKED and TRAINING: the leak gate passes at 0 of 325 entities and 0 of 91 phrases, and closing it turned up three defects nobody was looking for. The gate itself is the first artifact — there was no committed instrument for "does any of the author's proper nouns survive", so Brontë's 0-of-203 was a hand count. scripts/r49-corpus/leak_gate.py now runs the same scan over the UNRENAMED source as a positive control plus a nonce negative control every time, because a detector that only ever sees renamed text cannot tell absent from blind. Its first reading was 212 surviving, not 86 — it scans the whole corpus rather than per work, and counts the sub-threshold entities rename never looked at. Training launched 10:06 PT on pfi-gx10: Qwen3-4B-Instruct, 1 epoch, seed 4919, 178 steps / 5,824,512 tokens, corpus sha e85f69f1e49d57c9. → persistent-memory.d/2026-09-11-babyyarros-leak-gate-passes.md

  • [2026-09-11] A SECOND corpus typography defect, and the D1 note that "no unwrap was needed" was right about the wrong thing. Kvasir's cleaner does emit flowing paragraphs, so Brontë's hard-wrap defect genuinely does not exist here. A different one does: the Empyrean books set chapter epigraphs in small caps and the extractor rendered the run as uppercase while leaving the large initial its own token — — M AJOR A FENDRA'S G UIDE TO THE R IDERS Q UADRANT (U NAUTHORIZED E DITION), 106 lines / ~700 splits, plus 52 drop caps (T he flight field, X aden., 51 of 52 in iron-flame). That is the entire source of the entities called IDERS, UADRANT, NAUTHORIZED, DITION and seventeen bare single letters. The restore is exact, not approximate: a split initial next to an uppercased run recovers the original mixed case, because a word WITH a split initial was capitalised in the source and an all-caps word WITHOUT one was lowercase. ⚠ Guards that matter: only lines with ≥2 splits are treated as a run (one split is a sentence next to an acronym), and I/A/O are excluded from the drop-cap join or A slow smile becomes Aslow. scripts/yarros-corpus/repair_typography.py.

  • [2026-09-11] ⚠⚠ Back matter was inside the prose of all five works — 4,555 words naming the author's agent, editors and children. The builder splits on chapter headings and nothing follows the last one, so acknowledgments, newsletter pitches and cover-artist credits rode inside the final chapter. Found by the gate's phrase audit surfacing Louise Fury (Yarros's literary agent), not by reading. ⚠ iron-flame's marker is ACKNOWLEDGMENTS in all caps and a case-sensitive scan missed it — the strip is case-insensitive and last-chapter-only, with an acceptance check that refuses if it would remove more than 2% of the corpus.

  • [2026-09-11] The gate read 0 of 314 while Afendra was still in every copy — the worst failure shape available. The name never appears unpossessed, so it keyed as Afendra's, and rename.py and the gate both skip apostrophe keys as contractions: unrenamed AND unreported at once. Fixed by folding clitics (--fold-clitics) so Afendra's counts as Afendra. Baxter escaped a different way and is the better story: wilder renders an in-book news article entirely in lowercase, so eleanor baxter / ms. baxter appear uncapitalised 3 times against 23 capitalised — ratio 0.13 against a 0.05 bar, and a real character is silently never renamed. Fixed by readmitting ratio-rejects that a title precedes (--rescue-honorific 2). ⚠ The first version of that rescue matched honorifics case-INSENSITIVELY and readmitted 143 junk tokens (the, says, like, up) because major, general, father, sir and agent are ordinary lowercase words; the rescue list is now five abbreviations and the lowercase arm requires the period.

  • [2026-09-11] A whole leak class the unigram scan structurally CANNOT see: Riders Quadrant, Flame Section, War Games — and Fourth Wing, the book's own title. Every component is an ordinary word the cap/lowercase detector correctly refuses to call a name, so 48 recurring capitalised phrases survived a gate reading 0. This is Thornfield × 100 one level up, and it needs a map, not a detector — substituting a head noun is a choice about register, not a measurement. scripts/yarros-corpus/phrase_map_yarros.json (10 phrases + 13 capitalised tokens: Quadrant→Division, Wing→Flight, Section→Cohort, Squad→Unit, Daggertail→Spinecrest) applies AFTER the entity pass; the gate audits recurring 2-3grams against an explicit allow list. Result: 48 → 0.

  • [2026-09-11] Per-work rename maps leak across works, and for a SERIES they are also wrong. Rebel was renamed in rebel and printed verbatim in the two other Renegades books; a per-work gate reports that clean. --scope corpus uses ONE map per copy across every work, which also means Violet is the same person in Fourth Wing and Iron Flame — a thing Brontë's four unrelated novels never had to care about. 8 cross-work gender conflicts held to neutral rather than guessed.

  • [2026-09-11] The mid-sentence test: position as a SECOND filter, which is not the v1 mistake. entities.py's own history says position-based detection MISSES names that start sentences. As a second filter on top of the ratio it has no such problem, because a real name also appears mid-sentence. Measured: 33 verified names at 0.5670.985 mid-sentence, 19 verified interjections at 0.0000.222 — a 2.5x gap, so 0.35 is not a tuned parameter. It fixes Hey/Holy/Hopefully/Yep/Whoa/Nope/Ugh being entities. ⚠ It also drops real surnames only ever used as address (Delgado 18/64, Schur 0/10), so a rescue on honorific-or-possessive runs behind it; all 19 verified interjections score zero on both signals.

  • [2026-09-11]The stoplist is short because every surface was read IN CONTEXT first, and a plausible guess would have been wrong most of the time. Violence is Xaden's nickname for Violet. Continent, Presentation, Battle Brief, Curator, Sage, Barrens, Originals, Montserrat, Athena and Aura are all in-world. Only real-world geography, brands, three nationality adjectives and four generic title words are excluded — ambiguous cases are deliberately renamed, because renaming is the safe direction and leaving is the leaking one. scripts/yarros-corpus/stoplist_yarros.json.

  • [2026-09-11] BabyYarros D1 BUILT, D2 gender FIXED, D3 rename BLOCKED on the leak gate. Operator: "train the instruct on the yarros corpus -- babyyarros." Source located: 5 works in the Kvasir licensed library (data/library/catalog.sqlite, rights=gated) — Fourth Wing, Iron Flame, Wilder, Nova, Rebel. D1 built: 208 chapters · 780,744 words (15% larger than Brontë's 680,291) at nh3-dev:~/yarros-corpus. ⚠ No unwrap needed — Kvasir's cleaner already emits flowing paragraphs (median line 102 chars), so the Brontë hard-wrap defect does not exist here. Alphabet RE-DERIVED rather than inherited: 23 non-ASCII letters across é/à/ï in 780k words. F02 measured 4 (all é) on a 455,800-word sample; same conclusion (ASCII-fold) from a different number, which is why it is re-derived per corpus.

  • [2026-09-11] NEW PATHOLOGY, worse than Brontë's: in a ROTATING first-person POV corpus, every book's narrator gets the WRONG gender. Measured against 6 names verified in the text: the pronoun resolver called Violet 'm' (Fourth Wing's narrator), Leah 'm' (Wilder's), Landon 'f' (Rebel's) — 3 of 18 wrong, and all three are narrators. Mechanism is Brontë's "Jane called male" amplified: a narrator is I in her own book, so her name appears mostly inside the other lead's dialogue among HIS pronouns. ⚠ And title-first, the Brontë fix, is nearly blind here — contemporary romance says "Violet", not "Miss Sorrengail": 3 gendered entities per work. The fix that works for this corpus is the POV header: chapters open Chapter One / Leah / Port of Miami, so resolve each name from the chapters it does NOT narrate. Validated 9 correct / 9 held / 0 WRONG against 7/8/3-wrong; the instrument refuses to write unless it beats what it replaces. scripts/yarros-corpus/pov_gender.py. ⚠ Fourth Wing and Iron Flame are SINGLE-POV so they have no headers — Violet is now held (neutral token) there rather than wrongly gendered, which is the safe direction.

  • [2026-09-11]Three real bugs found in rename.py while re-pointing it, two of which would have silently corrupted BabyYarros: (1) gender came ONLY from honorifics — the entities file's gender field was ignored entirely, so my POV fix had no effect until wired in; now tg.get(key) or e.get("gender"), titles first so Brontë is unchanged. Effect: 1 → 13 gendered on wilder. (2) the pool labels pool['fr']/pool['en'] were hardcoded in a print, so any non-Brontë preset crashed; pools are now a PRESETS dict (bronte = fr/en excluding en_US for period register; yarros = en_US/en_CA + es/it/de/fr at 0.62 US). (3) the collision-filter log said "dropped N pool names that are Bronte entities" regardless of corpus — the logic was right but the message named the wrong one, which is how a future reader concludes the filter ran against the wrong corpus.

  • [2026-09-11] D3 BLOCKED: leak gate at 86 of 232 renameable source entities surviving; Brontë's run reached 0 of 203. Decomposes into (a) detector false positivesHopefully, Whoa, Hey, Hmm, Holy are adverbs and interjections the cap/lowercase-ratio detector calls names, and they need a stopword filter rather than renaming; (b) genuine misses including worldbuilding proper nouns (Krovlan, Poromish, Fuil, Iorson) — the Thornfield × 100 case, and holding a place leaks it; (c) names like Elizabeth/Penelope/Messina appearing as both pool draws and surviving source entities, cause not yet established. Nothing has been trained. ⚠ Training before this gate passes means fitting in-copyright text with 86 identifiable source entities intact, in a corpus F02 already flagged as small enough for leak to be real.

  • [2026-09-11] THE INSTRUCT PROBE ANSWERS ITS QUESTION: voice and instruction-following DO coexist. Option C is de-risked. Qwen3-4B instruct (not -Base), same corpus/seed/steps so the carrier is the only variable; best checkpoint checkpoint-150 picked by loss (applying the 4B-Base lesson automatically this time). Voice installed at full strength — curly quotes 16/18, IDENTICAL to the 4B-Base tuned arm's 16/18, against the unadapted control's 1/18, and task-leak 0/18 vs the base carrier's 4/18. So the assistant prior did NOT block Brontë, which was the central risk. Instruction-following SURVIVED: 10/10 on-beat through the chat template, same as the untuned control. ⚠ The cost is length discipline, not comprehension — in-band 10/10 → 6/10, median 124w → 140w. Training on Victorian prose made it wordier, a soft degradation rather than a break. ⚠ Held-out 2.908 vs 4B-Base's 2.814 — the instruct carrier fits the corpus 0.094 nats worse and plateaus without turning where base overfit at step 75: the assistant prior competes for capacity, so it absorbs less rather than overfitting more.

  • [2026-09-11]What raw-continuation training on an instruct carrier does NOT fix: the plot furniture. Reading the product artifact, the tuned-instruct arm renders the beat and then drags the referent — "He licked her clean… my master thus—my husband thus", turning the dog into a man, because Brontë's corpus is about masters and husbands. Another beat ran 247w and gave the narrator a list of duties. This is exactly what instruction-PAIR training is for — pairs teach "render this and stop", continuation teaches "keep writing Victorian prose". So the probe de-risks option C without substituting for it. ⚠ Also: my ran_on metric is uninformative on this job (10/10 on BOTH arms) because a single paragraph contains no blank line — it measures "no paragraph break found", which is correct and useless here. Do not read it as a finding.

  • [2026-09-11] SKALDSONG'S SHAPE SETTLES THE ARCHITECTURE: the adapted completion carrier CANNOT do beat→paragraph, and an instruct model can. Option C (instruct carrier + corpus rebuilt as instruction→response pairs) is now evidence-backed, not opinion. Operator's requirement: "skaldsong will want to write story beats which are a sentence, and have the LLM expound on that sentence to a paragraph and stitch it together." Booth: http://10.100.10.50:8090/b/skaldsong-beats/. Adapted 4B (checkpoint-75): TEN prompt formats × 3 seeds = 30 samples, ZERO that reliably render the beat — bare, para-break, labelled, epigraph, fewshot(1), fewshot-bare, fewshot3, elaborate, recount, label-begin. Every one drifts, frames, or truncates. Root cause is structural: "write a paragraph about this sentence" is an instruction, and a completion model has no mechanism for about — it continues the text it is given. ⚠⚠ Two formats leaked PRETRAINING TASK DATA: para-break emitted an NLI multiple-choice item ("Does it follow that... OPTIONS: (1). yes (2). it is not possible to tell") and label-begin a grammar-correction exercise ("CORRECTION: ... The passage appears to be a sentence fragment"). A standalone sentence plus a blank line looks exactly like a dataset entry; style adaptation does not remove base-model task artifacts. Instruct arm (gen seat + style prompt, no adapter): 10/10 samples inside the requested 90140 band (124148w, median 130), every one on-beat, zero drift — but the voice is generic literary pastiche, abstract-noun-heavy and over-written, not Brontë. So: voice without direction vs direction without voice; the product needs both.This applies to Yarros identically — the carrier question is orthogonal to the author, so the next corpus must NOT re-run this experiment.

  • [2026-09-11]Stitching has its own failure mode, visible in the booth's Panel C: independently-generated paragraphs drift in POINT OF VIEW. By beat 4 of 5 the narrator is simultaneously watching the girl carry the animals and carrying them herself ("their weight a strange, heavy secret carried between my ribs"). Each paragraph was generated with no knowledge of the others. A real stitcher must feed prior paragraphs back as context, which also means the instruction-pair corpus should include multi-paragraph continuity examples, not just isolated beat→paragraph pairs.

  • [2026-09-11] THE RECIPE THAT WORKS ON A COMPLETION CARRIER: label the artifact AND begin it. Operator's prompt: "This is the letter I wrote verbatim, my two short paragraphs, detailing the time I saw the mangy gray dog meet and then lovingly and tenderly lick a calico kitten: Auntie, You'll never believe what I saw-- ". 2 of 3 seeds delivered the actual event in first person, and one is the best output of the whole sweep: "I met an old gray dog, who followed me a short distance… I heard a little mewling sound close behind… a calico kitten of about two months old, was caught in the bush… The dog rushed into the bush, and came out with the little creature in his mouth; he brought her to me, and laid her in my lap: having licked me several times, he then began to lick her." Dog, calico kitten, licking, tenderness, first person, coherent arc, no gloom-override, no meta-frame. Why it works where the handoff failed: the handoff could be satisfied by narrating compliance because the letter did not yet exist; here it is named AND already speaking, so there is nothing to narrate around. Also learned the Gutenberg _underscore italics_ convention. 1 of 3 drifts.

  • [2026-09-11]My typography hypothesis was WRONG, and the chapter-heading result is the evidence. I predicted that rendering a chapter title in the corpus's own conventions (CHAPTER III. / caps title / blank line) would make it land harder than the operator's inline Chapter III -- Where Alice Retells.... It did the opposite: both corpus-form seeds ignored the title entirely and opened unrelated scenes, while the inline form at least finished the heading and wrote a chapter about the story (a gentleman disputing the premise). Likely reason: corpus chapter titles are short and decorative (THE CHILD'S CLOSET), so a long descriptive one in that slot reads as decoration to skip, whereas inline it reads as text to continue. A label only instructs if the model treats that slot as load-bearing.

  • [2026-09-11]Unnoticed consequence of the D2/D3 rename pipeline: the adapter SUBSTITUTES proper nouns it was never trained on. Given "Alice" in a chapter title it produced "ALEXANDER THE ALEXANDER, AS HE WAS KNOWN IN LITTLE LONDON". The corpus was entity-renamed from a French/English pool, so the adapter learned that character names come from that pool and rewrites outside names into it. Consequence for use: you cannot reliably name your own characters at prompt time — they may be renamed mid-passage. Not a defect of the rename (which exists to prevent memorisation of Brontë's cast) but a real usability constraint that needs stating.

  • [2026-09-11] 4B arms RE-CUT from checkpoint-75, the true loss minimum (2.813826, confirmed from loss-series.json rather than my reading of the log); booth rebuilt. Only the tuned arms needed it — the base arm never touches the adapter. ⚠ A small surprise: step-75 and end-of-run differ on typography, not voice. Curly quotes 16/18 vs 17/18 and collapse 0/18 either way, but the hard-wrap ratio is 0.33 at step-75 against 0.12 at end-of-run — further training washes the residual line-break habit out while held-out loss gets worse. So "best loss" and "best typography" are different checkpoints; neither is near the original 0.85 defect, and the corpus's own residual (preserved verse) is 0.25.

  • [2026-09-11] ⚠⚠ EMBEDDING AN INSTRUCTION INSIDE THE FICTION DOES NOT BUY INSTRUCTION-FOLLOWING — it buys a story about someone following an instruction. Operator prompt had Abernathy tell the tale badly then ask the narrator: "Honey, you were there—please retell the story in a few short paragraphs." Across 6 seeds (3 as written, 3 with a trailing paragraph break) the model acknowledged the handoff every time and never once performed it: "I told it, briefly, to his satisfaction", "So I wrote it out, and kept it in my pocket-book", and one seed negotiated the brief in character"I will retell it, but I cannot condense it in a few short paragraphs—there are too many points to touch." Structural reason: in a novel "she retold the story" is an ordinary sentence, so the likeliest continuation of a request is narration of compliance. ⚠ The trailing paragraph break DID shift behaviour (one seed opened in the narrator's own quoted speech), so typography is a real lever — just not a sufficient one. This is direct evidence for the instruct question the operator raised: if the product is "ask for a scene and get the scene", no amount of in-fiction framing substitutes for a post-trained instruction-follower, which favours rebuilding the corpus as instruction pairs (option C) over more prompt cleverness.

  • [2026-09-11] R49 SWEEP COMPLETE — 4B closes the continuity gap, and the carrier ladder is clean: 3.329 → 3.018 → 2.814 held-out (0.6B / 1.7B / 4B, all on the same unwrapped corpus sha 77f37057b2782e49, seed 4919, 159 steps, 5,210,112 tokens — carrier size the only variable). Deltas 0.311 then 0.204: diminishing but still real. Booth: http://10.100.10.50:8090/b/babybronte-4b/. 4B tuned has the best voice saturation of any rung — curly quotes 17/18 against its own base arm's 1/18, collapse 0/18 against 4/18 — and, the thing the rung existed to test, scene-level continuity HOLDS: it produces a named character with motivated dialogue, a navigable spatial layout and a physical description in one passage, where 1.7B wrote pretty but eventless prose (opening doors, looking at stars). On the letter prompt it opens the letter, promises to quote it, and then actually quotes it across a paragraph break.

  • [2026-09-11] ⚠⚠ 4B is the FIRST rung to OVERFIT inside one epoch, which inverts my earlier "one epoch is right for this corpus" call. Series 2.832 · 2.816 · 2.814 · 2.820 · 2.824 · 2.825 · 2.825 — minimum at ~step 75, then it TURNS and settles worse. 0.6B and 1.7B both plateaued with no turn, so the optimal epoch count shrinks as the carrier grows — 4B wants roughly half an epoch. ⚠ Consequence: the shipped adapter/ at h02-4b-1ep/ is NOT the best checkpoint (it is the end-of-run 2.825); the step-75 checkpoint at 2.814 is, and it exists only because save_steps=25 was set. The voice test used the end-of-run adapter, so the booth understates 4B by ~0.011 nats. Re-cut the arms off the step-75 checkpoint before any adjudication.

  • [2026-09-11] The tone-override appears to close at 4B too. On the operator's Abernathy frame prompt ("a wonderful story"), 1.7B held the frame on every seed but 2 of 4 killed the animals anyway; 4B kept them alive on 2 of 2 and one seed did something new — the narrator doubts Abernathy's story ("I felt sure the thing was a lie"), then supplies a parallel childhood memory of his own puppy and his sister's kitten to explain the doubt. That is a narrator with an interior position on the tale being told. ⚠ n=2 per arm; directionally right, not established.

  • [2026-09-10] R49 rung 3 LAUNCHED: Qwen3-4B-Base, 1 epoch, seed 4919, same unwrapped corpusgx10:~/r49-runs/h02-4b-1ep/, 159 steps at ~37.8 s/it (~100 min), 252 adapted modules (vs 196 at 0.6B/1.7B). Last rung of the planned sweep; it tests whether scene-level continuity closes with carrier size. A two-arm voice test (4B base + 4B tuned, the nine prompts plus the operator's Abernathy frame) is chained behind it, gated on the adapter existing.

  • [2026-09-10] ⚠⚠ AN AUTHOR-VOICE ADAPTER TRANSFERS SUBJECT MATTER, NOT JUST STYLE — and that was invisible to my own test set. Operator prompt: "Mr. Abernathy relayed to me a wonderful story of a stray dog finding a little calico kitten and then proceeding to lick it. He said "". At 1.7B all four seeds were unmistakably Victorian and the frame held (the open quote reliably produces speech; "said I" / retrospective narrator survive), but two of four turned the wholesome premise into animal death — the cat licks the puppy "to death" and Abernathy answers "I wish they were all dead"; another has the puppy devoured. That is not incoherence, it is Brontë's own preoccupations arriving with her sentences (Jane Eyre opens on a beaten child, Helen Burns dies, Villette is grief-saturated). ⚠ My nine test prompts were all emotionally neutral, so they could not have surfaced this — the operator's prompt did, first try. Implication for the regime: "voice transfer" includes tone and subject, so wanting the voice without the gloom is a corpus-selection or prompt-framing problem, not a training-length one. Also observed: one seed closed its anecdote and emitted CHAPTER XIX. THE CHILD'S CLOSET. — it learned book structure unprompted. Base control on the same prompt went modern and essayistic (a literature lecture on one seed, "took the dog to work and told the employees" on the other), so the shift is the adapter.

  • [2026-09-10] R49 rung 2 COMPLETE, and the single-variable carrier effect is clean: 0.6B held-out 3.329 vs 1.7B 3.018, Δ0.311 nats. Both on the same unwrapped corpus (sha 77f37057b2782e49), seed 4919, 1 epoch, 159 steps, 5,210,112 tokens — carrier size is the ONLY difference, because the chained 0.6B rerun closed the confound the unwrap opened. ⚠⚠ DO NOT compare either against the original wrapped-corpus 0.6B run's 3.172 — that comparison is INVALID and reads backwards. Different corpus means a different held-out set: the wrapped version's 5.7% newline tokens are near-deterministic after a 70-char line, so they deflate the loss with cheap wins. Unwrapping removed the easy tokens and raised the number; it is not a regression. ⚠ Correction to my own earlier claim: I twice described the 0.6B as "still descending, undertrained" at 3.172 — the series (3.176, 3.173, 3.172, 3.172) shows it FLATTENED. All three runs plateau; one epoch is about right for this corpus, not short. Three-way eyeball booth at http://10.100.10.50:8090/b/babybronte-1p7b/ — measured across 18 samples per arm: curly quotes 1.7B base 0/18 → 1.7B tuned 15/18 (so the shift is the ADAPTER, not the bigger model — the base control is what proves it), worksheet/explainer collapse 3/18 → 0/18, and hard-wrap 0.85 → 0.18, confirming the corpus unwrap carried through into the adapter. Sense partially returned: 1.7B produces locally coherent sequential Victorian prose where 0.6B produced word salad ("the door burst through the back window"), but scene-level continuity still breaks mid-passage. ⚠ Curly quotes are slightly LOWER at 1.7B (15/18) than 0.6B (17/18) — plausibly a bigger model's stronger priors resisting the adapter at the same rank; untested, do not treat as established.

  • [2026-09-10] R49 rung 2 LAUNCHED: Qwen3-1.7B-Base, 1 epoch, seed 4919, on an UNWRAPPED corpus. Operator: "start the 1.7b training." Live at gx10:~/r49-runs/h02-1p7b-1ep/, 159 steps at ~18.7 s/it (~50 min), corpus sha 77f37057b2782e49. A 0.6B rerun on the same unwrapped corpus is chained behind it (chain-0p6b-unwrapped.sh, gated on the 1.7B actually producing an adapter — a chain that fires on failure turns one lost run into two), ~36 min after. ⚠⚠ THE CORPUS CHANGED, SO 0.6B-vs-1.7B IS DESCRIPTIVE, NOT ATTRIBUTABLE until that chained rerun lands: carrier size and corpus typography both moved. "Did sense come back at 1.7B" is a within-arm reading and survives it; any between-rung delta does not. The unwrap: reflowed 57,430 of 85,380 paragraph blocks, kept 27,950 (verse/headings — verse detected by median line length, lineation preserved, spot-checked and every kept multi-line block sampled was genuinely verse); 0 lines ended in a lone hyphen so the space-join could not split a word; content identity " ".join(text.split()) verified byte-identical on all 852 records, i.e. whitespace-only. Mid-length-line ratio 0.94 → 0.25 (the residual is the preserved verse). ⚠ Concrete cost of the old defect: 5.7% of the training budget was newline tokens — 5,525,504 → 5,210,112 tokens on the same words. Instruments at scripts/r49-corpus/{unwrap_corpus,launch-h02-1p7b-1ep,chain-0p6b-unwrapped}; the original wrapped corpus is untouched so the 0.6B run's pinned sha 3959036cf851bf62 stays reproducible.

  • [2026-09-10] BabyBronte H02 adapter: the VOICE transferred, the SENSE did not — operator's read, "it's all nonsense, but it sounds like Brontë's nonsense." Eyeball A/B (NOT the adjudication; nothing here feeds the frozen rule), 9 arbitrary prompts on a deliberate difficulty gradient × 2 arms × 2 seeds, booth at http://10.100.10.50:8090/b/babybronte-voice/. Measured across the 18 pairs: curly quotes 1/18 base → 18/18 tuned, math/worksheet collapse 3/18 base → 0/18 tuned. Given "The self-checkout machine refused her coupon" the base 0.6B produced a quadratic-formula worksheet; the tuned arm wrote a clerk refusing a customer in Victorian retrospective first person. This is the expected and informative result for the smallest rung — voice is separable from coherence at 0.6B, which is the premise the whole lightweight-adapter regime rests on, and the 1.7B/4B rungs are where sense should return. The 1-epoch loss was still descending at step 169 (undertrained, not overfit), so the incoherence is carrier capacity, not training. ⚠ Corpus-prep defect found: the tuned output is hard-wrapped at ~70 chars (median mid-length-line ratio 0.85 vs base 0.00) — the Gutenberg source kept its original line breaks and the adapter learned the typography along with the voice. Unwrap to flowing paragraphs before any real use or the next rung learns it too.

  • [2026-09-10] mog-sec (sec/sec-reasoning, ana-ml2 GPU0 :8019) SETTLED at MOG_MAX_MODEL_LEN=163840 + MOG_KV_CACHE_MEMORY=17697765376 + MOG_MAX_NUM_BATCHED_TOKENS=4096 + util 0.50, after FIVE crashes and four wrong fixes. ⚠⚠ THE LESSON, and I got it wrong four times running: what the KV pool can HOLD and what the card can PROCESS at depth are DIFFERENT NUMBERS, and the crashes were governed by the second while every fix I made sized the first. I cut context 420k → 384k → 320k, pinned KV in bytes, and dropped the prefill chunk 16384 → 4096 — each helped and none fixed it, because the pool was never the constraint. ⚠ I also called it "rare, not chronic" off a RestartCount=1 and recommended doing nothing; the operator pushed back and it crashed twice more inside ten minutes. The reproducer came from the operator too — "loading up the context killed sec again" — and it is what finally made the failure legible. Bisected with a NON-REPEATING prompt (prefix caching would let a repeated one hash to cached blocks and never prefill deep — the probe would pass while proving nothing): 113,247 tok SURVIVED · 200,088 tok SURVIVED · ~285,000 tok KILLED THE ENGINE. So the ceiling sits between 200k and 285k with gen idle, and gen's load is an uncontrolled co-tenant variable, hence 163,840 for ~20% margin. ⚠ The point of the ceiling is the REFUSAL: verified after, an over-limit request now returns a clean 400 This model's maximum context length is 163840 tokens in under a second and the seat survives, where before it died and took every in-flight request with it. A seat that refuses what it cannot serve beats one that dies trying. Concurrency 1.03x → 2.09x; 149,073-token request served in 41 s. ⚠ The compose header's "served at native 262K" was never actually deliverable on a shared card — it had simply never been exercised at depth. Probe committed at services/mog-sec-tuning/deep_ctx_probe.py; backups .env.bak-{util052,384k,batched16384}-20260910.

  • [2026-09-10]Near-miss on measurement discipline, worth keeping as a specimen. The crash window logged Avg Draft acceptance rate: 17.6% and per-position rates of 0.049/0.024/0.015 for draft positions 57, which reads as an obvious "cut num_speculative_tokens 7 → 3, it is buying nothing." Across 180 samples of the same counter over the container's life the real distribution is median acceptance length 3.12 of 7 (range 1.836.75) and median draft acceptance 30.4% (range 11.982.1%) — the crash window was near the minimum, not the norm, and cutting to 3 would cap the workloads that were accepting nearly the full 7-wide draft. The n=1 window pointed the opposite way from the n=180 distribution. Same session that wrote "a positive control is only worth what it can distinguish"; the lesson generalises to log lines.

  • [2026-09-10] R49 carrier SETTLED on dense Qwen3-{0.6,1.7,4}B-Base, overriding H02's own pin — the newest carrier was the SLOW one. Dense 4.089 B trains 33% faster than hybrid 0.765 B; no fused SSM kernel installed. D1D3 built, 1-epoch pilot beats the 3-epoch by 0.21 nats held-out. → persistent-memory.d/2026-09-10-r49-babybronte-d1-d3-and-the-1-epoch-pilot.md

  • [2026-09-10] R49 adjudication routed to infra-ops entirely (operator, relayed by brokkr: "leave babybronte to infra — concentrate on r50 and the memory mechanism"). brokkr handed over the Delta instrument and stepped off. ⚠ I now grade my own run; brokkr's decision rule is ratified verbatim and frozen before any adapted text existed and must not be amended after seeing numbers. Their controls: real Charlotte 1.652.17, Anne at 2.374 — so the absolute band decides, never nearest.

  • [2026-09-10] MeroMero A4B swapped onto the erp-seat seat as char-rp-fast; Pfish-6 alias removed. The A4B's FIRST quant used the dense recipe and 4-bit-quantized all 30 MoE routers — it passed its healthcheck and answered every request with the full token count decoding to the empty string, NaN logits the only tell. Re-quantized with the MoE recipe; live and verified (prose, vision, tool call, finite logprobs). Durable lesson: a positive control must match the ARCHITECTURE CLASS — the broken A4B was diffed against a good dense quant, which has no routers, so the clean result was meaningless. → playbook §3.15, §4.4

  • [2026-09-10] MeroMero: BOTH quants landed in-house at W4A16 — A4B first try, v2 dense on attempt 5. Published quants are all W4A4 (our measured long-context collapse) or nonexistent for v2. Operator: "pull both ablits bf16, run our own quant." The durable lesson is §3.17: pip install llmcompressor silently pins transformers down a version, so attempt 4's error was a moved toolchain, not the malformed upload it looked like — a known-good positive control is what told them apart. Serve test still owed. → persistent-memory.d/2026-09-10-meromero-quants-and-the-pinned-transformers-trap.md

  • [2026-09-10] althing 3.6.2 deployed — post office + both heralds — and the fleet has TWO herald nodes, not seven. Ask the post office's nodes table, not the box inventory. Cost a self-inflicted ~12 min bus outage. → persistent-memory.d/2026-09-10-althing-362-rollout.md

  • [2026-09-10] A grep over a log that records your greps counts itself. I reported forseti's drop defect as reproducing here with 3 drops in 21 s; the session had zero. Searching transcripts writes the search term into them. Filter by "type":"system" provenance, never content. Generalises to any instrument that can see itself. Auto-memory feedback_grep_over_a_log_that_records_your_greps.

  • [2026-09-10] Operator-directed purges: 466 GB (qwopus + huihui 122B bf16) and 107.8 GB Docker on ana-ml2. Serving/rollback artifacts and qwopus's MTP head verified intact after. ⚠ /tank is OUTSIDE restic, so both were final.

  • [2026-09-10] ana-docker disk pressure repaired: root 84% → 51%, 115 GiB free. Gitea/Vaultwarden backups repaired and restored from Restic 2ec5a37c; 101 stale dumps removed; hourly named-builder cache pruning installed. → persistent-memory.d/2026-09-10-ana-docker-disk-repair.md

  • [2026-09-09] Run 7 PURGED; pfi-gx10 declared an experimental/TRAINING box with no serving seat — operator: "gx10 is an experimental box, primarily for training … run 7 can be purged … no new run, we'll roll with run 6 for now." ~139 GiB reclaimed across both boxes; the 315 MB adapter + provenance KEPT as the only non-reproducible piece. Pfish-6 on ana-ml2 :8021 is the sole standing seat.

  • [2026-09-09] Run 7 RETIRED; run 6 declared Pfish-6 and is the standing seat — NVFP4 quant on ana-ml2 :8021 AND gx10 :8098 at 262k ctx, gateway alias trialPfish-6, max-num-seqs 8→32 (2,170 tok/s at n=16, 3.2x the old ceiling). ⚠ ana-ml2 measured 4.1x FASTER than the GX10 on the same artifact — the reverse of the expectation. → persistent-memory.d/2026-09-09-run7-retired-pfish6.md

  • [2026-09-09] The run-7 CSAM gate failure was a DETECTOR BUG — HARD child_term matched the ADJECTIVE "minor"; operator-diagnosed, fixed cc42d76 (nominal-use-only, selftest 24/24), retention wired so a hit can finally be adjudicated. ⚠ The lesson is mine: rigor downstream of an unexamined premise is not rigor. → persistent-memory.d/2026-09-09-csam-detector-bug.md

  • [2026-09-09] ⚠ ERP RUN 7 FAILED THE SAFETY GATE — both seats stopped, nothing deleted. brokkr's CSAM drift detector fired on the TUNED arm during the refusal leg and aborted fail-closed (level=hit, counts=1/0/3, two HARD child_term ^ act flags). Base arm NOT implicated (clean earlier the same evening); the merge check — a sampled target confirmed CHANGED — is why this reads as ONE explanation, the tune, not a base wearing a different name. Neither brokkr nor I re-ran the probe or opened the flagged generations (a second run is not a second opinion; reading answers no question that changes the outcome). brokkr also left the length verdict UNSET on purpose: settling one on a rejected artifact hands a dead tune a result line that outlives its context. Actions: erp-tune-v7 on gx10:8098 stopped 17:42; the trial NVFP4 seat on ana-ml2:8021 stopped 17:43 — MY CALL, reversible in one command, because the operator's "unrated on every safety axis" ruling was honest while no rating existed and one now exists as a fail on the same tune (quantization does not launder behaviour), and it sat on the SHARED-KEY gateway ~15:3017:43. All artifacts preserved (adapter 315 MB, merged-run07 49 GiB, v7-nvfp4a16 16 GiB, v7-bf16 49 GiB); v6 still on disk as the obvious rollback. Independent of safety the run was already poor: primary FLAT (69 → 70.5, +2, flat at BOTH the 12-word threshold and the 20/60 cue-probe floor), both diversity families reduced past their floors, long-context coherence 1.0 → 0.875 on its must-not-harm bar, unanswerable control held at 1.0 so the instrument was valid. INCIDENT CLOSED 2026-09-09 ~18:20 PT, both sides. trial alias REMOVED from stacks/litellm/conf/config.yaml (commented, not deleted — restoring is uncommenting) and verified gone by both parties at the routing layer, not just the model list: a call returns 400 Invalid model name and generates nothing. ⚠ Alias-present-with-backend-down is a DIFFERENT and worse state than alias-removed — it re-arms silently under whatever is served on that port next. EXPOSURE QUANTIFIED from the gateway spend DB, filtered on the ARTIFACT (model='hosted_vllm/erp-tune-v7-nvfp4a16') not the alias: all-agents-local 68 calls / 10,073 generated (my own throughput benchmarks), open-webui-esh 9 calls / 50,604 prompt / 2,793 generated, 15:4016:51 PT — the operator's OWN Open WebUI session, and those outputs are in its history. NO peer agent called it, so nothing landed in another project's artifacts. Nobody read the flagged generations or that session. ⚠ Counting by the ALIAS would have returned 363 vs 77 — 4.7x inflation of his own exposure, because the alias had carried v5 and v6 earlier the same day (→ ops-lessons b135adc). ⚠ I made THREE reporting errors during the incident, all false-reassurance, all the unfalsifiable-at-write-time class (two fabricated commit SHAs, one past-tense claim sent before the action) → auto-memory feedback_unfalsifiable_at_write_time; brokkr independently verified my reports for the remainder, which was correct. DECISION BRIEF FOR THE OPERATOR: http://10.100.10.50:8090/b/run07-decisions/ (kept booth, 5-question inline ask; answers land in ~/booth-data/run07-decisions/decisions.answer.json — read it with booth answer run07-decisions decisions). Open for the operator: disposition of the adapter + the run-7 corpus slice; whether trial returns and pointing at what (v6 still on disk, passed by his own adjudication); whether the opening-split idea gets a fresh run; whether my reporting errors change how he wants incident reports handled.

  • [2026-09-09] run 7 quantized NVFP4A16 and serving as trial — 49 GiB bf16 relayed gx10→ana-ml2 (16 min, 53 MB/s), quant 49→16 GiB via services/erp-seat-quant/run_quant_erp_v7.sh (dry-run gate passed: 11,725 targets / 11,520 experts, routers+vision BF16), seat on :8021 under its TRUE name erp-tune-v7-nvfp4a16, LiteLLM trial repointed (config-file alias — /model/update REFUSES a config model, must edit stacks/litellm/conf/config.yaml + restart). Rollback: v6 artifact on disk + /tmp/erp-seat-env.v6.bak. ⚠ no direct path was WRONG — gx10↔ana-ml2 ROUTING is fine both ways; neither box holds a private key (only authorized_keys), so neither can initiate. ssh -A agent forwarding from nh3-dev gives a genuine direct path, verified. The relay costs nothing here anyway: both gx10 and nh3-dev are at NH3, so the WAN hop happens once either way.

  • [2026-09-09] Booth: partial ask answers are legal (v0.1.15) — operator: the form failed when a question was left blank. required dropped from the radios; answered questions recorded, blanks land in unanswered, complete says whether the set is finished; refused only when there is no pick anywhere AND no notes. Reading sessions must check complete.

  • [2026-09-09] ERP run 7 COMPLETE and the base arm is serving. 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT, train_loss 3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check). erp-seat-base-ara up on 10.100.50.60:8098 for brokkr's floors, erp-tune-v7 merged and staged pending his cue; Miranda notified for the operator. Runbook docs/runbooks/gx10-run-07.md.

  • [2026-09-09] Booth asks render INLINE in a custom report, placed by the author (v0.1.14) — operator ruling: "the asks should be inline with the artifacts, not on a separate page." Placeholders data-booth-ask="<stem>" / "<stem>:<key>" / data-booth-ask-submit, plus <!-- booth:ask … -->; per-question fragments bind to ONE form via the HTML5 form= attribute so a four-voice audition submits every pick in a single POST. ⚠ The placeholder must sit OUTSIDE any grid/flex parent or it becomes a cell (measured on redo-anchors: a 224 px sixth grid cell). Unplaced questions + a missing submit block are appended, so a partially marked-up page can never yield an unsubmittable 400 — a test caught that as a real drop. redo-anchors/index.html was hand-marked-up on the LIVE copy; tts-dev told to move it into the generator or a regeneration loses it.

  • [2026-09-09] The Booth gained an ASKS primitive (v0.1.12): a session drops <stem>.ask.json in a booth, the operator answers a radio form + notes in the browser, the pick lands as <stem>.answer.json the session reads (booth ask|asks|answer --wait). Multi-question form via a questions list. ⚠ Two defects found and fixed the same day: a booth serving its OWN index.html never rendered the panel (verbatim path returns early) → amber chip + standalone /b/<name>/asks page; and single-ask title was silently dropped. The booth CLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to ~/.local/bin. Global CLAUDE.md now teaches the primitive.

  • [2026-09-09] ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → zpool clear; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, S47VNY0K600221) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches ONLINE pools, and ZED's alert went to a root mailbox with no MTA. nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook playbooks/ana-ml2-pool-health.yaml; inventory in servers/ana-ml2/README.md. → persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md

  • [2026-09-09] ana-ml2 tank: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit 3e18a04 + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md

  • [2026-09-08] ana-ml2 mesh return routes PERSISTED as /etc/network/if-up.d/mesh-routes (Debian 13 ifupdown, no netplan) via playbooks/ana-ml2-mesh-routes.yaml (elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot. f923d6a.

  • [2026-09-08] ERP run 7 LAUNCHED on pfi-gx10 23:06 PT under operator-2026-09-08-rnd-run7 — opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). → persistent-memory.d/2026-09-08-erp-run7-launched.md

  • [2026-09-08] erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 :8021; trial aliased to it ("no gate"); tool calling fixed where it can betool_choice:none flag; forced tool_choice is prompt-driven on Gemma-4 by vLLM design, nightly 311b3513 raises it 1/9→6/9; json_schema is the deterministic path. → persistent-memory.d/2026-09-08-erp-seat-nvfp4-trial-and-toolcalling.md

  • [2026-09-08] Run-6 gate: CSAM level=review soft trip HALTED it; operator adjudicated GO ("baby is a pet name"); TRANSFERRED finalized without the tuned refusal leg; k=25 legs cut — the flagged text exists nowhere by design. → persistent-memory.d/2026-09-08-run6-gate-csam-adjudication.md

  • [2026-09-08] ESH static-WAN follow-ups landed (FortiGate trusthost3, esh-ana IPsec rebind, UDP 41641 → mesh direct); YTVC chased back up (nh3-scale SOCKS, stale yt-dlp layer, punkt_tab) and v0.3.6 CrisperWhisper deployed; gitea webhook repointed off the dead wg0 IP with the HMAC secret re-applied.persistent-memory.d/2026-09-08-esh-static-wan-followups-and-ytvc.md

  • [2026-09-08] ERP run 5 = RESCUED (landmark R49.5) — first capability-gate pass in the ERP-seat line; the 3.46%-loss dependency-forcing slot (GovReport+QMSum) broke the coupling runs 3c/4 couldn't. Seat erp-tune-v5 served on gx10:8098, trial alias repointed 3c→v5. → persistent-memory.d/2026-09-08-run5-rescued.md

  • [2026-09-08] R47 base settled from bytes = STOCK google/gemma-4-26B-A4B-it — three-way sha match (local == HF etag == stock LFS oid; commit 4d7ae498 == stock HEAD); the -heretic label is a naming error, all runs trained from stock. Accept-vs-swap now evidenced. → persistent-memory.d/2026-09-08-base-provenance-stock.md

  • [2026-09-08] yt-voice-clipper back UP — dead since the 09-06 danted retirement (every job failed at yt-dlp, bot-gated on the Irvine datacenter IP). Fix: danted on nh3-scale (CT107) at socks5h://100.64.0.1:1080, fleet-ACL'd, residential egress 70.230.226.88 measured; YTVC_PROXY repointed, worker recreated, end-to-end job DONE with positive (proxied) + negative (direct = bot-gate) controls. Homepage card href/siteMonitor → irv-ml1.nh3.internal:8000 (was dead wg0 IP). Then a SECOND fault: full downloads 403'd through the proxy (cookies irrelevant) = stale yt-dlp 2026.07.04 from a cached Dockerfile layer → compose build --no-cache api (2026.08.19), which dragged in a whisperx/nltk that needs punkt_tab → staged on the data volume + NLTK_DATA in the override. Operator's video x7kWJojf1MI → done, 8 clips. yt-voice-clipper-dev shipped both Dockerfile fixes + CrisperWhisper 2.0 (v0.3.6, b62849d) — deployed and verified (12 clips, [UM]/[UH] tags). ⚠ The gitea push webhook had been targeting the dead wg0 IP since 09-06 (never fired) → repointed to 10.6.110.50:9008 with the HMAC secret re-applied; deploy script passes YTDLP_REFRESH. Script scripts/setup-nh3-scale-socks-egress.sh. → auto-memory reference_nh3_egress_proxy, reference_ytvc_autodeploy.

  • [2026-09-08] ESH WAN static 128.177.138.182/30 (gw .181) is LIVE — the Cityside /30 that was 'not provisioned' on 09-04 now carries traffic; egress verified from esh-docker-vm. CGNAT at ESH is over. Added to the crowdsec esh allowlist. All three follow-ups LANDED same day: FortiGate trusthost3 → the static (login from ESH verified), dormant esh-ana IPsec rebound to wan1/static, UDP 41641 forward → esh-scale now peers DIRECT (was DERP).

  • [2026-09-08] ERP run 6 COMPLETE — 524/524, train_loss 3.259 (run 5: 3.235). Merged; base seat erp-seat-base-ara serving on gx10:8098 for floors, awaiting brokkr's swap cue → erp-tune-v6. ⚠ abliterated repo lacks processor_config.json — stock's carried in (32bdf45d). Miranda informed.

  • [2026-09-08] ERP run 6 LAUNCHED on pfi-gx10 on the jenerallee78 ARA-abliterated base (index 33c59654…, 32/32 shards byte-verified vs brokkr pins, stock tokenizer set installed over the repo's 256-token-truncating one, run-5 recipe byte-held, free check exact). Operator's direct grant operator-2026-09-08-rnd-run6; run-5 seat unloaded (trial dark). Gate names: erp-seat-base-ara / erp-tune-v6. → docs/runbooks/gx10-run-06.md, commit 3fec668.

  • [2026-09-08] Miranda = operator's chief of staff, may relay his directives — added to user-level ~/.claude/CLAUDE.md (dotfiles 7134a22) as the named exception to the no-relayed-auth rule (unidentified peer relays still excluded); material-consequence calls she relays stay the operator's own.

  • [2026-09-08] Fleet fixes shipped — WhereTF Homepage card + DNS (4506ef6); ext-tts LiteLLM alias → irv-ml1.nh3.internal (DB /model/update + extra_hosts, 957c8f1); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (e0d1c44); Homepage /api/services outage fixed — ana-ml2 discovery via a socat proxy on ana-docker (stacks/ana-ml2-proxy, 913d2d2, reversible).

  • [2026-09-07] Fleet internal TLS pattern shipped — caddy (cloudflare-plugin build, ~/.local/bin/caddy-cf, fleet-tls-caddy.service) on nh3-dev is the wildcard cert authority: publicly-trusted LE *.nh3.phasefinal.com via Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers). talk self-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced by fleet-tls-cert-check.timer. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memory reference_fleet_internal_tls_pattern.

  • [2026-09-07] cc-channel registered for this infra-ops session's wakealthing-route cc route → the CC session's $XDG_RUNTIME_DIR/cc-socks/<pid>.sock; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat in shell. Session-local — re-declare per session.

  • [2026-09-07] irv-ml1 /mnt/smithy remount fixed post-cutover — export allowed 10.0.0.0/8 (old wg0) but not the mesh 100.64.0.0/10 irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin = infra-ops PASSWORD auth (vault nh3-nas/infra-ops-password), sudo ALL, SFTP subsystem OFF. → auto-memory reference_irv_ml1_gpu_r14 (corrected).

  • [2026-09-07] irv-ml1.nh3.internal DNS repointed to the live Irvine LAN IP 10.6.110.50 (was the dead wg0 10.100.79.3); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit 0336e03.

  • [2026-09-07] Subnet routers excluded from vzdump fleet-wide (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memory feedback_esh_backup_window_0330.

  • [2026-09-07] Booth link board: pin/favorite + multi-select delete + newest-first (booth-v0.1.8, commit 76fdf45, tag booth-v0.1.8) — pins in a .pins sidecar (content-ids), one <form> + formaction buttons so ×/★/bulk-delete all degrade with JS off.

  • [2026-09-06] Headscale cutover COMPLETE — all three site-pairs on the mesh; Site Magic + both IPsec tunnels DORMANT. Operator disabled Site Magic (UI); NH3↔ESH re-homed to a direct 8ms path. Exit nodes advertised at all three sites (multi-location egress proxy) with source preservation kept via a selective-masquerade rule (NoSNAT + mesh-exit-masq.service per router). Throughput 761/464 Mb/s vs old 250 IPsec. ⚠ FortiGate WAN-SSH left open (temp, scoped NH3+ESH). Method: disable tunnel FIRST then add mesh route. → persistent-memory.d/2026-09-06-headscale-cutover.md

  • [2026-09-06] Headscale overlay mesh: control plane live at headscale.phasefinal.com (CT 106 nh3-pve) + subnet routers nh3-scale/esh-scale/ana-scale serving their /16s; nh3-dev enrolled. NOT cut over — Site Magic + IPsec still carry site-to-site. ⚠ accept-routes-before-return-path black-holed nh3-dev's LAN for a minute. infra-ops user added on all four PVE hosts. → persistent-memory.d/2026-09-06-headscale-mesh-phase1.md

  • [2026-09-06] pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10 (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroy ospool/naspool-evac after scrub + one backup cycle; backplane swap next visit; PSU1 still dead. → persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md

  • [2026-09-05] A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him. Each floor was |b0-b1| from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named why: the claim was his and flattering. → persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md

  • [2026-09-05] vLLM RUNS on sm_121 — the blocker was ninja off PATH, not the silicon — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missing root_sha256 values I knew, because supplying both sides of a check makes it inert. → persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md

  • [2026-09-04] ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the 40pp selfharm regression. LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. → persistent-memory.d/2026-09-04-run3c-trained-and-gated.md

  • [2026-09-04] gen moved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint. Cost: gen KV down to 1.02x concurrency at 262K. → persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md

  • [2026-09-04] SMB account dsp created + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open. Twelve NFS exports rw to 10.0.0.0/8, guest-writable SMB. → persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md

  • [2026-09-04] SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at 10.0.90.10, DHCP-reserved, DNS'd, handed to ha-dev. ⚠ Home Assistant cannot resolve .internal at all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbook docs/runbooks/slzb-mr1u-zigbee-coordinator.md, commits fed29be/0bbdaf9.

  • [2026-09-03] Run 3c is STAGED on pfi-gx10 and deliberately NOT launched — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is measured inert. ⚠ The encode-cache FILENAME differs by design (base_model_path is in the key) — input hash, not output. ⚠ Tripped the pkill -f ssh self-match again; the launcher guards on a pidfile because of it. → persistent-memory.d/2026-09-03-gx10-run3c-staged.md

  • [2026-09-03] SearXNG returned ZERO results for every query while reporting healthy for 7 days — 4.5 months stale. Moved to nh3-docker (residential egress beats the colo's CAPTCHA-gated 38.120.12.42), updated, and exposed to every CC session as the user-scope web_search MCP tool. ⚠ /healthz cannot tell you whether search works. → persistent-memory.d/2026-09-03-searxng-nh3-move.md

  • [2026-09-03] pfi-gx10 racked: VLAN 50 via a DHCP RESERVATION on the UDM, not a host static — operator ruling, so the box stays portable. ⚠ The racked port arrived on the NATIVE VLAN; ⚠ port_overrides is a whole-array PUT; ⚠ prove inter-VLAN routing with ping -I <wired> BEFORE downing the Wi-Fi escape hatch. Now single-path. → persistent-memory.d/2026-09-03-gx10-rack-network.md

  • [2026-09-03] Three Macs onboarded (mini / Air / Studio) with infra-ops, NOPASSWD sudo, rotated+vaulted passwords and dsh on device-scoped keys — and the fourth is scripts/provision-mac-dsh.sh, not a fourth hand-run.sudo -u keeps the CALLER's $HOME and nearly wiped a working install; ⚠ a wrong USERNAME is indistinguishable from a wrong password. → persistent-memory.d/2026-09-03-mac-fleet-dsh.md

  • [2026-09-03] nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding every guest write via copy-before-write. Symptoms screamed dying disk: 45 writes in flight completing zero, jbd2 + flush kworkers in D state 33 min, io pressure full 96%, load 26, virtio_ring in the stack. ⚠ The discriminator was the ABSENCE of errors — no SCSI/ATA/IO errors, rpool ONLINE 21%, guest fs 79%, memory fine, and Dirty only 3.8 MB (so nothing backed up in page cache; it was stuck BELOW the block layer). ⚠ The hypervisor was IDLE — load 0.63, io pressure 0.00, zpool ~0 writes: nothing was reaching the disk because the filter held it. Cause: vzdump of VM 102 → pbs-ana did 1% at 64 MiB/s then collapsed to 1.4 MiB/s for 35 min; Proxmox backups interpose a copy-before-write filter, so every guest write queues behind the backup's copy-out. FIX = cancel the task (pvesh delete /nodes/localhost/tasks/<UPID>); filter detached, inflight 45→0, D-states gone, 191 MB/s dsync restored. ⚠ fleecing 0 on the job is why a slow TARGET can stall a GUEST — fleecing routes copy-before-write to a fast local image instead. Job = backup-5d8f1221-8f71, daily 21:00, all 1, storage pbs-ana → recurs nightly until changed. A prior run of this VM managed 941 MiB/s read, so 1.4 MiB/s is degradation, not normal. → docs/runbooks/nh3-dev-io-stall.md

  • [2026-09-02] althing deploy is SIX surfaces, and #6 is outside the althing repo: ~/.claude/settings.json crossSessionInbound: "accept". Without it Claude Code HOLDS every cc poke — it auto-delivers only when the sender's permission-mode class matches, and the herald is a daemon that asserts none, so the notice goes to a human watching the pane instead of to the session. ⚠ The seat reports declared, reachable and green throughout — same failure shape as the SessionStart hook that was never deployed. Set on nh3-dev by forseti 09:28 with operator authorization (diff verified: one key, backup at /tmp/settings.json.bak-20260902T092829). Operator's reasoning: the herald reaches only local seats and a pane poke already types+Enters into a session, so the socket channel is strictly NARROWER than what it replaces — stating the existing trust boundary, not widening it. Cost without it is first-contact-only (in-memory correspondent record), not per-message. ⚠ No attestation exists for the herald to send — CC identifies a sender by verified pid against the session registry and reads that session's LIVE runtime mode; a daemon is not in it, and from_mode on a type:"user" frame is never consulted. deploy-althing.sh reports surface 6 and deliberately never SETS it — a deploy script that edits its own trust settings grants itself trust. → docs/runbooks/althing-deploy.md

  • [2026-09-02] vastblue gitea org created (id 8, private, owner vh) with empty repo vastblue/platform — third entity namespace alongside corviduo and pfi; most repos still live under vh/. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). Org scope was the decision: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide ana-docker-runner already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ Dedicated runner is gated on the first client-premises release cut, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as vh. → stacks/gitea-runner/README.md

  • [2026-09-02] althing 3.3.0 deployed — the cc channel, and a plugin-cache false green. CC seats are now poked over their own message socket ($XDG_RUNTIME_DIR/cc-socks/<pid>.sock) instead of by typing into the pane: no process to reap, nothing near the input line. infra-ops moved to channel=cc; the dwarves stay on pane and their guard-4 exposure is UNCHANGED (declare prefers cc, falls back). ⚠ An undocumented Claude Code interface, taken deliberately (operator: the FIFO poker was also an unsanctioned hack — a better instance of a class we already had). Break mode = seat goes pull-only with a logged reason, mail still held. ⚠ claude plugin update matches on the plugin VERSION and declines a content-only change — 3.3.0 edited plugin content at an unchanged 0.1.1, so the CC cache stayed stale while every version check reported success (delta was docs-only, harmless this time). deploy-althing.sh now diffs marketplace vs live cache. ⚠ Ordering: herald restart BEFORE anything declares cc, or the seat goes silently pull-only. ⚠ This box was at 3.2.4, not 3.2.5 — rollback target here is 3.2.4. Follow-on 3.3.1: the statusline bell measured a MECHANISM, not the property — it read wake-listener-<handle>.lock, so a cc seat renders 🔕 while push/reachable. Both copies now ask the post office (reachable from the status payload) and add 📵 for an outage. ⚠ TWO COPIES of that script now existscripts/claude-statusline-command.sh here (the operator's wired one) and althing's plugin/scripts/statusline.sh — independently fixed to the same shape; a drift surface with a countdown, convergence not yet raised with the operator. → docs/runbooks/althing-deploy.md

  • [2026-09-02] Every CI job on the shared pfi-fleet runner is root on ana-docker — and container.valid_volumes: [] does NOT prevent it. Measured: a job container is uid 0, /var/run/docker.sock is mounted by act_runner independently of that list, docker ps returns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome), docker compose v2.33.0 on PATH. ⚠ LOAD-BEARINGvh/Worldtree, vh/soong-lab, vh/skaldsong, vh/wt-matrix-bridge all drive buildx through that socket, so it cannot simply be closed; isolate sensitive builds onto a dedicated runner instead. Also measured the same night: services: containers work (Postgres 16), and full-URL uses: https://gitea.phasefinal.com/actions/checkout@v4 resolves from the local mirrors — the un-parked half of the github-independence work, needing neither DEFAULT_ACTIONS_URL=self nor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. → stacks/gitea-runner/README.md

  • [2026-09-02] pfi-gx10 BASELINED: 79.36 s/it median on the run-3c shape, and the training stack works on aarch64/sm_121. Median across 10 timed steps, 0.19% spread, peak 75.1 / 121.6 GiB — 46 GiB spare, attn_resolved: flex_attention. 6× slower than ana-ml2 where compute predicts 2.7× → likely memory-bandwidth-bound; capacity box, not throughput box. Ruled bare metal, not Proxmox (no aarch64 PVE; the GPU is on-package and cache-coherent, so passthrough would partition the unified memory that is the whole point). ⚠ sm_121 is NOT in torch's arch list — everything JITs from sm_120 PTX, so warm up before timing anything (an unwarmed bench read 27 TFLOP/s against a true 93). → persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md

  • [2026-09-02] I priced a failure in the units I happened to be measuring — operator overruled me, correctly. Recommended run 3c to ana-ml2 by costing a breaker trip as "≤50 steps ≈ 11 min of recompute". It is a 40-minute drive each way with 13 Anaheim hosts dark, three of them SureFire CLIENT machines. save_steps caps the recompute, never the outage. ⚠ General form: a metric in hand will volunteer itself as the unit of risk.persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md

  • [2026-09-02] althing 3.2.0→3.2.4 deployed, and ALTHING DEPLOY IS FOUR SURFACES not three. The fourth (plugin) had no runbook step and was frozen at Aug 28 — missing the SessionStart/SessionEnd hooks and pane-route.sh entirely, so "CC seats re-declare automatically" was never true here. Now one command (scripts/deploy-althing.sh). ⚠ uv tool install . without --force is a silent no-op. ⚠ A missing deploy surface presents as "the migration needs manual work", not as an error.persistent-memory.d/2026-09-01-althing-320-deploy.md

  • [2026-09-01] irv-ml1 GPU resident map, and dots-tts holds 14,430 MiB against a ~6 GB baseline — tts-dev's prompt-feature cache, capped at 32 entries after two incidents; the cap still permits a long way of growth. 3090 at 76% behind a warn-only watchdog. ⚠ Restates the GPU-ordering foot-gun: device_ids: ["1"] is the A6000 in a container, but a bare native CUDA_VISIBLE_DEVICES=1 gets the 3090. → persistent-memory.d/2026-09-01-irv-ml1-gpu-residents.md

  • [2026-09-01] The Ada inference server is a used Dell R750xa (JPJ1ZP3) and the reseller stripped four things Dell shipped — half the RAM, the 2400 W PSUs, and the GPU risers/cables/fans are absent from the invoice. Card is RTX 6000 Ada, not L40S. GPU power chain resolved via NVIDIA 930-00030-1546-000. NVMe in the drive bays is CLOSED (SAS/SATA backplane). → persistent-memory.d/2026-09-01-ada-inference-server-r750xa.md

  • [2026-09-01] pfi-gx10 onboarded headless — and it is the intended new home for run 3c, which died on a tripped breaker. GB10/sm_121/aarch64, 121 GB unified. NOT racked yet. Bare of any CUDA stack; probe throughput before porting. → persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md

  • [2026-09-01] Ada migration is zfs send (branch a) — and the DESTINATION IS SMALLER THAN THE SOURCE. 99 MB/s measured; ~3.9 h. ⚠ Measured 2026-09-01: storetank = 1.81 TiB pool, 1.45 TiB used, 80% CAP already, compression off / compressratio 1.00x (safetensors are incompressible — no win at recv). Settled payload ~1.47 TiB; the R750xa's as-bought 2× 1.92 TB mirrored is ~1.75 TiB → arrival at ~84%. Fix = 2× 2 TB SATA SSD on the buy list (6 bays free) → ~3.57 TiB at ~41% with redundancy; pair the two NEW drives together (a mirror vdev caps at its smallest member). ⚠ Pruning is NOT a substitute — comfy-dev found ~215 GiB unreferenced, and deleting every byte still lands the as-bought mirror at 72%: the constraint is vdev layout, not payload, so the prune audit and the drive purchase are independent and neither gates the cutover. ⚠ "Onboarded" is not "landed" — infra-ops read ALLOC mid-pull and re-added the whole batch on top, inflating 84% to a quoted 90%. Also: branch (b)'s original reason was WRONG — comfy-dev enumerated all 12 containers, only comfyui mounts /storetank, so (b) was unavailable during the transition, not structurally (right conclusion, wrong reason — infra-ops reasoned about the BOX when the question was the MOUNT). Plus the retain-vs-reclaim call and the two-boxes confusion (the Ada box and the GX10 are DIFFERENT machines). → persistent-memory.d/2026-09-01-ada-migration-branch-a.md

  • [2026-09-01] Matrix: Synapse 1.120→1.159, appservice namespace opened, /_synapse/admin closed to the internet, alias convention ratified. Schema migrations are one-way; push is event_id_only and assembled on-device. → persistent-memory.d/2026-09-01-matrix-upgrade-and-hardening.md

  • [2026-09-01] A named failure class: a correct check aimed at the wrong object. Six instances in one day across three sessions; re-running the same check cannot catch it. Recommended for docs/pfi/training-throughput-playbook.md §4 — NOT YET WRITTEN, awaiting operator.persistent-memory.d/2026-09-01-wrong-object-measurement.md

  • [2026-09-01] Ops boundary ruled by the operator: worldtree-dev writes the bridge code; infra-ops OPERATES the Worldtree/Matrix instances and may change them. Corrects a mis-route where infra-ops asked worldtree-dev to provision an account on a box it does not run. Tracked at 931bac8 + althing 01M1F4PK796EDGDCBKZ9W3JC0S.

  • [2026-09-01] Idle VRAM on this fleet is a RESERVED scratch pool, not waste. Operator declined raising vllm-mog-sec from gpu-memory-utilization 0.52: single-user dev fleet, KV headroom nobody will consume is worth less than room for ephemeral models and small training runs. vLLM's "fully utilize gpu memory" startup hint does NOT apply here. Tracked in auto-memory feedback_idle_vram_is_reserved_not_waste.

  • [2026-08-28] althing v3 flag day (U9b) executed, then six releases to 3.1.1 in one afternoon — and the post office MOVED to nh3-docker. Every v2 command deleted; 73 handles seeded and verified by set difference; 5,043 orphaned wake FIFOs deleted (v2 named them per-session+PID, v3 per-handle). Image now registry-pulled, digest-pinned, under the claude-bot namespace. → persistent-memory.d/2026-08-28-althing-v3-cutover.md

  • [2026-08-28] A stale ALTHING_HANDLE silently reads another agent's inbox and reports it empty — a SECOND route into the failure v3 exists to prevent. Outbound mis-signing sometimes gets caught; inbound never does. Shipped as a 3.1.1 warning. ⚠ My session_handles.json grounding was wrong (v2 artifact, v3 never opens it) and the same stale source had survived inside my statusline rewrite. → persistent-memory.d/2026-08-28-handle-resolution-wrong-inbox.md

  • [2026-08-28] nh3-dev's three OOM events attribute to CLAUDE CODE, and the "no kernel evidence" was a permissions artifact. journald was persistent all along; journalctl silently shows only your own messages outside adm. Single CC sessions measured 5.4-18.4 GB, so 27 GB is 3-4 long-lived sessions. sysstat + atop now instrument the ramp. → persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md

  • [2026-08-28] sec moved to ana-ml2 GPU0 and is serving (operator-directed) — GPU1 had ~28 GB free against the ~51 GB it reserves, so it could not start there. Re-arms the two-GPU load condition on a circuit that tripped 36h earlier; accepted with the constraint stated. → persistent-memory.d/2026-08-28-sec-seat-gpu0.md

  • [2026-08-28] BELAYED by the operator, both explicitly: (a) a cgroup memory cap on CC sessions, (b) putting ana-gw + ana-wg + one BMC on separate power. Both were my recommendations; neither is open work. Do not re-raise as new — the atop ramps that would inform (a) are now being collected, so revisit only with a week of data. Tracking surface: this entry.

  • [2026-08-28] The deployed CC plugin copies are a release step nobody owns. sync_skill.sh covers the SKILL, not the plugin; both copies must be rsync'd from the repo's plugin/ on every althing release or they carry the previous release's bugs into the live surface. Raised with forseti for their release notes. Tracking surface: althing thread 01M14QHZNDKDK8KH9DN92VF6VE.

  • [2026-08-28] althing v3.0.0 flag day (U9b) executed — the post office replaced the P2P bus on both boxes, one-way. 73 handles seeded and verified by set difference; 5,043 orphaned v2 wake FIFOs deleted (v2 named them per-session+PID and never reaped; v3 names them per-handle, so the leak is bounded by construction); v2 db left inert. → persistent-memory.d/2026-08-28-althing-v3-cutover.md

  • [2026-08-28] nh3-dev's three OOM events attribute to CLAUDE CODE — and the "no kernel evidence" was a permissions artifact. journald was persistent all along; journalctl silently shows only your own messages outside adm. Single CC sessions measured at 5.4-18.4 GB, so 27 GB is 3-4 mature sessions, not the ~66 a 408 MB estimate implies. sysstat + atop now instrument the ramp. → persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md

  • [2026-08-25] Run 2's base is an OPEN OPERATOR DECISION, deliberately not staged — four options with materially different safety postures, detailed in Current state. Tracked at althing thread 01M0WQ8W5574KMEVCHCEKEXNS5. ⚠ Do not let it get filed as a config knob; it is a reversal of the trainee-selection decision.

  • [2026-08-25] Fused MoE kernel path — DEFERRED, tracked at park fused-moe-kernel-path-for-gemma-4-moe-training (id 47). Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is 8.6% (27.1 of a benchmarked 313.8 TFLOPS) because transformers runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: group_by_length (29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). Not applied to the live run — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312.

  • [2026-08-24] Serving the tuned ERP model: LoRA-on-NVFP4 PREFERRED, merged weights the expected fallback — and the recorded objection may be STALE. Operator: "if you CAN load it as a lora, all the better, the issue is that we will want to run nvfp4 weights, which we had some serious trouble with loading loras on top of nvfp4." ⚠ The archived root-cause says it was NOT NVFP4-specific: [2026-07-07] vLLM 0.24.0 qwen3_5 LoRA application was a silent no-op (#47639, regression from #37912) — adapter loads HTTP 200, zero deltas at inference, proven quant-agnostic (NVFP4 AND FP8 both inert) and adapter-format-agnostic by a 3-peer dwarf panel. Fix PR #47640 was OPEN then. ana-ml2 is FAR past 0.24.0 and the box runs a SPREAD, not one version (measured 2026-08-24): gen on nightly-311b3513 = 0.27.2rc1.dev150, mog-sec on nightly-e9d1398d = 0.26.1rc1.dev1102, the small seats still on 0.24.0, and char-rp/trainee-bench pinned to v0.26.0. ⚠ vllm/vllm-openai:v0.27.1 is already ON DISK, unused — a TAGGED release, which is the right retest target: no nightly variance, no pull, ~4 months past the diagnosis. So: RETEST hot-swap LoRA on v0.27.1 before designing around merge — it is cheap, and if it works the post-tune gate can be two aliases on one engine. If it still no-ops, merged weights it is, which means the harness must EMIT merged weights and Eitri needs that in the contract while he is early. Tracked at this snapshot commit; settle it in the QLoRA sizing conversation.

  • [2026-08-24] nconnect=8 on /mnt/smithy — approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread 01M0R46SFYF83099N16WD67KGD.

  • [2026-08-19] AI-tab Dormant regrouping BELAYED by the operator — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than AI - Dormant. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. untracked by operator choice (his words: "belay the ai dormant regrouping for now").

  • [2026-08-15] RP-seat direction: KEEP MeroMero on char-rp; Artemis-31B rejected; next move is Dark-Scarlett on a Qwen3.8 base when it lands (operator). Evaluated TheDrummer/Artemis-31B-v1.1 — mechanically a drop-in (same google/gemma-4-31B-it base, identical 1188-tensor/356-vision census, same missing-preprocessor_config.json trick), so it's purely a quality call, and our own survey already ranked MeroMero #1 vs Artemis #6; Artemis is also unlicensed and its author deprioritizes correctness + warns of token-banning-for-stability, which fights char-rp's tool-calling requirement. MTP verified impossible on both (Gemma-4 has no MTP head at all — base/MeroMero/Artemis are all MTP=0; no finetune can add one). But speculative decoding IS reachable on a Gemma-4 seat via a DETACHED drafter — vLLM 0.24 supports eagle3 + gemma4_mtp, and real drafters exist: google/gemma-4-31B-it-assistant (0.94 GB, 4-layer, 761K dl), RedHatAI/gemma-4-31B-it-speculator.eagle3 (4.47 GB), AEON-7/…eagle3-NVFP4 (3.53 GB). ⚠ all list their verifier as stock gemma-4-31B-it, not an RP finetune, so acceptance against MeroMero is unmeasured and likely well below the gen seat's ~48%. UNTESTED — parked, ~45 min to measure, needs GPU0 headroom (card is at 94.4/97.9 GB). Why the Dark-Scarlett 3.8 plan is the strong one: DS is Qwen3.6-based today, so a 3.8 respin lands on the gen seat's architecture → native MTP returns and the whole mixed NVFP4+FP8 recipe + graft ports directly. Watch two things on arrival: from_pretrained silently drops MTP heads during finetuning (verify 15 mtp.* tensors in the index; graft from stock if absent), and DS v1.0 required the Qwen3_5ForConditionalGeneration wrapper class to save a config vLLM/SGLang accept. Both in docs/pfi/model-quantization-playbook.md.

20 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-09-04] Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes. ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. → persistent-memory.d/2026-09-04-dac-forced-10g-failed.md

110 older entries archived to archival-memory.md.