Files
esh-pfi-infrastructure/persistent-memory.md
T

66 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.

(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the handoff unread across any overnight gap, which is the exact case it exists for. 8 h also matches the global CLAUDE.md and the /snapshot skill default.)

Repo purpose

  • 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up. /tank and other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. See persistent-memory.d/2026-09-10-beszel-fleet-wiring.md and stacks/beszel/README.md.

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) — [2026-09-24] MOTHBALLED by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (e6da607) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecated — althing-cli→postbox, althing-wake-listener→althing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild)
vh/soong-lab Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16)
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path) — [2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev (fleet ownership audit; also CLAUDE.md 2026-09-05). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-10-01 ~0420 PT.

Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)

  • INCIDENT 04:21 PT 2026-10-01: FIXED by nemo-0.1.1 (infra-hermes 963f9ed, live ~0441, infra-ops spot audit passed).
    • Cause: gen-small hit CUDA OOM on a 394 MiB lazily allocated workspace while the seat was parked at its 3,582 MiB window cache (GPU 0 Free 388).
    • The fix: empty_cache around every window; MEM_CAP_MIB=3840 (a hard process ceiling); CUDA_GRAPHS=0, because empty_cache poisons the graph pool (an illegal memory access on the next request).
    • Result: the seat rests at 2,108 MiB and peaks at 3,028 during a 12-min file. gen-small is steady at 36,116 (its runtime growth landed at init, with zero movement on later requests). GPU 0 Free is ~1.05 GB at rest.
    • Cost: graphs-off is ~2–8 ms slower at short clips and ~27 ms at 20–60 s (35 / 40 / 50 / 98 ms against 33 / 36 / 42 / 71), still 4–15× faster than the old seat.
    • Lesson: a shared-card tenant must carry a HARD cap, whether or not it returns memory.
  • Rollback: docker stop parakeet-nemo && docker start parakeet. The old container and image are kept.
  • GPU 0 is FULL:
    • The seat's steady state is 3,582 MiB (its cached window peak); Free is 385 MiB.
    • vllm-gen-small runs at util 0.36, and its .env holds 0.33 for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×.
    • Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01.
  • Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with --http h11 (httptools 0.8.0 emits HTTP/1.1 200\x00OK, which LiteLLM/httpx rejects).
  • NVIDIA Open Model License accepted for internal use. → persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md

Worldtree U11 memory cutover (demo + personal)

  • Legacy plane OFF since 0115/0120 PT 2026-09-30 (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md
  • Daily gate batches: infra-hermes runs scripts/wt-memory-gate-batch from 2026-10-01 and copies me on every verdict. The count is 1 of 3 consecutive PASS at off (20260930T090608Z); a FAIL restarts it.
  • ⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:
    1. Run docker exec -i <api> python - < scripts/wt-h2-count.py VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output.
    2. Delete LIVE, with the api running, using literal paths only: agents/{forseti,lofn,mimir}/memory/<agent>.chroma and memory/context_promotion, on BOTH instances.
    3. Send worldtree-dev the stamp; b193 ships after it.
    • A b192 restart re-creates an empty schema-only ledger.db. That is residue: say so in the stamp and remove it after b193.
  • Legacy archive: DESTROY it whole by 2026-10-30, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file.
  • TODO: re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys.

fv-ml1 GPU layout (as of 2026-10-01)

  • GPU 0: cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
  • GPU 1: vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
  • GPU 3: the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to irv-ml1's A6000, not back to GPU 1.

intern-decision (replaced SemIf on 2026-09-30)

  • LIVE 0.1.3 at intern-decision.fv.internal:8033: semif-compatible /decide plus Jev /v1/systemone, 32k tokens, a Triton cache volume. Run scripts/intern-decision-warmup after an IMAGE change. → persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md
  • Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.

Scriberr (fv-ml1 GPU 3)

  • LIVE scriberr:local-blackwell-a353078-dropout2: upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, PARAKEET_MODEL_PATH), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. scripts/scriberr-rebuild re-applies both. → persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md

nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)

  • nh3-pve: Secure Boot OFF; IGFX restored; NVIDIA 580.178.04 DKMS; kernel 6.8.12-43 (benign btmtk oops every boot). AMT live: static 10.100.250.61 on nh3-mgmt (UDM port 6), KVM on, Opt-in None, password in the vault as nh3-pve/amt-admin. → servers/nh3-pve/README.md
  • nh3-ml1 (CT 109 @ 10.100.50.80):
    • TEI embed/rerank is load-shared with esh-ml1 through the gateway, with router_settings.enable_weighted_failover.
    • brokkr's foundry seats:
      • lfm-vl :8030 (gateway lfm25-vl-3b);
      • lfm-vl-uncensored :8032 (direct; passed brokkr's eval);
      • vibevoice-asr :8031 (audio.cpp, Q8_0).
    • GPU ~11.3/16 GB. → servers/nh3-ml1/README.md
  • Not yet proven: nh3-ml1 coming back by itself after an nh3-pve reboot (only the config was checked).

esh-matter: Matter server for Home Assistant (2026-09-26)

  • CT 111 @ 10.0.90.20, on VLAN 90 (esh-iot) only. matter.js 1.4.0; :5580 is firewalled to HA 10.0.50.46. HA's matter integration is loaded (ha-dev).
  • Next is Prime's: share the Aqara W200 into HA via Matter multi-admin. If commissioning misbehaves, the UniFi mDNS reflector on esh-iot is the first knob (a Prime/infra-ops change). → servers/esh-matter/README.md

Worldtree reward path: code + config live (2026-09-27, Prime)

  • d8726a13 (Domari's Skywork/Selene fix) was ALREADY on origin/main when Prime said to push it: it was pushed as part of 79cce92f at 1221. Demo runs image 79cce92f. Personal still runs 51168a2e, which is older and lacks the fix; its image follows worldtree CI, never a manual up.
  • deploy-wt-config deployed demo and personal (1437): providers.yaml reward_models.scalar-judge (6d4ac44), plus a77639d's corrected selene description, which had never reached the hosts. Both are healthy, with 0 drift. It is inert on personal until that image picks up d8726a13.
  • worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were pushed on Prime's go-ahead (a9d091e..6d4ac44). Origin, the repo and both hosts now agree.

Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime)

  • Prime: "go ahead with gpu 3, both", and he does not use Blender, so agents drive it through MCP. Blender 5.2.2 LTS (linuxserver/selkies image, digest-pinned) with a web desktop at https://10.251.50.54:3001 (vault fv-ml1/blender-web-password). On demand only: scripts/blender-mcp up|down|status. It was left DOWN (GPU 3 back to 2 MiB).
  • MCP: mcp-for-blender 2.1.1 (Dvalin's research pick). The add-on is vendored at 41a18432. The server runs INSIDE the container, and agents reach it as ssh+docker-exec stdio via scripts/blender-mcp. No port is published, and screenshots work because server and Blender share a filesystem. Telemetry is off and safe mode is on. Verified end to end (render, screenshot, safe-mode refusal). Batch path: scripts/blender-run (one-shot docker run --rm, --job staging; first user is draupnir). Access: the shared fleet infra-ops login, with no render-only key (Prime, 2026-09-28). Registration + up/down: PER WORKING SESSION (Prime, 2026-09-28, relayed by draupnir; superseded per-task of 09-27): blender-mcp up + claude mcp add at session start, claude mcp remove + down at its end. Never always-on, never user- or project-wide config. → stacks/blender/README.md
  • Extensions (2026-09-28, draupnir; Prime ruled Blender a MANDATORY pipeline stage): 8 pinned add-ons (stacks/blender/extensions.lock) built by scripts/blender-extensions sync into fv-ml1:/tank/blender-extensions/5.2/system (LIVE), mounted read-only as the System repo; fleet_extensions.py enables them (GUI startup timer; blender-run --extensions). Headless acceptance 8/9 (CAD Sketcher sketching is GUI-only); MCP 9/9 after Prime ran the deploy himself (1505; the classifier had refused mine). SurfacePsycho's eval() is patched to literal_eval (held under MCP). Open: agent-drawn CAD Sketcher geometry (its stateful ops want point picks), and MeasureIt overlays not seen in MCP screenshots. GUI left DOWN.

Zigbee2MQTT on esh-docker-vm (2026-09-27, Prime go-ahead; ha-dev request)

  • LIVE since 1240: Z2M 2.14.1 (digest-pinned) on http://10.0.50.45:8099 (auth token), radio SLZB-MR1U chip 0 tcp://10.0.90.10:6638 (ember), channel 25, PAN 0xCFF4. State and the network key are in /opt/docker/data/zigbee2mqtt (root 0700, restic via /opt/docker, never in git). Vault: esh-docker-vm/zigbee2mqtt-{network-key,frontend-token,mqtt-password}. Broker user zigbee2mqtt added to mosquitto (passwd backup .bak-20260927-z2m).
  • ha-dev deleted the ZHA entry; pairing the Aqara T1 is theirs. ⚠ The "bridge in HA" acceptance first FAILED because HA had had no MQTT since 2026-09-25 06:13. Cause: the esh-docker-vm macvlan shim had no host route to HA (two equal /24s, ens18 wins). Fixed at 1249 with a /32 via the shim (if-up.d hook, playbooks/esh-docker-vm-macvlan-shim-route.yaml), and HA reconnected. My first report called HA "connected" from a client count that included my own probe; $SYS counts are not identities.
  • ⚠ The first credential-mint attempt was blocked by the auto-mode classifier on a peer-relayed approval. It went ahead only on Prime's own go-ahead in this session. Treat that as correct. → stacks/zigbee2mqtt/README.md

restic: credential leak fixed (2026-09-27, Prime)

  • Seven hosts were moved from env-file to repository-file (playbooks/restic-repository-file.yaml), so the systemd units no longer carry the rest-server password. Secrets are vaulted as <host>/etc/restic/{repository,password}. restic.env is KEPT for manual snippets, so a rotation must update the vault, restic.env and repository.
  • vm-esh-nas done too (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). All 8 hosts are clean and vaulted.
  • The passwords were readable until today, so the rotation (Prime's) remains the real fix.
  • 2026-09-27 0859 freshness check: all restic and PBS repos ✅. Those 0100 runs PRE-date the move (done ~0140–0200).
  • VERIFIED 2026-09-28 (infra-hermes, read-only): the first nightly on repository-file (0100 PT) succeeded on all 8 hosts (unit Result=success AND a today-dated snapshot in each repo), and the 0800 freshness check was all-green, PBS included. nh3-dev also passed a content restore (af5580c3).

augaman: face recognition for Cicada

  • v0.1.3 LIVE on esh-ml1:8040 (v0.1.1 first deployed 2026-09-26 2347 PT), stacks/augaman, healthy on CUDA, in nvidia-smi. Built on-box from a git archive of the tag. pytest -m gpu tests/vision 3/3 PASS on v0.1.3.
  • Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27 (Prime OK'd off-site): restic daily 0100 PT → rest-server-ana esh-ml1/, fail-closed hook, secrets vaulted under esh-ml1/etc/restic/. Identity-level restore matched the live gallery (canary, snapshot fd3061a1). The backup gate for real enrollments is MET.
  • Deploy CLOSED by augaman-dev 2026-09-27 0009 PT: HTTP checks passed; the canary survived the recreate and was then deleted, so the gallery is empty and ready for real enrollments. /recognize p50 186 ms (1080p, one face). The detector-latency follow-up is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops.
  • The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT) after the bench. esh-ml1 is the only instance: in the house, backed up, 48 ms per face on v0.1.3.
  • v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later. Speed bench (docs/pfi/augaman-speed-bench/), server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to augaman-dev; the deployment does not use CPU mode.

esh-ml1

  • The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the gateway /scalar-judge passthrough, which is key-gated per key via allowed_passthrough_routes, granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image (see "Worldtree reward path").
  • The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or drop.
  • The Beszel superuser password was echoed into a session transcript (local only, not in memory). Rotation offered to Prime.

Live threads

  • git: eshpfi-management pushed to 6b66207 and worldtree-instance-configs to b6fdd81 (2026-10-01 ~0418, Prime's go). Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with git commit -- <paths> (auto-memory feedback_shared_git_index_commit_pathspecs). graphify-out/GRAPH_REPORT.md stays modified and uncommitted on purpose: it is auto-regenerated.
  • nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
  • Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth 50bfc7b). Pushing it is booth's call, per Prime; it is not ours.
  • ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.

Recent decisions

  • [2026-10-01] ⚠ Gitea's [webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8 silently REJECTS headscale mesh IPs (100.64.0.0/10). The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN 10.6.110.50:9009; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without secret and branch_filter drops both, so always resend them.

  • [2026-10-01] vh/arbo push webhook repointed from the retired wg0 lifeline 10.100.79.3:9009 to irv-ml1's mesh address 100.64.0.6:9009. It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner irv-ml1-arbo was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.

  • [2026-10-01] Prime: delete the bench leftovers, no upstream for Scriberr, push. DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; unified-en KEPT, the live seat mounts it), plus /tank/spikes/scriberr-slicer (including the private copies of Prime's recordings) and /tank/spikes/parakeet-ab. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.

  • [2026-10-01] irv-ml1 /storetank reclaim done: Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread 01M3TCSYRSFNAPA9BPQTMFQ6KJ).

  • [2026-09-30] Parakeet speech seat → parakeet-unified-en-0.6b under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137. → persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md

  • [2026-09-30] Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30. → persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md

  • [2026-09-30] SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev /v1/systemone at 32k tokens on GPU 1, with a Triton warm-up cache volume. → persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md

  • [2026-09-30] Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept. → persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md

  • [2026-09-29] Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches). → persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md

  • [2026-09-28] Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors. → persistent-memory.d/2026-09-28-worldtree-u10-backfill.md

  • [2026-09-28] Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired. → persistent-memory.d/2026-09-28-bonsai-ternary-spike.md

  • [2026-09-28] blender-run gained --cpu (no GPU attached) and a fixed hostname fv-ml1-blender (draupnir). The design stage never renders, so it stays off GPU 3.

  • [2026-09-28] Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (--extensions). SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → stacks/blender/README.md § Extensions

  • [2026-09-27] hermes-gateway restarted 0401 for highseat-dev (SVOS v2.1.12: propose_decision gained seat_up, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling zellij-fleet@Claude at boot remains Prime's call.

  • [2026-09-27] SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime): wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md

  • [2026-09-27] Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port. → stacks/blender/README.md

  • [2026-09-27] esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August). → servers/esh-docker-vm/README.md

  • [2026-09-27] Zigbee2MQTT live on esh-docker-vm :8099 (PAN 0xCFF4, ch 25) replacing ZHA; network key vaulted, data host-only. Prime go-ahead in this session; a peer-relayed approval was blocked by the permission gate. → stacks/zigbee2mqtt/README.md

  • [2026-09-27] semif-serve 0.1.4: object states ending in ), ; or } no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart. Prime ruled; heid bug hunt folded. → stacks/semif/README.md

  • [2026-09-27] SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win. Build nothing (Prime). Henge 88 carries it. → persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md

  • [2026-09-27] SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21. Recommendations await Prime. → persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md

  • [2026-09-27] SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels. Tracked in the in-flight SemIf section. → persistent-memory.d/2026-09-27-semif-order-averaging.md — DONE: 0.1.3 live, with fast kernels adopted (77b8cb4).

  • [2026-09-27] restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime). Rotation stays Prime's. → persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md

  • [2026-09-27] infra-ops bootstrapped on vm-esh-nas by Prime (the first account at the fleet-pinned uid/gid 850; bootstrap-infra-ops-user.yaml now pins 850 when free, d775a01).

  • [2026-09-27] augaman: second instance on fv-ml1 benched, then REMOVED (Prime). esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (ad484c3, bench docs/pfi/augaman-speed-bench/).

  • [2026-09-27] Created empty private repo corviduo/svoperatingsystem (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. svos-dev → the OS session; The High Seat's session is highseat-dev again.

  • [2026-09-26] augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy. → persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md — [2026-09-27] DONE: deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (d8f59a1).

  • [2026-09-26] zellij-fleet@.service installed, NOT enabled (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling @Claude at boot is Prime's call when seat_up ships. Tracked: 0ad7799, services/zellij-fleet/README.md.

  • [2026-09-26] btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority). Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked: servers/nh3-pve/README.md.

  • [2026-09-26] GPU-LXC Temperature alerts watch the GPU, not the host CPU (SENSORS=-coretemp_*,acpitz); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump. 6d901c0.

  • [2026-09-26] Abliterated LFM2.5-VL-3B parallel seat (:8032) passed brokkr's 6-image eval and stays (direct only; stock seat kept for SFW A/B). 9cc3824, 944bb36.

  • [2026-09-25] AMT follow-ups PARKED until the two MS-03s arrive and esh-pve's MS-01 is on AMT (Prime 2326). — Parked: MeshCommander test; dummy HDMI plug, then NanoKVM → gx10; an OOB path not through nh3-scale… → persistent-memory.d/2026-09-25-amt-follow-ups-parked-until-the-two-ms-03s-arrive-and-esh.md

  • [2026-09-25] nh3-pve AMT LIVE: static 10.100.250.61 on nh3-mgmt (UDM port 6), KVM on, Opt-in None — (Prime). → persistent-memory.d/2026-09-25-nh3-pve-amt-live-static-10-100-250-61-on-nh3-mgmt-udm-port.md

  • [2026-09-25] pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot. servers/pfi-gx10/README.md, servers/esh-docker-vm/README.md.

  • [2026-09-26] VibeVoice ASR → Q8_0 (Prime). WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed.

  • [2026-09-26] esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only — , for ha-dev (operator-approved, relayed). → persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md

  • [2026-09-26] Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime). — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… → persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md

  • [2026-09-26] Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed): — LFM2.5-VL-3B on llama.cpp :8030 (gateway lfm25-vl-3b, LiteLLM restarted 36 s at 0039) and… → persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md

  • [2026-09-26] Coder seat STAYS on fv-ml1 (Prime). — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… → persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md

  • [2026-09-25] nh3-ml1 LIVE after the NH3 visit (SB off, IGFX restored): driver + CT 109 + TEI, parity vs esh-ml1 indistinguishable (cos min 0.999993 = own floor; overlap@10 1.000 vs MRL-256 control 0.684), same speed; monitoring/DNS wired. Autonomous: lxc-pve 6.0.0-2 upgrade on nh3-pve (Docker-in-LXC fix #7006). Gateway routing + AMT follow-up are Prime's. → persistent-memory.d/2026-09-25-nh3-ml1-live.md

  • [2026-09-25] nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls. Tracked: 7ddd116, Current state. → persistent-memory.d/2026-09-25-nh3-ml1-standup.md

  • [2026-09-25] nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333). Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked: servers/nh3-pve/README.md (2b49be4, cdd7605).

  • [2026-09-25] MS-01 foot-gun: a GPU in the PCIe slot renames every NIC — (the slot's root port takes bus 01, so the X710 goes enp2s0f0np0→enp3s0f0np0). → persistent-memory.d/2026-09-25-ms-01-foot-gun-a-gpu-in-the-pcie-slot-renames-every-nic.md

  • [2026-09-25] Direct ESH→esh-ml1 consumer path — PARKED (no ESH-side callers in 7 days; consumers go through the gateway). Tracked at persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md.

  • [2026-09-25] Created empty private repo corviduo/norn (git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; the norn-dev handle is the operator's to declare.

  • [2026-09-25] Reward seat audited (Skywork-Reward-V2-Llama-3.1-8B still #1 of 188 on RewardBench 2; our AWQ ≈ bf16 within noise; double BOS costs ~2.7 pts) and moved fv-ml1 → esh-ml1; esh-ml1 monitoring wired (Beszel+GPU, Kuma, Homepage, Dozzle); Dozzle hub's stale agent IPs fixed; nh3-dev Beszel agent revived. → persistent-memory.d/2026-09-25-reward-move-and-esh-ml1-monitoring.md

  • [2026-09-25] TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend for qwen3-embedding + reranker; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1. → persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md

  • [2026-09-24] esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM order: 2 failover; dead DB alias reranker-a3-bge-v2-m3 (pre-relocation IP) repaired. → persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md

  • [2026-09-24] esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank (implementation deferred to next session, tracked here + servers/esh-pve/README.md). → persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md

  • [2026-09-24] pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25 (tracking: servers/pfi-gx10/README.md, e5197a3). — [2026-09-25] ✅ VALIDATED: Prime pulled and replugged AC, and it came up by itself at 1108:54 (2b49be4). → persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md

  • [2026-09-24] NH3 power outage recovered — pbs-nh3 had no onboot (set), NFS boot race fixed with automount (1cbde50), every other Claude session on nh3-dev died. → persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md

  • [2026-09-24] Miranda standing order is a repo CLAUDE.md operating parameter — (4b29492, aligned to the global send protocol in bcf3342): high-urgency matters go to her, fixed or not… → persistent-memory.d/2026-09-24-miranda-standing-order-is-a-repo-claude-md-operating.md

  • [2026-09-24] task-board mothballed (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (e6da607). Its hooks had sent no traffic in 30 days.

  • [2026-09-24] Military 24-hour Pacific clock times carried into the Codex/Grok shared bootstrap docs/fleettools/AGENT-BOOTSTRAP.md (ad2b4d9); Claude seats get it from the global CLAUDE.md.

  • [2026-09-24] Worldtree admin.memory.forget stays OFF on demo/personal until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config.

  • [2026-09-24] A git checkout under root:docker needs safe.directory for its deploy user — the 09-14 normalization (826a63b) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (eea9eb2). Sweep found no other case.

  • [2026-09-23] elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts. → persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md

  • [2026-09-23] headscale-ddns hardened — (fedd4b6): Cloudflare calls retry and validate every body (pick()), no write without both IDs, the run… → persistent-memory.d/2026-09-23-headscale-ddns-hardened.md

  • [2026-09-23] esh-docker-vm restic was skipped 09-22..23 by my own Kuma move — a dead uptime-kuma lookup aborted pre-backup.sh under set -e (25e41d2). → persistent-memory.d/2026-09-23-esh-docker-vm-restic-was-skipped-09-22-23-by-my-own-kuma.md

  • [2026-09-23] hermes-gateway restart exit-1 is a Hermes race, not a crash — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in… → persistent-memory.d/2026-09-23-hermes-gateway-restart-exit-1-is-a-hermes-race-not-a-crash.md

  • [2026-09-23] Booth link board cleared to 14 durable links (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted at scriberr.fv.internal:8080). Prime's rule: durable debug links only.

  • [2026-09-22] Both carried calls approved — build the NRestarts flap sampler (163bb97); the restic content-assertion ruling is ratified and stays (ba60fda).

  • [2026-09-22] safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes. — ⚠ The package installs INERT and looks fine — Debian's /etc/zsh/zprofile has 0 non-comment lines so the… → persistent-memory.d/2026-09-22-safe-rm-installed-on-nh3-dev-and-delegated-fleet-wide-to.md

  • [2026-09-22] The acceptance probe for a guard must not be able to destroy what it tests — (infra-hermes). → persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md

  • [2026-09-22] ⚠ {"sent": true} is a claim about transmission, never about effect. pane_send structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable /clear and its receipt reported success. Consumed an operator approval. → persistent-memory.d/2026-09-22-instrument-errors.md

  • [2026-09-22] D-0010/D-0011 were misrouted to this seat by pane_find matching a ROLLING PANE TITLE. — Genuine and operator-approved, wrong seat; fleet_telemetry held the right mapping and carries the warning… → persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md

  • [2026-09-22] headscale-ddns exited 1 silently and the alarm carried no cause — both failure paths were || exit 1 with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. 30517fd.

  • [2026-09-22] ⭐⭐ Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject. --state=running skipped the units most needing hooks; awk '{print $1}' dropped systemd's ●-decorated FAILED rows; grep -ic restic on the wrapper missed the check script; restic ls's header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. Every one reported cleanly while looking at the wrong thing. The rule is not "verify" — it is verify, then ask what the verification could not have seen. → persistent-memory.d/2026-09-22-instrument-errors.md

  • [2026-09-22] Fleet alert bridge generalized — beszel-althing → althing-alert-bridge, route registry (/beszel + /kuma), each with its own… → persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md

  • [2026-09-22] Uptime Kuma rebuilt from scratch on 2.5.5, moved esh-docker-vm → ana-docker. — ⚠ :latest is a TRAP — it tracks 1.x, so an Aug-2026 pull gave a Dec-2024 build. → persistent-memory.d/2026-09-22-uptime-kuma-rebuilt-from-scratch-on-2-5-5-moved-esh-docker.md

  • [2026-09-22] Beszel and Uptime Kuma are DISJOINT, not redundant — Beszel's alerts bind to a system with a threshold; there is no URL column, so it is structurally incapable… → persistent-memory.d/2026-09-22-beszel-and-uptime-kuma-are-disjoint-not-redundant.md

  • [2026-09-22] Failed-START alarms on 23 nh3-dev units — (services/althing-notify-failure/). → persistent-memory.d/2026-09-22-failed-start-alarms-on-23-nh3-dev-units.md

  • [2026-09-22] Backup coverage is a property of the SYSTEM, never one job's scope — establish it by querying the repo for the path in a real snapshot, never by reading a job's SRC=. → persistent-memory.d/2026-09-22-backup-coverage-is-a-property-of-the-system-never-one-job-s.md

  • [2026-09-22] restic checks now assert CONTENT and are DISCOVERED not enumerated — conjunctive (recent AND present AND restores non-zero bytes); repos discovered from each NAS, hand list… → persistent-memory.d/2026-09-22-restic-checks-now-assert-content-and-are-discovered-not.md

  • [2026-09-22] irv-ml1: Irvine is a TENANCY behind a Fortinet PFI does not control — its TLS inspection breaks Tailscale relay/control intermittently (41 cert warnings/week). → persistent-memory.d/2026-09-22-irv-ml1-irvine-is-a-tenancy-behind-a-fortinet-pfi-does-not.md

  • [2026-09-21] ⭐⭐⭐ lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone. 720 generations, 3 arms. Voice +0.152 at 2.9× floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and ~3/4 of the gain survives stripping every punctuation mark, so it is not the cheap win. ⭐ Memorisation: ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12, all 96 matches READ and every one stock grammar (he looked at the wolf and he looked at him); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: 20% / 28% of generations overshoot the 90–140 band against base's 1%, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribed score_beats.py's v1 criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. → persistent-memory.d/2026-09-21-lv-mccarthy-gate.md

  • [2026-09-21] The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. — On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one… → persistent-memory.d/2026-09-21-the-two-epoch-recipe-is-now-0-for-3-and-this-time-the-loss.md

  • [2026-09-21] The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass. → persistent-memory.d/2026-09-21-the-memorisation-control-that-shipped-broken-is-now.md

  • [2026-09-21] My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug. → persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md

  • [2026-09-21] ⭐⭐⭐ The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable. Claim dropped by a sub-tool, hook behind graphify's eight exit 0s, no handle in the env, ssh-target written as a hostname. The general shape is configured ≠ effective; twelve instruments reported confidently and wrongly across three days, five of them mine. → persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md

  • [2026-09-21] ⭐⭐ The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing — a reveal handler Jinja discarded for sitting after {% endblock %}, and a × a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. scripts/layout-probe.py took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → persistent-memory.d/2026-09-21-booth-two-dead-controls.md

  • [2026-09-21] ⭐⭐ claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to pfi; vh-token use is now standing-authorized from the vault. ⚠ vh is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ ~/.config/claude-bot/gitea-token is DEAD and had been misreporting permissions; the working one is gitea-token-repo-create. → persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md

  • [2026-09-21] ⭐⭐ nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from /tmp, which this box sweeps at 3 days. babyyarros existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md

  • [2026-09-21] Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested — (34c4179, e574b91, 2e08edc). → persistent-memory.d/2026-09-21-draupnir-geometry-engine-provisioned-on-irv-ml1-and.md

  • [2026-09-21] Backup alarm verdict split: STALE (exit 1) vs ERRORED-JOBS (exit 3) — (7fe4102). → persistent-memory.d/2026-09-21-backup-alarm-verdict-split-stale-exit-1-vs-errored-jobs.md

  • [2026-09-21] vh/forgefirm mirrored — from github.com/openglow-org/forgefirm, following the house convention read off the existing 17: vh/… → persistent-memory.d/2026-09-21-vh-forgefirm-mirrored.md

  • [2026-09-20] ravenpen.com REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered. — infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a… → persistent-memory.d/2026-09-20-ravenpen-com-registered-by-the-operator-the-09-18-hold-is.md

  • [2026-09-19] FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed). → persistent-memory.d/2026-09-19-fv-is-a-dedicated-20-a-circuit-carrying-fv-ml1-and-the-r420.md

  • [2026-09-19] althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning. → persistent-memory.d/2026-09-19-althing-3-7-0-rolled-out-to-both-infra-ops-surfaces-and-the.md

  • [2026-09-19] Three agents commit as one git author, and closing that gap took three instruments to get right. — An unattributable commit (e43e262) appeared in the push set between two of mine — unidentifiable from git… → persistent-memory.d/2026-09-19-three-agents-commit-as-one-git-author-and-closing-that-gap.md

  • [2026-09-19] The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking. → persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md

  • [2026-09-19] elway evaluated when: / creates: / removes: / changed_when: WITHOUT the step's sudo, and it fails silently in the dangerous direction. → persistent-memory.d/2026-09-19-elway-evaluated-when-creates-removes-changedwhen-without.md

  • [2026-09-19] ESH VM 102 (esh-vm-workstation) excluded from the nightly backup job — operator ruling. — It is a Windows 11 Parsec/RDP sandbox (no password, no state to recover), and its vzdump had failed… → persistent-memory.d/2026-09-19-esh-vm-102-esh-vm-workstation-excluded-from-the-nightly.md

  • [2026-09-19] The ops log is BUILT — scripts/ops-log, automatic writers, and a detector for the path they cannot cover. → persistent-memory.d/2026-09-19-the-ops-log-is-built-scripts-ops-log-automatic-writers-and.md

  • [2026-09-19] infra-hermes is this session's ASSISTANT, and the division of labour is now standing policy. — infra-ops keeps improving infrastructure tooling plus the hard calls; infra-hermes does **day-to-day… → persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md

  • [2026-09-18] ⭐⭐⭐ NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure. Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; tailscale ping 373–522 ms → 6 ms direct, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs central-nat, so a policy dstaddr is the REAL internal address, not the VIP. No OOB access — back up with show to a local file and make additive changes ONLY. irv-ml1 still relayed. → persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md

  • [2026-09-18] ⭐⭐⭐ .internal DNS was failing ~10% of lookups fleet-wide, from two independent causes. A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped ratelimit: 20 shared across an entire /24, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ resolv.conf is DHCP-managed — change it at the UDM/FortiGate, not the file. → persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md

  • [2026-09-18] ⭐⭐ SearXNG had ONE working general web engine and every health check said fine. 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — inactive: false only makes an engine SELECTABLE, disabled: false puts it in the DEFAULT set. Now seven. ⭐ This stack tracks :latest ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md

  • [2026-09-18] ⭐⭐ althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames. FTS5 read the hyphen as a column filter; search() caught the error, its probe passed, and it returned []. Routed to forseti (they own the code, I own rollout) → v3.6.3 deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → persistent-memory.d/2026-09-18-althing-363-hyphen-search.md

  • [2026-09-18] ⭐ FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file. docs/fleettools/ + ~/FLEETTOOLS.md; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → persistent-memory.d/2026-09-18-fleettools-agent-index.md

  • [2026-09-18] Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy — (operator-approved). → persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md

  • [2026-09-18] Worldtree's env.sh secrets are vaulted — 10 entries under worldtree/ (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md

  • [2026-09-17] dragonfireacoustics.com expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way. → persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md ⚠ If it is ever transferred, DNS does NOT come with the registration — the nameservers are eNom's name-services.com and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: * (WILDCARD) → 199.250.192.76 which is dead (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), www → 38.120.12.45 (us), and 7 Google Workspace MX records that must not be lost. No DNSSEC (delegationSigned: false), so no transfer complication. ⚠ Also found: no SPF and no DMARC at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the TAC/EPP code from the eNom account, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).

  • [2026-09-17] dragonfireacoustics.com IS configured on pfi-ana-webhost, and the whole thing is dead — a forgotten public-facing VM. → persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md

  • [2026-09-17] PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… → persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md

  • [2026-09-17] The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev — dev_launch.py has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md

  • [2026-09-17] PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… → persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md

  • [2026-09-15] ⚠⚠ --gpu-memory-utilization DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it. Wrong in both directions on fv-ml1: vllm-cyberprev util 0.40 (expect ~39,155 MiB) holds 47,124 (+8 GB over); vllm-gen-small util 0.48 (expect ~46,986) holds 36,942 (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read nvidia-smi --query-compute-apps. Full per-seat residency table + the breeze shuffle arithmetic → persistent-memory.d/2026-09-15-breeze-placement-sizing.md

  • [2026-09-15] breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, move-the-tts-stack-breeze-tts-bragi-tts-gateway), triggered on evacuating embed/rerank/reward. ⚠ Trigger as stated says "gpu0" but those three are on GPU 1 (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only breeze-tts is GPU-resident (~10.3 GiB, growing) while bragi and tts-gateway are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. breeze-tts sizing — original recommendation NOT to move it. ~10.3 GiB measured under load at 53 min uptime, up from 9.2 GiB shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a 1.7 GB margin and shrinking, on the live chat serving path. ⚠ Two measurement traps: it reports nothing at idle on the wrong card (BREEZE_GPU_DEVICES=0 = the 3090, not the A6000), and an early reading understates it. ⭐ The real objection is topology: tts-gateway is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → persistent-memory.d/2026-09-15-breeze-placement-sizing.md

  • [2026-09-15] Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos 7165272) → persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md

  • [2026-09-13] STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. → persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md

  • [2026-08-25] Fused MoE kernel path — DEFERRED, tracked at park fused-moe-kernel-path-for-gemma-4-moe-training (id 47). → persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md

  • [2026-08-24] nconnect=8 on /mnt/smithy — approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread 01M0R46SFYF83099N16WD67KGD.

167 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-10-01] Size + mtime scans to pick reclaimable "staging" on a model store (infra-hermes, irv-ml1). _inbound/retro-diffusion looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.

  • [2026-09-30] Shorter Parakeet slices as Scriberr's memory fix. I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. expandable_segments:True was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.

  • [2026-09-30] Whole-file local-attention Parakeet in Scriberr (context 255/255). OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.

  • [2026-09-30] Start-time midpoint stitching of overlapped Parakeet chunks. It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.

  • [2026-09-30] int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime. The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.

  • [2026-09-30] GPU budgets computed as total − used. nvidia-smi Free is ~640 MiB lower per card (driver reserve). Budget from Free.

  • [2026-09-27] The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel. Every tool worked except get_viewport_screenshot, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → stacks/blender/README.md

  • [2026-09-27] A TCP connect as the "is Blender ready" probe. docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to ping. The same trap applies to any service behind a published port.

  • [2026-09-27] log.exception() in a GPU failure path — the record keeps exc_info, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log traceback.format_exc() text instead (semif-serve engine._guard).

  • [2026-09-27] fla/triton in a slim image without gcc — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.

  • [2026-09-27] uv export → requirements.txt → uv pip install for a torch-from-cu128 project — the export drops uv's per-package index routing; with --extra-index-url … --index-strategy unsafe-best-match, triton came from the pytorch index and failed the lock's hash. Keep uv sync and blank the project version in a deps-only stage instead (services/semif-serve/Dockerfile).

  • [2026-09-27] Translating a CUDA OOM by raising inside except (or from exc) — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, gc.collect() before empty_cache().

  • [2026-09-27] deploy-stack.sh --yes / dns-sync.py fed a blind y — the harness refuses a blind apply. Review with echo n | first, then apply with echo y |.

  • [2026-09-27] Relying on the build cache surviving on esh-ml1 — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the docker builder prune runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.

  • [2026-09-26] Greedy exact-match as parity for a small generative seat (coder) — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md

  • [2026-09-26] zsh traps in ad-hoc test loops: set -- $a does NOT word-split, so alert POSTs went out with empty fields; and local path=… inside a function CLOBBERS $PATH (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and ${=var}.

  • [2026-09-25] Concluding "not on the UDM" from a port table read 15 s after link-up — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory feedback_polled_stats_lag_the_event.

  • [2026-09-25] The esh-pve NVIDIA DKMS recipe on nh3-pve — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flights mokutil --sb-state (7ddd116).

  • [2026-09-25] elway against an unpinned host-key name — every when: hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 in when: should fail.

  • [2026-09-24] Testing "Restore AC Power Loss" with an OS shutdown — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button.

  • [2026-09-23] booth link --help — there is no help flag; it posts --help to the operator's link board as a link. Read booth with no args for usage.

  • [2026-09-21] Using directory mtime as a liveness test when pruning session scratchpads — find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7 deleted an ACTIVE session's working… → persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md

  • [2026-09-18] Routing SearXNG's egress through a SOCKS5 proxy on esh-scale — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md

  • [2026-09-18] api_key: !ENV SEARXNG_BRAVE_API_KEY in searxng settings — this build has NO !ENV YAML constructor, so the file was unparseable and the container crash-looped ten… → persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md

118 older entries archived to archival-memory.md.