Files
esh-pfi-infrastructure/persistent-memory.md
T

82 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-10-01 ~0446 PT (gen-small OOM fixed by parakeet-nemo 0.1.1, with a hard cap and cache return; Prime: leave the gen-small KV and graphs-off as is; arbo webhook fixed (Gitea allow-list); leftovers deleted, repos pushed. Prior: Parakeet seat → unified-en LIVE, U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.

(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the handoff unread across any overnight gap, which is the exact case it exists for. 8 h also matches the global CLAUDE.md and the /snapshot skill default.)

Repo purpose

  • 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up. /tank and other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. See persistent-memory.d/2026-09-10-beszel-fleet-wiring.md and stacks/beszel/README.md.

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) — [2026-09-24] MOTHBALLED by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (e6da607) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecated — althing-cli→postbox, althing-wake-listener→althing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild)
vh/soong-lab Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16)
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path) — [2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev (fleet ownership audit; also CLAUDE.md 2026-09-05). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-10-01 ~0446 PT.

Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)

  • INCIDENT 04:21 PT 2026-10-01: FIXED by nemo-0.1.1 (infra-hermes 963f9ed, live ~0441, infra-ops spot audit passed).
    • Cause: gen-small hit CUDA OOM on a 394 MiB lazily allocated workspace while the seat was parked at its 3,582 MiB window cache (GPU 0 Free 388).
    • The fix: empty_cache around every window; MEM_CAP_MIB=3840 (a hard process ceiling); CUDA_GRAPHS=0, because empty_cache poisons the graph pool (an illegal memory access on the next request).
    • Result: the seat rests at 2,108 MiB and peaks at 3,028 during a 12-min file. gen-small is steady at 36,116 (its runtime growth landed at init, with zero movement on later requests). GPU 0 Free is ~1.05 GB at rest.
    • Cost: graphs-off is ~2–8 ms slower at short clips and ~27 ms at 20–60 s (35 / 40 / 50 / 98 ms against 33 / 36 / 42 / 71), still 4–15× faster than the old seat.
    • Lesson: a shared-card tenant must carry a HARD cap, whether or not it returns memory.
    • Prime ~0445 2026-10-01: leave both as is. gen-small keeps its 8 GiB KV pin (no insurance trim), and CUDA graphs stay OFF (the few-ms cost is accepted; no graph-pool rework).
  • Rollback: docker stop parakeet-nemo && docker start parakeet. The old container and image are kept.
  • GPU 0 is FULL:
    • The seat's steady state is 3,582 MiB (its cached window peak); Free is 385 MiB.
    • vllm-gen-small runs at util 0.36, and its .env holds 0.33 for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×.
    • Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01.
  • Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with --http h11 (httptools 0.8.0 emits HTTP/1.1 200\x00OK, which LiteLLM/httpx rejects).
  • NVIDIA Open Model License accepted for internal use. → persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md

Worldtree U11 memory cutover (demo + personal)

  • Legacy plane OFF since 0115/0120 PT 2026-09-30 (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md
  • U11b STEP 5 DONE 2026-10-02 14:54:03Z on Prime's direct go (in this channel, ~0750 PT). H2 re-check exit 0 on both (demo rows 4/12/5, personal 0/6/2200, agent_self 0); deleted agents/{forseti,lofn,mimir}/memory/<agent>.chroma + memory/context_promotion on demo and personal with literal paths; apis healthy, no restart; stamp sent to worldtree-dev (b193 is theirs to cut). Demo had already lost 22 context_promotion JSONL (2026-07-03) to ledger.py's 90-day JSONL_RETENTION_DAYS sweep on 10-01 07:10:30 — TTL, not an erasure; flagged to worldtree-dev as retirement owner.
  • b193 is cut and pushed (c2d87263, worldtree-dev 2026-10-02); demo deploys via CI, personal on Prime's staging tag. scripts/wt-memory-gate-batch now passes --legacy-mode only when the DEPLOYED sha's harness still has it (git grep at that sha; tested b192 → passed, b193 → omitted), so no timed edit is needed. ⚠ It will REFUSE (exit 3) while demo's deployed sha lags a tree HEAD whose pyproject/uv.lock/packages changed (b193, then the albok commit): by design. Still mine after b193 deploys: remove the residue memory/context_promotion/ledger.db a b192 restart may have re-created; then the parity config retirements (detail file). Verdict ccs come FROM hermes-gateway (never read); reply to infra-hermes.
  • Legacy archive: DESTROY it whole by 2026-10-30, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file.
  • TODO: re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys (list: detail file, "Post-b193 config retirements").

fv-ml1 GPU layout (as of 2026-10-01)

  • GPU 0: cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
  • GPU 1: vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
  • GPU 3: the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, Scriberr (0 idle, ~5.5 GB per job), and MIA auto-rig (one-shot scripts/mia-run, peak 3.4 GB, added 2026-10-01 for dread-dev). When a full-size seat claims GPU 3, Scriberr steps aside to irv-ml1's A6000, not back to GPU 1.

intern-decision (replaced SemIf on 2026-09-30)

  • LIVE 0.1.3 at intern-decision.fv.internal:8033: semif-compatible /decide plus Jev /v1/systemone, 32k tokens, a Triton cache volume. Run scripts/intern-decision-warmup after an IMAGE change. → persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md
  • Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.

Scriberr (fv-ml1 GPU 3)

  • LIVE scriberr:local-blackwell-a353078-dropout2: upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, PARAKEET_MODEL_PATH), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. scripts/scriberr-rebuild re-applies both. → persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md
  • URL is https://scriberr.nh3.phasefinal.com since 2026-10-01 (fronted by the fleet TLS caddy on nh3-dev; Caddyfile backup .bak-20261001-scriberr). Prime's "can't initialize recorder" was the SecureContext rule: getUserMedia is refused on http://. The HTTPS name is in ALLOWED_ORIGINS (playbooks/scriberr-https-origin.yaml); http://10.251.50.54:8080 still works except for recording. ⚠ /opt/docker/compose/scriberr is lkraven-owned, so deploy-stack.sh as infra-ops FAILS there (Permission denied); upload single files with elway --upload --sudo --owner lkraven:lkraven.

nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)

  • nh3-pve: Secure Boot OFF; IGFX restored; NVIDIA 580.178.04 DKMS; kernel 6.8.12-43 (benign btmtk oops every boot). AMT live: static 10.100.250.61 on nh3-mgmt (UDM port 6), KVM on, Opt-in None, password in the vault as nh3-pve/amt-admin. → servers/nh3-pve/README.md
  • nh3-ml1 (CT 109 @ 10.100.50.80):
    • TEI embed/rerank is load-shared with esh-ml1 through the gateway, with router_settings.enable_weighted_failover.
    • brokkr's foundry seats:
      • lfm-vl :8030 (gateway lfm25-vl-3b);
      • lfm-vl-uncensored :8032 (direct; passed brokkr's eval);
      • vibevoice-asr :8031 (audio.cpp, Q8_0).
    • GPU ~11.3/16 GB. → servers/nh3-ml1/README.md
  • Not yet proven: nh3-ml1 coming back by itself after an nh3-pve reboot (only the config was checked).

esh-matter: Matter server for Home Assistant (2026-09-26)

  • CT 111 @ 10.0.90.20, on VLAN 90 (esh-iot) only. matter.js 1.4.0; :5580 is firewalled to HA 10.0.50.46. HA's matter integration is loaded (ha-dev).
  • Next is Prime's: share the Aqara W200 into HA via Matter multi-admin. If commissioning misbehaves, the UniFi mDNS reflector on esh-iot is the first knob (a Prime/infra-ops change). → servers/esh-matter/README.md

Worldtree reward path: code + config live (2026-09-27, Prime)

  • d8726a13 (Domari's Skywork/Selene fix) was ALREADY on origin/main when Prime said to push it: it was pushed as part of 79cce92f at 1221. Demo runs image 79cce92f. Personal still runs 51168a2e, which is older and lacks the fix; its image follows worldtree CI, never a manual up.
  • deploy-wt-config deployed demo and personal (1437): providers.yaml reward_models.scalar-judge (6d4ac44), plus a77639d's corrected selene description, which had never reached the hosts. Both are healthy, with 0 drift. It is inert on personal until that image picks up d8726a13.
  • worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were pushed on Prime's go-ahead (a9d091e..6d4ac44). Origin, the repo and both hosts now agree.

Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime)

  • Prime: "go ahead with gpu 3, both", and he does not use Blender, so agents drive it through MCP. Blender 5.2.2 LTS (linuxserver/selkies image, digest-pinned) with a web desktop at https://10.251.50.54:3001 (vault fv-ml1/blender-web-password). On demand only: scripts/blender-mcp up|down|status. It was left DOWN (GPU 3 back to 2 MiB).
  • MCP: mcp-for-blender 2.1.1 (Dvalin's research pick). The add-on is vendored at 41a18432. The server runs INSIDE the container, and agents reach it as ssh+docker-exec stdio via scripts/blender-mcp. No port is published, and screenshots work because server and Blender share a filesystem. Telemetry is off and safe mode is on. Verified end to end (render, screenshot, safe-mode refusal). Batch path: scripts/blender-run (one-shot docker run --rm, --job staging; first user is draupnir). Access: the shared fleet infra-ops login, with no render-only key (Prime, 2026-09-28). Registration + up/down: PER WORKING SESSION (Prime, 2026-09-28, relayed by draupnir; superseded per-task of 09-27): blender-mcp up + claude mcp add at session start, claude mcp remove + down at its end. Never always-on, never user- or project-wide config. → stacks/blender/README.md
  • Extensions (2026-09-28, draupnir; Prime ruled Blender a MANDATORY pipeline stage): 8 pinned add-ons (stacks/blender/extensions.lock) built by scripts/blender-extensions sync into fv-ml1:/tank/blender-extensions/5.2/system (LIVE), mounted read-only as the System repo; fleet_extensions.py enables them (GUI startup timer; blender-run --extensions). Headless acceptance 8/9 (CAD Sketcher sketching is GUI-only); MCP 9/9 after Prime ran the deploy himself (1505; the classifier had refused mine). SurfacePsycho's eval() is patched to literal_eval (held under MCP). Open: agent-drawn CAD Sketcher geometry (its stateful ops want point picks), and MeasureIt overlays not seen in MCP screenshots. GUI left DOWN.

Zigbee2MQTT on esh-docker-vm (2026-09-27, Prime go-ahead; ha-dev request)

  • LIVE since 1240: Z2M 2.14.1 (digest-pinned) on http://10.0.50.45:8099 (auth token), radio SLZB-MR1U chip 0 tcp://10.0.90.10:6638 (ember), channel 25, PAN 0xCFF4. State and the network key are in /opt/docker/data/zigbee2mqtt (root 0700, restic via /opt/docker, never in git). Vault: esh-docker-vm/zigbee2mqtt-{network-key,frontend-token,mqtt-password}. Broker user zigbee2mqtt added to mosquitto (passwd backup .bak-20260927-z2m).
  • ha-dev deleted the ZHA entry; pairing the Aqara T1 is theirs. ⚠ The "bridge in HA" acceptance first FAILED because HA had had no MQTT since 2026-09-25 06:13. Cause: the esh-docker-vm macvlan shim had no host route to HA (two equal /24s, ens18 wins). Fixed at 1249 with a /32 via the shim (if-up.d hook, playbooks/esh-docker-vm-macvlan-shim-route.yaml), and HA reconnected. My first report called HA "connected" from a client count that included my own probe; $SYS counts are not identities.
  • ⚠ The first credential-mint attempt was blocked by the auto-mode classifier on a peer-relayed approval. It went ahead only on Prime's own go-ahead in this session. Treat that as correct. → stacks/zigbee2mqtt/README.md

restic: credential leak fixed (2026-09-27, Prime)

  • Seven hosts were moved from env-file to repository-file (playbooks/restic-repository-file.yaml), so the systemd units no longer carry the rest-server password. Secrets are vaulted as <host>/etc/restic/{repository,password}. restic.env is KEPT for manual snippets, so a rotation must update the vault, restic.env and repository.
  • vm-esh-nas done too (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). All 8 hosts are clean and vaulted.
  • The passwords were readable until today, so the rotation (Prime's) remains the real fix.
  • 2026-09-27 0859 freshness check: all restic and PBS repos ✅. Those 0100 runs PRE-date the move (done ~0140–0200).
  • VERIFIED 2026-09-28 (infra-hermes, read-only): the first nightly on repository-file (0100 PT) succeeded on all 8 hosts (unit Result=success AND a today-dated snapshot in each repo), and the 0800 freshness check was all-green, PBS included. nh3-dev also passed a content restore (af5580c3).

augaman: face recognition for Cicada

  • v0.1.3 LIVE on esh-ml1:8040 (v0.1.1 first deployed 2026-09-26 2347 PT), stacks/augaman, healthy on CUDA, in nvidia-smi. Built on-box from a git archive of the tag. pytest -m gpu tests/vision 3/3 PASS on v0.1.3.
  • Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27 (Prime OK'd off-site): restic daily 0100 PT → rest-server-ana esh-ml1/, fail-closed hook, secrets vaulted under esh-ml1/etc/restic/. Identity-level restore matched the live gallery (canary, snapshot fd3061a1). The backup gate for real enrollments is MET.
  • Deploy CLOSED by augaman-dev 2026-09-27 0009 PT: HTTP checks passed; the canary survived the recreate and was then deleted, so the gallery is empty and ready for real enrollments. /recognize p50 186 ms (1080p, one face). The detector-latency follow-up is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops.
  • The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT) after the bench. esh-ml1 is the only instance: in the house, backed up, 48 ms per face on v0.1.3.
  • v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later. Speed bench (docs/pfi/augaman-speed-bench/), server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to augaman-dev; the deployment does not use CPU mode.

esh-ml1

  • The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the gateway /scalar-judge passthrough, which is key-gated per key via allowed_passthrough_routes, granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image (see "Worldtree reward path").
  • The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or drop.
  • The Beszel superuser password was echoed into a session transcript (local only, not in memory). Rotation offered to Prime.

Live threads

  • git: eshpfi-management pushed to 6b66207 and worldtree-instance-configs to b6fdd81 (2026-10-01 ~0418, Prime's go). Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with git commit -- <paths> (auto-memory feedback_shared_git_index_commit_pathspecs). graphify-out/GRAPH_REPORT.md stays modified and uncommitted on purpose: it is auto-regenerated.
  • nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
  • Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth 50bfc7b). Pushing it is booth's call, per Prime; it is not ours.
  • ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.

Recent decisions

  • [2026-10-03] nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 nic3, AMT on nic0 with auto kept up, 1 TB wiped into vmstore (913 GiB), stale boot entry gone. My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed auto nic0 and /etc/hosts at 1316 (no reload). Lesson: auto-memory feedback_no_remote_reip_thrash_when_operator_has_console. → servers/nh3-pve-2/README.md

  • [2026-10-03] CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232) fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and ss -Ks any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → services/cira-tunnel-watchdog/. Same day: nh3-pve-2's host moved to its X710 10G port (38:05:25:3b:a0:a8, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS nh3-pve-2.nh3.internal = 10.100.250.62). ⚠ At 1233 its AMT (:a6, UDM port 5) had NO link, probably because the I226 port has no auto line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.

  • [2026-10-03] Worldtree memory-gate duty REFRAMED: regression check, not a gate (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.

  • [2026-10-03] albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1. 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and /health//search answered 500. 0.1.2 (sha256:6f91a33f…, built from albok-dev's LOCAL 544ecfd before the push; origin's tags verified = 544ecfd at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing personal/agent-feedback (group albok-feedback 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (albok/token-nemi; old token still live until albok-dev revokes it) and LiteLLM key albok-nemi (gen-small/summarizer/qwen3-embedding; albok/litellm-key-nemi). Nemi timer INSTALLED 0559 (user units albok-nemi.{service,timer} on nh3-dev, hourly; OnFailure → infra-ops inbox; services/albok-nemi/). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through systemctl --user start albok-nemi.service (no overlap). → stacks/albok-service/README.md

  • [2026-10-02] esh-pve-2 (MS-03 at ESH, future esh-dev host) onboarded: infra-ops (Prime), 1 TB → LVM-thin vmstore, AMT phoning home as esh-pve-2-amt. The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as vmstore (913 GiB; playbooks/esh-pve-2-disk-prep.yaml); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted esh-pve-2/amt-admin) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then amt-cira-setup.py --apply. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (ss -K), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC 38:05:25:3b:9c:12. → servers/esh-pve-2/README.md, servers/pfi-tacticalrmm/README.md

  • [2026-10-02] Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684). Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated worldtree-api only (compose up -d --no-deps) on the .env-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds :latest); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.

  • [2026-10-02] nh3-pve's AMT is in TacticalRMM's MeshCentral (group PFI-AMT, device nh3-pve-amt). The adds had failed silently because TRMM installs MeshCentral WANonly (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). UPDATE 0859: it now PHONES HOME (CIRA) — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via scripts/amt-cira-setup.py, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body servers/nh3-pve/amt-ethernet-static-revert.xml. 2FA still not forced. Prime's MeshCentral login token is vaulted pfi-tacticalrmm/meshcentral-login-token. → servers/pfi-tacticalrmm/README.md

  • [2026-10-02] Prime: dev-backup gets dailies + weeklies — DONE. Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (retention.py beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.

  • [2026-10-02] ✅ ESH tank (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629. raidz2-0 disk wwn-0x5000c500c91df554: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. Prime (via Miranda) 1925: verify, then clear + full scrub. Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. zpool clear 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107. History explains it: after the Aug 20 replace, the resilver ended errors=6676436, someone ran only an error-scrub (zpool scrub -e, 4 s) and then zpool clear 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse 6.19M as an integer); it still fires on completion. 0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m. 0118: SECOND full scrub started WITHOUT zpool clear, so any new error shows as count > 6494530 (zpool status -p); watcher tested with +1/null/finished controls; ETA ~0700. RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. zpool clear 0629 → pool 'tank' is healthy. Reported to Miranda once. PVE 8→9 plan for esh-pve-cluster written, NOT executed (Prime via Miranda): docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md. Also found: remote access to ESH rides CT 108 on pve (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.

  • [2026-10-02] albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): http://albok.nh3.internal:8392 (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key albok-service (qwen3-embedding only) + bootstrap admin token in the vault under albok/. Upgraded to 0.1.1 (e349d50) at 1805: /health now ok (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted albok/token-albok-dev). → stacks/albok-service/README.md

  • [2026-10-02] nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03. IDE-R from MeshCentral (ISO proxmox-ve_9.2-1.iso in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (…:a0:a9, USW port 1) only links during firmware, so it likely has no driver. On-site checklist: USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (wipefs + zpool labelclear if ZFS: two rpools collide), reserve .62, onboard infra-ops, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.

  • [2026-10-02] Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as nh3-pve-2 at NH3 (Prime, ~0910). Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then scripts/amt-cira-setup.py (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE). Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted nh3-pve-2/amt-admin): KVM on, listener on, OptIn 0 (ACM), amt-cira-setup.py --apply → phoning home at once; MeshCentral device nh3-pve-2-amt (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts esh-dev at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software). Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the *.nh3.phasefinal.com rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the nh3-dev/ vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.

  • [2026-10-02] nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot). Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB /swapfile (the old sda5 blocked growth), RESUME=none, all initrds rebuilt (playbooks/nh3-dev-grow-root.yaml). ⚠ TODO: delete VM snapshot pre-rootgrow-20261002 on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot). Trigger: uptime -s on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (swapon --show = /swapfile; systemd-analyze shows no ~30 s stall; journalctl -b | grep -i resume has no 'waiting for resume device'), and only then qm delsnapshot 102 pre-rootgrow-20261002.

  • [2026-10-01] MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev). scripts/mia-run --job DIR -- name=in.glb ..., image local/mia:0.1.0 (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset jasongzy/Mixamo, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → stacks/mia/README.md

  • [2026-10-01] Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed. They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (vastblue/praxis/references/PraxisPM_Rev0_07.11.26, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs chmod -R u+w before rm, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot 2026-09-29_0602 with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at ~/.config/dev-backup/dev-backup.log.20261001-0530.gz until someone deletes it.

  • [2026-10-01] Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.

  • [2026-10-01] ⚠ Gitea's [webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8 silently REJECTS headscale mesh IPs (100.64.0.0/10). The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN 10.6.110.50:9009; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without secret and branch_filter drops both, so always resend them.

  • [2026-10-01] vh/arbo push webhook repointed from the retired wg0 lifeline 10.100.79.3:9009 to irv-ml1's mesh address 100.64.0.6:9009. It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner irv-ml1-arbo was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.

  • [2026-10-01] Prime: delete the bench leftovers, no upstream for Scriberr, push. DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; unified-en KEPT, the live seat mounts it), plus /tank/spikes/scriberr-slicer (including the private copies of Prime's recordings) and /tank/spikes/parakeet-ab. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.

  • [2026-10-01] irv-ml1 /storetank reclaim done: Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread 01M3TCSYRSFNAPA9BPQTMFQ6KJ).

  • [2026-09-30] Parakeet speech seat → parakeet-unified-en-0.6b under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137. → persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md

  • [2026-09-30] Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30. → persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md

  • [2026-09-30] SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev /v1/systemone at 32k tokens on GPU 1, with a Triton warm-up cache volume. → persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md

  • [2026-09-30] Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept. → persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md

  • [2026-09-29] Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches). → persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md

  • [2026-09-28] Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors. → persistent-memory.d/2026-09-28-worldtree-u10-backfill.md

  • [2026-09-28] Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired. → persistent-memory.d/2026-09-28-bonsai-ternary-spike.md

  • [2026-09-28] blender-run gained --cpu (no GPU attached) and a fixed hostname fv-ml1-blender (draupnir). The design stage never renders, so it stays off GPU 3.

  • [2026-09-28] Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (--extensions). SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → stacks/blender/README.md § Extensions

  • [2026-09-27] hermes-gateway restarted 0401 for highseat-dev (SVOS v2.1.12: propose_decision gained seat_up, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling zellij-fleet@Claude at boot remains Prime's call.

  • [2026-09-27] SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime): wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md

  • [2026-09-27] Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port. → stacks/blender/README.md

  • [2026-09-27] esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August). → servers/esh-docker-vm/README.md

  • [2026-09-27] Zigbee2MQTT live on esh-docker-vm :8099 (PAN 0xCFF4, ch 25) replacing ZHA; network key vaulted, data host-only. Prime go-ahead in this session; a peer-relayed approval was blocked by the permission gate. → stacks/zigbee2mqtt/README.md

  • [2026-09-27] semif-serve 0.1.4: object states ending in ), ; or } no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart. Prime ruled; heid bug hunt folded. → stacks/semif/README.md

  • [2026-09-27] SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win. Build nothing (Prime). Henge 88 carries it. → persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md

  • [2026-09-27] SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21. Recommendations await Prime. → persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md

  • [2026-09-27] SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels. Tracked in the in-flight SemIf section. → persistent-memory.d/2026-09-27-semif-order-averaging.md — DONE: 0.1.3 live, with fast kernels adopted (77b8cb4).

  • [2026-09-27] restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime). Rotation stays Prime's. → persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md

  • [2026-09-27] infra-ops bootstrapped on vm-esh-nas by Prime (the first account at the fleet-pinned uid/gid 850; bootstrap-infra-ops-user.yaml now pins 850 when free, d775a01).

  • [2026-09-27] augaman: second instance on fv-ml1 benched, then REMOVED (Prime). esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (ad484c3, bench docs/pfi/augaman-speed-bench/).

  • [2026-09-27] Created empty private repo corviduo/svoperatingsystem (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. svos-dev → the OS session; The High Seat's session is highseat-dev again.

  • [2026-09-26] augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy. → persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md — [2026-09-27] DONE: deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (d8f59a1).

  • [2026-09-26] zellij-fleet@.service installed, NOT enabled (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling @Claude at boot is Prime's call when seat_up ships. Tracked: 0ad7799, services/zellij-fleet/README.md.

  • [2026-09-26] btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority). Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked: servers/nh3-pve/README.md.

  • [2026-09-26] GPU-LXC Temperature alerts watch the GPU, not the host CPU (SENSORS=-coretemp_*,acpitz); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump. 6d901c0.

  • [2026-09-26] Abliterated LFM2.5-VL-3B parallel seat (:8032) passed brokkr's 6-image eval and stays (direct only; stock seat kept for SFW A/B). 9cc3824, 944bb36.

  • [2026-09-25] AMT follow-ups PARKED until the two MS-03s arrive and esh-pve's MS-01 is on AMT (Prime 2326). — Parked: MeshCommander test; dummy HDMI plug, then NanoKVM → gx10; an OOB path not through nh3-scale… → persistent-memory.d/2026-09-25-amt-follow-ups-parked-until-the-two-ms-03s-arrive-and-esh.md

  • [2026-09-25] nh3-pve AMT LIVE: static 10.100.250.61 on nh3-mgmt (UDM port 6), KVM on, Opt-in None — (Prime). → persistent-memory.d/2026-09-25-nh3-pve-amt-live-static-10-100-250-61-on-nh3-mgmt-udm-port.md

  • [2026-09-25] pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot. servers/pfi-gx10/README.md, servers/esh-docker-vm/README.md.

  • [2026-09-26] VibeVoice ASR → Q8_0 (Prime). WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed.

  • [2026-09-26] esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only — , for ha-dev (operator-approved, relayed). → persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md

  • [2026-09-26] Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime). — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… → persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md

  • [2026-09-26] Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed): — LFM2.5-VL-3B on llama.cpp :8030 (gateway lfm25-vl-3b, LiteLLM restarted 36 s at 0039) and… → persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md

  • [2026-09-26] Coder seat STAYS on fv-ml1 (Prime). — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… → persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md

  • [2026-09-25] nh3-ml1 LIVE after the NH3 visit (SB off, IGFX restored): driver + CT 109 + TEI, parity vs esh-ml1 indistinguishable (cos min 0.999993 = own floor; overlap@10 1.000 vs MRL-256 control 0.684), same speed; monitoring/DNS wired. Autonomous: lxc-pve 6.0.0-2 upgrade on nh3-pve (Docker-in-LXC fix #7006). Gateway routing + AMT follow-up are Prime's. → persistent-memory.d/2026-09-25-nh3-ml1-live.md

  • [2026-09-25] nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls. Tracked: 7ddd116, Current state. → persistent-memory.d/2026-09-25-nh3-ml1-standup.md

  • [2026-09-25] nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333). Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked: servers/nh3-pve/README.md (2b49be4, cdd7605).

  • [2026-09-25] MS-01 foot-gun: a GPU in the PCIe slot renames every NIC — (the slot's root port takes bus 01, so the X710 goes enp2s0f0np0→enp3s0f0np0). → persistent-memory.d/2026-09-25-ms-01-foot-gun-a-gpu-in-the-pcie-slot-renames-every-nic.md

  • [2026-09-25] Direct ESH→esh-ml1 consumer path — PARKED (no ESH-side callers in 7 days; consumers go through the gateway). Tracked at persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md.

  • [2026-09-25] Created empty private repo corviduo/norn (git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; the norn-dev handle is the operator's to declare.

  • [2026-09-25] Reward seat audited (Skywork-Reward-V2-Llama-3.1-8B still #1 of 188 on RewardBench 2; our AWQ ≈ bf16 within noise; double BOS costs ~2.7 pts) and moved fv-ml1 → esh-ml1; esh-ml1 monitoring wired (Beszel+GPU, Kuma, Homepage, Dozzle); Dozzle hub's stale agent IPs fixed; nh3-dev Beszel agent revived. → persistent-memory.d/2026-09-25-reward-move-and-esh-ml1-monitoring.md

  • [2026-09-25] TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend for qwen3-embedding + reranker; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1. → persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md

  • [2026-09-24] esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM order: 2 failover; dead DB alias reranker-a3-bge-v2-m3 (pre-relocation IP) repaired. → persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md

  • [2026-09-24] esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank (implementation deferred to next session, tracked here + servers/esh-pve/README.md). → persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md

  • [2026-09-24] pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25 (tracking: servers/pfi-gx10/README.md, e5197a3). — [2026-09-25] ✅ VALIDATED: Prime pulled and replugged AC, and it came up by itself at 1108:54 (2b49be4). → persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md

  • [2026-09-24] NH3 power outage recovered — pbs-nh3 had no onboot (set), NFS boot race fixed with automount (1cbde50), every other Claude session on nh3-dev died. → persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md

  • [2026-09-24] Miranda standing order is a repo CLAUDE.md operating parameter — (4b29492, aligned to the global send protocol in bcf3342): high-urgency matters go to her, fixed or not… → persistent-memory.d/2026-09-24-miranda-standing-order-is-a-repo-claude-md-operating.md

  • [2026-09-24] task-board mothballed (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (e6da607). Its hooks had sent no traffic in 30 days.

  • [2026-09-24] Military 24-hour Pacific clock times carried into the Codex/Grok shared bootstrap docs/fleettools/AGENT-BOOTSTRAP.md (ad2b4d9); Claude seats get it from the global CLAUDE.md.

  • [2026-09-24] Worldtree admin.memory.forget stays OFF on demo/personal until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config.

  • [2026-09-24] A git checkout under root:docker needs safe.directory for its deploy user — the 09-14 normalization (826a63b) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (eea9eb2). Sweep found no other case.

  • [2026-09-23] elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts. → persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md

  • [2026-09-23] headscale-ddns hardened — (fedd4b6): Cloudflare calls retry and validate every body (pick()), no write without both IDs, the run… → persistent-memory.d/2026-09-23-headscale-ddns-hardened.md

  • [2026-09-23] esh-docker-vm restic was skipped 09-22..23 by my own Kuma move — a dead uptime-kuma lookup aborted pre-backup.sh under set -e (25e41d2). → persistent-memory.d/2026-09-23-esh-docker-vm-restic-was-skipped-09-22-23-by-my-own-kuma.md

  • [2026-09-23] hermes-gateway restart exit-1 is a Hermes race, not a crash — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in… → persistent-memory.d/2026-09-23-hermes-gateway-restart-exit-1-is-a-hermes-race-not-a-crash.md

  • [2026-09-23] Booth link board cleared to 14 durable links (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted at scriberr.fv.internal:8080). Prime's rule: durable debug links only.

  • [2026-09-22] Both carried calls approved — build the NRestarts flap sampler (163bb97); the restic content-assertion ruling is ratified and stays (ba60fda).

  • [2026-09-22] safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes. — ⚠ The package installs INERT and looks fine — Debian's /etc/zsh/zprofile has 0 non-comment lines so the… → persistent-memory.d/2026-09-22-safe-rm-installed-on-nh3-dev-and-delegated-fleet-wide-to.md

  • [2026-09-22] The acceptance probe for a guard must not be able to destroy what it tests — (infra-hermes). → persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md

  • [2026-09-22] ⚠ {"sent": true} is a claim about transmission, never about effect. pane_send structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable /clear and its receipt reported success. Consumed an operator approval. → persistent-memory.d/2026-09-22-instrument-errors.md

  • [2026-09-22] D-0010/D-0011 were misrouted to this seat by pane_find matching a ROLLING PANE TITLE. — Genuine and operator-approved, wrong seat; fleet_telemetry held the right mapping and carries the warning… → persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md

  • [2026-09-22] headscale-ddns exited 1 silently and the alarm carried no cause — both failure paths were || exit 1 with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. 30517fd.

  • [2026-09-22] ⭐⭐ Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject. --state=running skipped the units most needing hooks; awk '{print $1}' dropped systemd's ●-decorated FAILED rows; grep -ic restic on the wrapper missed the check script; restic ls's header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. Every one reported cleanly while looking at the wrong thing. The rule is not "verify" — it is verify, then ask what the verification could not have seen. → persistent-memory.d/2026-09-22-instrument-errors.md

  • [2026-09-22] Fleet alert bridge generalized — beszel-althing → althing-alert-bridge, route registry (/beszel + /kuma), each with its own… → persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md

  • [2026-09-22] Uptime Kuma rebuilt from scratch on 2.5.5, moved esh-docker-vm → ana-docker. — ⚠ :latest is a TRAP — it tracks 1.x, so an Aug-2026 pull gave a Dec-2024 build. → persistent-memory.d/2026-09-22-uptime-kuma-rebuilt-from-scratch-on-2-5-5-moved-esh-docker.md

  • [2026-09-22] Beszel and Uptime Kuma are DISJOINT, not redundant — Beszel's alerts bind to a system with a threshold; there is no URL column, so it is structurally incapable… → persistent-memory.d/2026-09-22-beszel-and-uptime-kuma-are-disjoint-not-redundant.md

  • [2026-09-22] Failed-START alarms on 23 nh3-dev units — (services/althing-notify-failure/). → persistent-memory.d/2026-09-22-failed-start-alarms-on-23-nh3-dev-units.md

  • [2026-09-22] Backup coverage is a property of the SYSTEM, never one job's scope — establish it by querying the repo for the path in a real snapshot, never by reading a job's SRC=. → persistent-memory.d/2026-09-22-backup-coverage-is-a-property-of-the-system-never-one-job-s.md

  • [2026-09-22] restic checks now assert CONTENT and are DISCOVERED not enumerated — conjunctive (recent AND present AND restores non-zero bytes); repos discovered from each NAS, hand list… → persistent-memory.d/2026-09-22-restic-checks-now-assert-content-and-are-discovered-not.md

  • [2026-09-22] irv-ml1: Irvine is a TENANCY behind a Fortinet PFI does not control — its TLS inspection breaks Tailscale relay/control intermittently (41 cert warnings/week). → persistent-memory.d/2026-09-22-irv-ml1-irvine-is-a-tenancy-behind-a-fortinet-pfi-does-not.md

  • [2026-09-21] ⭐⭐⭐ lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone. 720 generations, 3 arms. Voice +0.152 at 2.9× floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and ~3/4 of the gain survives stripping every punctuation mark, so it is not the cheap win. ⭐ Memorisation: ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12, all 96 matches READ and every one stock grammar (he looked at the wolf and he looked at him); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: 20% / 28% of generations overshoot the 90–140 band against base's 1%, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribed score_beats.py's v1 criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. → persistent-memory.d/2026-09-21-lv-mccarthy-gate.md

  • [2026-09-21] The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. — On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one… → persistent-memory.d/2026-09-21-the-two-epoch-recipe-is-now-0-for-3-and-this-time-the-loss.md

  • [2026-09-21] The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass. → persistent-memory.d/2026-09-21-the-memorisation-control-that-shipped-broken-is-now.md

  • [2026-09-21] My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug. → persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md

  • [2026-09-21] ⭐⭐⭐ The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable. Claim dropped by a sub-tool, hook behind graphify's eight exit 0s, no handle in the env, ssh-target written as a hostname. The general shape is configured ≠ effective; twelve instruments reported confidently and wrongly across three days, five of them mine. → persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md

  • [2026-09-21] ⭐⭐ The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing — a reveal handler Jinja discarded for sitting after {% endblock %}, and a × a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. scripts/layout-probe.py took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → persistent-memory.d/2026-09-21-booth-two-dead-controls.md

  • [2026-09-21] ⭐⭐ claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to pfi; vh-token use is now standing-authorized from the vault. ⚠ vh is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ ~/.config/claude-bot/gitea-token is DEAD and had been misreporting permissions; the working one is gitea-token-repo-create. → persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md

  • [2026-09-21] ⭐⭐ nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from /tmp, which this box sweeps at 3 days. babyyarros existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md

  • [2026-09-21] Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested — (34c4179, e574b91, 2e08edc). → persistent-memory.d/2026-09-21-draupnir-geometry-engine-provisioned-on-irv-ml1-and.md

  • [2026-09-21] Backup alarm verdict split: STALE (exit 1) vs ERRORED-JOBS (exit 3) — (7fe4102). → persistent-memory.d/2026-09-21-backup-alarm-verdict-split-stale-exit-1-vs-errored-jobs.md

  • [2026-09-21] vh/forgefirm mirrored — from github.com/openglow-org/forgefirm, following the house convention read off the existing 17: vh/… → persistent-memory.d/2026-09-21-vh-forgefirm-mirrored.md

  • [2026-09-20] ravenpen.com REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered. — infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a… → persistent-memory.d/2026-09-20-ravenpen-com-registered-by-the-operator-the-09-18-hold-is.md

  • [2026-09-19] FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed). → persistent-memory.d/2026-09-19-fv-is-a-dedicated-20-a-circuit-carrying-fv-ml1-and-the-r420.md

  • [2026-09-19] althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning. → persistent-memory.d/2026-09-19-althing-3-7-0-rolled-out-to-both-infra-ops-surfaces-and-the.md

  • [2026-09-19] Three agents commit as one git author, and closing that gap took three instruments to get right. — An unattributable commit (e43e262) appeared in the push set between two of mine — unidentifiable from git… → persistent-memory.d/2026-09-19-three-agents-commit-as-one-git-author-and-closing-that-gap.md

  • [2026-09-19] The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking. → persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md

  • [2026-09-19] elway evaluated when: / creates: / removes: / changed_when: WITHOUT the step's sudo, and it fails silently in the dangerous direction. → persistent-memory.d/2026-09-19-elway-evaluated-when-creates-removes-changedwhen-without.md

  • [2026-09-19] ESH VM 102 (esh-vm-workstation) excluded from the nightly backup job — operator ruling. — It is a Windows 11 Parsec/RDP sandbox (no password, no state to recover), and its vzdump had failed… → persistent-memory.d/2026-09-19-esh-vm-102-esh-vm-workstation-excluded-from-the-nightly.md

  • [2026-09-19] The ops log is BUILT — scripts/ops-log, automatic writers, and a detector for the path they cannot cover. → persistent-memory.d/2026-09-19-the-ops-log-is-built-scripts-ops-log-automatic-writers-and.md

  • [2026-09-19] infra-hermes is this session's ASSISTANT, and the division of labour is now standing policy. — infra-ops keeps improving infrastructure tooling plus the hard calls; infra-hermes does **day-to-day… → persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md

  • [2026-09-18] ⭐⭐⭐ NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure. Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; tailscale ping 373–522 ms → 6 ms direct, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs central-nat, so a policy dstaddr is the REAL internal address, not the VIP. No OOB access — back up with show to a local file and make additive changes ONLY. irv-ml1 still relayed. → persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md

  • [2026-09-18] ⭐⭐⭐ .internal DNS was failing ~10% of lookups fleet-wide, from two independent causes. A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped ratelimit: 20 shared across an entire /24, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ resolv.conf is DHCP-managed — change it at the UDM/FortiGate, not the file. → persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md

  • [2026-09-18] ⭐⭐ SearXNG had ONE working general web engine and every health check said fine. 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — inactive: false only makes an engine SELECTABLE, disabled: false puts it in the DEFAULT set. Now seven. ⭐ This stack tracks :latest ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md

  • [2026-09-18] ⭐⭐ althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames. FTS5 read the hyphen as a column filter; search() caught the error, its probe passed, and it returned []. Routed to forseti (they own the code, I own rollout) → v3.6.3 deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → persistent-memory.d/2026-09-18-althing-363-hyphen-search.md

  • [2026-09-18] ⭐ FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file. docs/fleettools/ + ~/FLEETTOOLS.md; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → persistent-memory.d/2026-09-18-fleettools-agent-index.md

  • [2026-09-18] Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy — (operator-approved). → persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md

  • [2026-09-18] Worldtree's env.sh secrets are vaulted — 10 entries under worldtree/ (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md

  • [2026-09-17] dragonfireacoustics.com expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way. → persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md ⚠ If it is ever transferred, DNS does NOT come with the registration — the nameservers are eNom's name-services.com and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: * (WILDCARD) → 199.250.192.76 which is dead (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), www → 38.120.12.45 (us), and 7 Google Workspace MX records that must not be lost. No DNSSEC (delegationSigned: false), so no transfer complication. ⚠ Also found: no SPF and no DMARC at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the TAC/EPP code from the eNom account, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).

  • [2026-09-17] dragonfireacoustics.com IS configured on pfi-ana-webhost, and the whole thing is dead — a forgotten public-facing VM. → persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md

  • [2026-09-17] PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… → persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md

  • [2026-09-17] The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev — dev_launch.py has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md

  • [2026-09-17] PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… → persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md

  • [2026-09-15] ⚠⚠ --gpu-memory-utilization DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it. Wrong in both directions on fv-ml1: vllm-cyberprev util 0.40 (expect ~39,155 MiB) holds 47,124 (+8 GB over); vllm-gen-small util 0.48 (expect ~46,986) holds 36,942 (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read nvidia-smi --query-compute-apps. Full per-seat residency table + the breeze shuffle arithmetic → persistent-memory.d/2026-09-15-breeze-placement-sizing.md

  • [2026-09-15] breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, move-the-tts-stack-breeze-tts-bragi-tts-gateway), triggered on evacuating embed/rerank/reward. ⚠ Trigger as stated says "gpu0" but those three are on GPU 1 (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only breeze-tts is GPU-resident (~10.3 GiB, growing) while bragi and tts-gateway are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. breeze-tts sizing — original recommendation NOT to move it. ~10.3 GiB measured under load at 53 min uptime, up from 9.2 GiB shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a 1.7 GB margin and shrinking, on the live chat serving path. ⚠ Two measurement traps: it reports nothing at idle on the wrong card (BREEZE_GPU_DEVICES=0 = the 3090, not the A6000), and an early reading understates it. ⭐ The real objection is topology: tts-gateway is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → persistent-memory.d/2026-09-15-breeze-placement-sizing.md

  • [2026-09-15] Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos 7165272) → persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md

  • [2026-09-13] STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. → persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md

  • [2026-08-25] Fused MoE kernel path — DEFERRED, tracked at park fused-moe-kernel-path-for-gemma-4-moe-training (id 47). → persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md

  • [2026-08-24] nconnect=8 on /mnt/smithy — approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread 01M0R46SFYF83099N16WD67KGD.

167 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-10-01] Size + mtime scans to pick reclaimable "staging" on a model store (infra-hermes, irv-ml1). _inbound/retro-diffusion looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.

  • [2026-09-30] Shorter Parakeet slices as Scriberr's memory fix. I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. expandable_segments:True was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.

  • [2026-09-30] Whole-file local-attention Parakeet in Scriberr (context 255/255). OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.

  • [2026-09-30] Start-time midpoint stitching of overlapped Parakeet chunks. It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.

  • [2026-09-30] int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime. The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.

  • [2026-09-30] GPU budgets computed as total − used. nvidia-smi Free is ~640 MiB lower per card (driver reserve). Budget from Free.

  • [2026-09-27] The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel. Every tool worked except get_viewport_screenshot, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → stacks/blender/README.md

  • [2026-09-27] A TCP connect as the "is Blender ready" probe. docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to ping. The same trap applies to any service behind a published port.

  • [2026-09-27] log.exception() in a GPU failure path — the record keeps exc_info, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log traceback.format_exc() text instead (semif-serve engine._guard).

  • [2026-09-27] fla/triton in a slim image without gcc — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.

  • [2026-09-27] uv export → requirements.txt → uv pip install for a torch-from-cu128 project — the export drops uv's per-package index routing; with --extra-index-url … --index-strategy unsafe-best-match, triton came from the pytorch index and failed the lock's hash. Keep uv sync and blank the project version in a deps-only stage instead (services/semif-serve/Dockerfile).

  • [2026-09-27] Translating a CUDA OOM by raising inside except (or from exc) — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, gc.collect() before empty_cache().

  • [2026-09-27] deploy-stack.sh --yes / dns-sync.py fed a blind y — the harness refuses a blind apply. Review with echo n | first, then apply with echo y |.

  • [2026-09-27] Relying on the build cache surviving on esh-ml1 — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the docker builder prune runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.

  • [2026-09-26] Greedy exact-match as parity for a small generative seat (coder) — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md

  • [2026-09-26] zsh traps in ad-hoc test loops: set -- $a does NOT word-split, so alert POSTs went out with empty fields; and local path=… inside a function CLOBBERS $PATH (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and ${=var}.

  • [2026-09-25] Concluding "not on the UDM" from a port table read 15 s after link-up — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory feedback_polled_stats_lag_the_event.

  • [2026-09-25] The esh-pve NVIDIA DKMS recipe on nh3-pve — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flights mokutil --sb-state (7ddd116).

  • [2026-09-25] elway against an unpinned host-key name — every when: hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 in when: should fail.

  • [2026-09-24] Testing "Restore AC Power Loss" with an OS shutdown — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button.

  • [2026-09-23] booth link --help — there is no help flag; it posts --help to the operator's link board as a link. Read booth with no args for usage.

  • [2026-09-21] Using directory mtime as a liveness test when pruning session scratchpads — find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7 deleted an ACTIVE session's working… → persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md

  • [2026-09-18] Routing SearXNG's egress through a SOCKS5 proxy on esh-scale — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md

  • [2026-09-18] api_key: !ENV SEARXNG_BRAVE_API_KEY in searxng settings — this build has NO !ENV YAML constructor, so the file was unparseable and the container crash-looped ten… → persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md

118 older entries archived to archival-memory.md.