Files
esh-pfi-infrastructure/persistent-memory.md
T

473 lines
67 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Persistent memory — eshpfi-management
_Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
> handoff from the previous session), then delete it. Older than 8 hours:
> stale — delete it unread.
>
> _(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the
> handoff unread across any overnight gap, which is the exact case it exists for.
> 8 h also matches the global CLAUDE.md and the `/snapshot` skill default.)_
## Repo purpose
- **2026-09-10 Beszel fleet wiring:** all seven requested hosts plus existing corviduo-dev report up. `/tank` and other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to **infra-ops**, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. See `persistent-memory.d/2026-09-10-beszel-fleet-wiring.md` and `stacks/beszel/README.md`.
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
`/opt/docker/compose/<stack>/`; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. **It was originally
spun up to handle the fleet backups** — keep that lens when triaging
backup/storage issues.
## Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) — **[2026-09-24] MOTHBALLED** by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (`e6da607`) | push-to-main → CI deploys (2026-04-29) |
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
| `vh/althing` | Lean trusted inter-agent message bus — **v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback)**: ONE container on nh3-dev at `http://10.100.50.40:8390` is the only stateful component; `althing-po-herald` one per box; `althing-listen` one per session; `postbox` is the client. **Every v2 command was DELETED, not deprecated** — `althing-cli`→`postbox`, `althing-wake-listener`→`althing-listen`, `althing-light-monitor`/`althing-receiver` gone. Sessions need BOTH `ALTHING_POST_OFFICE` and `ALTHING_HANDLE`; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md` | per-box install (NOT CI-deploy); **nh3-dev** = container host + repo; **nh3-extdev** = system WHEEL at `/opt/uv-tools`, needs its own wheel install (`playbooks/nh3-extdev-althing-v3.yaml`) |
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
| `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker**; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
| `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` |
| `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
| `vh/zonos-gateway` | OpenAI-compatible TTS gateway over stock ZONOS2 (`:8890` irv-ml1); emotion **dials-first** + voice mapping; reached via LiteLLM `ext-tts` alias. **v0.2.1 (2026-07-18): voice-resolved emotion presets** (`resolve_preset(name,voice)`; angry/happy/startled_happy per-voice). 8 voices incl. 4 clones | pushed to gitea (main `8f1885b`/`v0.2.1`); **deployed irv-ml1 tree still NON-git** (hand-updated build context — CI-wire = open follow-up). Spec `docs/EMOTION-DIALS-SPEC.md`; host-managed voices bind-mount (`./voices:/app/voices`, drop wav + restart, no rebuild) |
| `vh/soong-lab` | Noonien Soong character-design studio (SPA + /api + WT `/bifrost/tool-call`); **containerized 2026-07-18**, LIVE on corviduo-dev `:8443` (image `vh/soong-lab:latest`). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host | CI = Gitea Actions build+push+**DEPLOY** on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; **auto-redeploy LIVE 2026-07-18** — runner SSHes corviduo-dev as `deploy`, `compose pull && up -d` from **/opt/soong-lab**, health-gated on /api/version). Manual redeploy `sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'`. → `archival-memory.md` (archived 2026-08-16) |
| `model-training-forge` (mtf-dev) | Fine-tuning recipe forge; **T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06)** (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(`vh/volva` + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev `.service` units were removed —
no longer deployed sidecars here. See Recent decisions.)
- **Two-layer backups** — Backrest orchestrates restic for file+DB (5
fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for
VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore +
cross-site restic targets — see `docs/runbooks/disaster-recovery.md`
for the blast-radius matrix. **⚠️ The restic file+DB layer routes
through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 →
ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @
nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS
export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.)
- **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace
model/dataset onto ana-ml2's shared cache at
`/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns.
- **Worldtree admin auth — per-instance.** Each Worldtree deployment
(demo :8080, personal :8081, pinned :8082) has its own Heimdall
registry and its own bootstrap admin key. Infra-ops's stored
long-lived admin key (`key_id 61419c92`) at
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`
auths against **demo only**. Personal-instance admin (the
`~/.config/worldtree/personal-admin-token`, mode 600) POSTs
`/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no
scope param** — scopes are tier-derived). **On-instance mint recipe
(cleaner than DB-manip):** `docker exec worldtree-worldtree-api-1` POST
`/admin/keys` with the in-container `WORLDTREE_BOOTSTRAP_ADMIN_KEY`; cleartext
once in `.key`=`wt_live_+16hex`. auto-memory `reference_worldtree_demo_key_mint`.
- **Per-project user keys against personal Worldtree** (issued
2026-05-19): `skaldsong:79744637`, `skaldsong:7c1dbbbe`,
`althing:50d85460`, `mead-hall:a360822d`. Mint via `/admin/keys`, drop
value to `/tmp/wt-personal-<name>.key` mode 600, dev collects +
shreds (DO NOT cat to chat transcript).
- **Skaldsong CD pattern (registry-pull).** vh/skaldsong's CI builds and
pushes `gitea.phasefinal.com/vh/skaldsong:<sha>` + `:latest`;
`playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates.
SHA-pin only. Prereq: host needs `docker login gitea.phasefinal.com` once.
- **gitea internal route for fleet hosts.** gitea is a container on
**ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo
hosts must use this internal route, NOT public `gitea.phasefinal.com`
(`38.120.12.44`) — the public path fail2bans the host egress IP. Full
gotcha in `docs/orientation.md` → Git/gitea.
- **docker-as-root pattern** (for ops with no admin API, or to edit
deploy-owned/root-owned files without sudo): `docker run --rm -v
<target-dir>:/wt docker:cli sh -c "..."`. docker-group membership is
effectively root via bind-mount. **Foot-gun: relative paths in compose.yaml
resolve against the sandbox CWD but the daemon interprets them against the
HOST fs — always pass `-e VAR=/abs/path` for any relative-default config dir.**
- **`scripts/elway` sudo handling** — elway prompts for the sudo password
ONCE via `getpass` before the first `sudo: true` step → can't run
unattended from a non-TTY tool if any step needs sudo. Sudo-free
playbooks run fully non-interactive over key SSH.
- **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo
on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On
ana-docker: **default `ssh ana-docker` = `lkraven`** (docker-group, NO
passwordless sudo); **`ssh infra-ops@ana-docker` HAS NOPASSWD root**. **→
For any sudo op on ana-docker, use `ssh infra-ops@ana-docker`.** `ssh
infra-ops@10.100.10.50` (nh3-dev) ALSO NOPASSWD sudo; on **nh3-extdev** infra-ops
is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path) — **[2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev** (fleet ownership audit; also CLAUDE.md 2026-09-05). **irv-ml1:
`ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-10-01 ~0420 PT._
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
- **⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:**
- `vllm-gen-small`'s EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache.
- vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23.
- I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB.
- **infra-hermes is tasked with nemo-0.1.1:** `empty_cache` after windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth.
- **Awaiting Prime:** trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3.
- Until fixed, a long transcription can re-grow the seat's cache and starve gen-small.
- **LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, `local/parakeet-nemo:nemo-0.1.0`, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLM `ext-stt`/`whisper-1` unchanged. **infra-ops audit PASSED 0137.**
- p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat.
- WER: LibriSpeech clean 1.965, other 3.026.
- Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word.
- **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept.
- **GPU 0 is FULL:**
- The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**.
- `vllm-gen-small` runs at util 0.36, and its `.env` holds **0.33** for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×.
- Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01.
- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects).
- NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
### Worldtree U11 memory cutover (demo + personal)
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- **Daily gate batches:** infra-hermes runs `scripts/wt-memory-gate-batch` from 2026-10-01 and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it.
- **⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:**
1. Run `docker exec -i <api> python - < scripts/wt-h2-count.py` VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output.
2. Delete LIVE, with the api running, using literal paths only: `agents/{forseti,lofn,mimir}/memory/<agent>.chroma` and `memory/context_promotion`, on BOTH instances.
3. Send worldtree-dev the stamp; b193 ships after it.
- A b192 restart re-creates an empty schema-only `ledger.db`. That is residue: say so in the stamp and remove it after b193.
- **Legacy archive:** DESTROY it whole by **2026-10-30**, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file.
- **TODO:** re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys.
### fv-ml1 GPU layout (as of 2026-10-01)
- **GPU 0:** cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
- **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
- **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1.
### intern-decision (replaced SemIf on 2026-09-30)
- **LIVE 0.1.3** at `intern-decision.fv.internal:8033`: semif-compatible `/decide` plus Jev `/v1/systemone`, 32k tokens, a Triton cache volume. Run `scripts/intern-decision-warmup` after an IMAGE change. → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
- **Open:** label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.
### Scriberr (fv-ml1 GPU 3)
- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. `scripts/scriberr-rebuild` re-applies both. → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
- **nh3-pve:** Secure Boot OFF; IGFX restored; NVIDIA 580.178.04 DKMS; kernel
6.8.12-43 (benign btmtk oops every boot). **AMT live**: static
`10.100.250.61` on nh3-mgmt (UDM port 6), KVM on, Opt-in None, password in the
vault as `nh3-pve/amt-admin`. → `servers/nh3-pve/README.md`
- **nh3-ml1** (CT 109 @ `10.100.50.80`):
- TEI embed/rerank is **load-shared with esh-ml1** through the gateway, with
`router_settings.enable_weighted_failover`.
- brokkr's foundry seats:
- `lfm-vl` `:8030` (gateway `lfm25-vl-3b`);
- `lfm-vl-uncensored` `:8032` (direct; passed brokkr's eval);
- `vibevoice-asr` `:8031` (audio.cpp, **Q8_0**).
- GPU ~11.3/16 GB. → `servers/nh3-ml1/README.md`
- Not yet proven: nh3-ml1 coming back by itself after an nh3-pve reboot (only the
config was checked).
### esh-matter: Matter server for Home Assistant (2026-09-26)
- CT 111 @ `10.0.90.20`, on VLAN 90 (esh-iot) only. matter.js 1.4.0; `:5580` is
firewalled to HA `10.0.50.46`. HA's matter integration is loaded (ha-dev).
- **Next is Prime's:** share the Aqara W200 into HA via Matter multi-admin. If
commissioning misbehaves, the UniFi mDNS reflector on esh-iot is the first knob
(a Prime/infra-ops change). → `servers/esh-matter/README.md`
### Worldtree reward path: code + config live (2026-09-27, Prime)
- `d8726a13` (Domari's Skywork/Selene fix) was ALREADY on origin/main when Prime said to push it: it
was pushed as part of `79cce92f` at 1221. Demo runs image `79cce92f`. **Personal still runs
`51168a2e`**, which is older and lacks the fix; its image follows worldtree CI, never a manual up.
- `deploy-wt-config` deployed demo and personal (1437): `providers.yaml` `reward_models.scalar-judge`
(6d4ac44), plus a77639d's corrected selene description, which had never reached the hosts. Both
are healthy, with 0 drift. It is inert on personal until that image picks up d8726a13.
- worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were
**pushed on Prime's go-ahead** (a9d091e..6d4ac44). Origin, the repo and both hosts now agree.
### Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime)
- Prime: "go ahead with gpu 3, both", and he does not use Blender, so **agents drive it through MCP**.
Blender 5.2.2 LTS (linuxserver/selkies image, digest-pinned) with a web desktop at
`https://10.251.50.54:3001` (vault `fv-ml1/blender-web-password`). **On demand only:**
`scripts/blender-mcp up|down|status`. It was left DOWN (GPU 3 back to 2 MiB).
- MCP: `mcp-for-blender` 2.1.1 (Dvalin's research pick). The add-on is vendored at `41a18432`. The
server runs INSIDE the container, and agents reach it as ssh+docker-exec stdio via
`scripts/blender-mcp`. No port is published, and screenshots work because server and Blender
share a filesystem. Telemetry is off and safe mode is on. Verified end to end (render, screenshot,
safe-mode refusal). **Batch path: `scripts/blender-run`** (one-shot `docker run --rm`, `--job` staging; first user is draupnir). **Access: the shared fleet `infra-ops` login, with no render-only key (Prime, 2026-09-28).** **Registration + up/down: PER WORKING SESSION (Prime, 2026-09-28, relayed by draupnir; superseded per-task of 09-27)**: `blender-mcp up` + `claude mcp add` at session start, `claude mcp remove` + `down` at its end. Never always-on, never user- or project-wide config.
→ `stacks/blender/README.md`
- **Extensions (2026-09-28, draupnir; Prime ruled Blender a MANDATORY pipeline stage):** 8 pinned
add-ons (`stacks/blender/extensions.lock`) built by `scripts/blender-extensions sync` into
`fv-ml1:/tank/blender-extensions/5.2/system` (LIVE), mounted read-only as the System repo;
`fleet_extensions.py` enables them (GUI startup timer; `blender-run --extensions`). Headless
acceptance 8/9 (CAD Sketcher sketching is GUI-only); **MCP 9/9 after Prime ran the deploy himself
(1505; the classifier had refused mine)**. SurfacePsycho's eval() is patched to literal_eval
(held under MCP). Open: agent-drawn CAD Sketcher geometry (its stateful ops want point picks),
and MeasureIt overlays not seen in MCP screenshots. GUI left DOWN.
### Zigbee2MQTT on esh-docker-vm (2026-09-27, Prime go-ahead; ha-dev request)
- **LIVE since 1240:** Z2M 2.14.1 (digest-pinned) on `http://10.0.50.45:8099` (auth token), radio
SLZB-MR1U chip 0 `tcp://10.0.90.10:6638` (ember), channel 25, **PAN 0xCFF4**. State and the network
key are in `/opt/docker/data/zigbee2mqtt` (root 0700, restic via /opt/docker, never in git). Vault:
`esh-docker-vm/zigbee2mqtt-{network-key,frontend-token,mqtt-password}`. Broker user `zigbee2mqtt`
added to mosquitto (passwd backup `.bak-20260927-z2m`).
- **ha-dev deleted the ZHA entry**; pairing the Aqara T1 is theirs. ⚠ The "bridge in HA" acceptance first
FAILED because HA had had no MQTT since 2026-09-25 06:13. Cause: the esh-docker-vm macvlan shim had no
host route to HA (two equal /24s, ens18 wins). Fixed at 1249 with a /32 via the shim (if-up.d hook,
`playbooks/esh-docker-vm-macvlan-shim-route.yaml`), and HA reconnected. My first report called HA
"connected" from a client count that included my own probe; `$SYS` counts are not identities.
- ⚠ The first credential-mint attempt was blocked by the auto-mode classifier on a peer-relayed
approval. It went ahead only on Prime's own go-ahead in this session. Treat that as correct.
→ `stacks/zigbee2mqtt/README.md`
### restic: credential leak fixed (2026-09-27, Prime)
- Seven hosts were moved from `env-file` to `repository-file` (`playbooks/restic-repository-file.yaml`),
so the systemd units no longer carry the rest-server password. Secrets are vaulted as
`<host>/etc/restic/{repository,password}`. `restic.env` is KEPT for manual snippets, so
**a rotation must update the vault, `restic.env` and `repository`.**
- **vm-esh-nas done too** (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). **All 8 hosts are clean
and vaulted.**
- The passwords were readable until today, so the rotation (Prime's) remains the real fix.
- 2026-09-27 0859 freshness check: all restic and PBS repos ✅. Those 0100 runs PRE-date the move
(done ~0140–0200).
- **VERIFIED 2026-09-28 (infra-hermes, read-only):** the first nightly on `repository-file` (0100 PT)
succeeded on all 8 hosts (unit Result=success AND a today-dated snapshot in each repo), and the
0800 freshness check was all-green, PBS included. nh3-dev also passed a content restore (`af5580c3`).
### augaman: face recognition for Cicada
- **v0.1.3 LIVE on esh-ml1:8040** (v0.1.1 first deployed 2026-09-26 2347 PT), `stacks/augaman`,
healthy on CUDA, in `nvidia-smi`. Built on-box from a `git archive` of the tag.
`pytest -m gpu tests/vision` 3/3 PASS on v0.1.3.
- **Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27** (Prime OK'd off-site): restic
daily 0100 PT → rest-server-ana `esh-ml1/`, fail-closed hook, secrets vaulted under
`esh-ml1/etc/restic/`. Identity-level restore matched the live gallery (canary,
snapshot `fd3061a1`). **The backup gate for real enrollments is MET.**
- **Deploy CLOSED by augaman-dev 2026-09-27 0009 PT:** HTTP checks passed; the canary
survived the recreate and was then deleted, so the gallery is empty and ready for real
enrollments. `/recognize` p50 186 ms (1080p, one face). The detector-latency follow-up
is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops.
- **The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT)** after the bench. esh-ml1 is the only
instance: in the house, backed up, 48 ms per face on v0.1.3.
- **v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later.** Speed bench (`docs/pfi/augaman-speed-bench/`),
server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode
REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to
augaman-dev; the deployment does not use CPU mode.
### esh-ml1
- The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the
gateway `/scalar-judge` passthrough, which is key-gated per key via `allowed_passthrough_routes`,
granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image
(see "Worldtree reward path").
- The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or
drop.
- The Beszel superuser password was echoed into a session transcript (local
only, not in memory). Rotation offered to Prime.
### Live threads
- git: **eshpfi-management pushed to `6b66207` and worldtree-instance-configs to `b6fdd81` (2026-10-01 ~0418, Prime's go).** Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- <paths>` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated.
- nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`).
Pushing it is booth's call, per Prime; it is not ours.
- ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
## Recent decisions
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
- `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
- `[2026-09-29]` **Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches).** → `persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md`
- `[2026-09-28]` **Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors.** → `persistent-memory.d/2026-09-28-worldtree-u10-backfill.md`
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
- `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3.
- `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
- `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md`
- `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md`
- `[2026-09-27]` **Zigbee2MQTT live on esh-docker-vm :8099 (PAN 0xCFF4, ch 25) replacing ZHA; network key vaulted, data host-only.** Prime go-ahead in this session; a peer-relayed approval was blocked by the permission gate. → `stacks/zigbee2mqtt/README.md`
- `[2026-09-27]` **semif-serve 0.1.4: object states ending in `)`, `;` or `}` no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart.** Prime ruled; heid bug hunt folded. → `stacks/semif/README.md`
- `[2026-09-27]` **SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win.** Build nothing (Prime). Henge 88 carries it. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md`
- `[2026-09-27]` **SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21.** Recommendations await Prime. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md`
- `[2026-09-27]` **SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels.** Tracked in the in-flight SemIf section. → `persistent-memory.d/2026-09-27-semif-order-averaging.md` — **DONE:** 0.1.3 live, with fast kernels adopted (`77b8cb4`).
- `[2026-09-27]` **restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime).** Rotation stays Prime's. → `persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md`
- `[2026-09-27]` **infra-ops bootstrapped on vm-esh-nas by Prime** (the first account at the fleet-pinned uid/gid 850; `bootstrap-infra-ops-user.yaml` now pins 850 when free, `d775a01`).
- `[2026-09-27]` **augaman: second instance on fv-ml1 benched, then REMOVED (Prime).** esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (`ad484c3`, bench `docs/pfi/augaman-speed-bench/`).
- `[2026-09-27]` Created empty private repo `corviduo/svoperatingsystem` (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. `svos-dev` → the OS session; The High Seat's session is `highseat-dev` again.
- `[2026-09-26]` **augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy.** → `persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md` — `[2026-09-27]` **DONE:** deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (`d8f59a1`).
- `[2026-09-26]` **`zellij-fleet@.service` installed, NOT enabled** (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling `@Claude` at boot is Prime's call when seat_up ships. Tracked: `0ad7799`, `services/zellij-fleet/README.md`.
- `[2026-09-26]` **btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority).** Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked: `servers/nh3-pve/README.md`.
- `[2026-09-26]` **GPU-LXC Temperature alerts watch the GPU, not the host CPU** (`SENSORS=-coretemp_*,acpitz`); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump. `6d901c0`.
- `[2026-09-26]` **Abliterated LFM2.5-VL-3B parallel seat (:8032) passed brokkr's 6-image eval and stays** (direct only; stock seat kept for SFW A/B). `9cc3824`, `944bb36`.
- `[2026-09-25]` **AMT follow-ups PARKED until the two MS-03s arrive and esh-pve's MS-01 is on AMT (Prime 2326).** — Parked: MeshCommander test; dummy HDMI plug, then NanoKVM → gx10; an OOB path not through nh3-scale… → `persistent-memory.d/2026-09-25-amt-follow-ups-parked-until-the-two-ms-03s-arrive-and-esh.md`
- `[2026-09-25]` **nh3-pve AMT LIVE: static `10.100.250.61` on nh3-mgmt (UDM port 6), KVM on, Opt-in None** — (Prime). → `persistent-memory.d/2026-09-25-nh3-pve-amt-live-static-10-100-250-61-on-nh3-mgmt-udm-port.md`
- `[2026-09-25]` **pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot.** `servers/pfi-gx10/README.md`, `servers/esh-docker-vm/README.md`.
- `[2026-09-26]` **VibeVoice ASR → Q8_0 (Prime).** WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed.
- `[2026-09-26]` **esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only** — , for ha-dev (operator-approved, relayed). → `persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md`
- `[2026-09-26]` **Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime).** — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… → `persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md`
- `[2026-09-26]` **Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed):** — LFM2.5-VL-3B on llama.cpp `:8030` (gateway `lfm25-vl-3b`, LiteLLM restarted 36 s at 0039) and… → `persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md`
- `[2026-09-26]` **Coder seat STAYS on fv-ml1 (Prime).** — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… → `persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md`
- `[2026-09-25]` **nh3-ml1 LIVE after the NH3 visit (SB off, IGFX restored): driver + CT 109 + TEI, parity vs esh-ml1 indistinguishable (cos min 0.999993 = own floor; overlap@10 1.000 vs MRL-256 control 0.684), same speed; monitoring/DNS wired. Autonomous: lxc-pve 6.0.0-2 upgrade on nh3-pve (Docker-in-LXC fix #7006). Gateway routing + AMT follow-up are Prime's.** → `persistent-memory.d/2026-09-25-nh3-ml1-live.md`
- `[2026-09-25]` **nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls.** Tracked: `7ddd116`, Current state. → `persistent-memory.d/2026-09-25-nh3-ml1-standup.md`
- `[2026-09-25]` **nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333).** Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked: `servers/nh3-pve/README.md` (`2b49be4`, `cdd7605`).
- `[2026-09-25]` **MS-01 foot-gun: a GPU in the PCIe slot renames every NIC** — (the slot's root port takes bus 01, so the X710 goes `enp2s0f0np0`→`enp3s0f0np0`). → `persistent-memory.d/2026-09-25-ms-01-foot-gun-a-gpu-in-the-pcie-slot-renames-every-nic.md`
- `[2026-09-25]` **Direct ESH→esh-ml1 consumer path — PARKED** (no ESH-side callers in 7 days; consumers go through the gateway). Tracked at `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`.
- `[2026-09-25]` Created empty private repo `corviduo/norn` (git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; the `norn-dev` handle is the operator's to declare.
- `[2026-09-25]` **Reward seat audited (Skywork-Reward-V2-Llama-3.1-8B still #1 of 188 on RewardBench 2; our AWQ ≈ bf16 within noise; double BOS costs ~2.7 pts) and moved fv-ml1 → esh-ml1; esh-ml1 monitoring wired (Beszel+GPU, Kuma, Homepage, Dozzle); Dozzle hub's stale agent IPs fixed; nh3-dev Beszel agent revived.** → `persistent-memory.d/2026-09-25-reward-move-and-esh-ml1-monitoring.md`
- `[2026-09-25]` **TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend for `qwen3-embedding` + `reranker`; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1.** → `persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md`
- `[2026-09-24]` **esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM `order: 2` failover; dead DB alias `reranker-a3-bge-v2-m3` (pre-relocation IP) repaired.** → `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`
- `[2026-09-24]` **esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank** (implementation deferred to next session, tracked here + `servers/esh-pve/README.md`). → `persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md`
- `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). — `[2026-09-25]` ✅ **VALIDATED**: Prime pulled and replugged AC, and it came up by itself at 1108:54 (`2b49be4`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
- `[2026-09-24]` **NH3 power outage recovered** — pbs-nh3 had no `onboot` (set), NFS boot race fixed with automount (`1cbde50`), every other Claude session on nh3-dev died. → `persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md`
- `[2026-09-24]` **Miranda standing order is a repo CLAUDE.md operating parameter** — (`4b29492`, aligned to the global send protocol in `bcf3342`): high-urgency matters go to her, fixed or not… → `persistent-memory.d/2026-09-24-miranda-standing-order-is-a-repo-claude-md-operating.md`
- `[2026-09-24]` **task-board mothballed** (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (`e6da607`). Its hooks had sent no traffic in 30 days.
- `[2026-09-24]` **Military 24-hour Pacific clock times** carried into the Codex/Grok shared bootstrap `docs/fleettools/AGENT-BOOTSTRAP.md` (`ad2b4d9`); Claude seats get it from the global CLAUDE.md.
- `[2026-09-24]` **Worldtree `admin.memory.forget` stays OFF on demo/personal** until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config.
- `[2026-09-24]` **A git checkout under `root:docker` needs `safe.directory` for its deploy user** — the 09-14 normalization (`826a63b`) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (`eea9eb2`). Sweep found no other case.
- `[2026-09-23]` **elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts.** → `persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md`
- `[2026-09-23]` **headscale-ddns hardened** — (`fedd4b6`): Cloudflare calls retry and validate every body (`pick()`), no write without both IDs, the run… → `persistent-memory.d/2026-09-23-headscale-ddns-hardened.md`
- `[2026-09-23]` **esh-docker-vm restic was skipped 09-22..23 by my own Kuma move** — a dead uptime-kuma lookup aborted `pre-backup.sh` under `set -e` (`25e41d2`). → `persistent-memory.d/2026-09-23-esh-docker-vm-restic-was-skipped-09-22-23-by-my-own-kuma.md`
- `[2026-09-23]` **hermes-gateway restart exit-1 is a Hermes race, not a crash** — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in… → `persistent-memory.d/2026-09-23-hermes-gateway-restart-exit-1-is-a-hermes-race-not-a-crash.md`
- `[2026-09-23]` **Booth link board cleared to 14 durable links** (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted at `scriberr.fv.internal:8080`). Prime's rule: durable debug links only.
- `[2026-09-22]` **Both carried calls approved** — build the `NRestarts` flap sampler (`163bb97`); the restic content-assertion ruling is ratified and stays (`ba60fda`).
- `[2026-09-22]` **safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes.** — ⚠ The package installs INERT and looks fine — Debian's `/etc/zsh/zprofile` has 0 non-comment lines so the… → `persistent-memory.d/2026-09-22-safe-rm-installed-on-nh3-dev-and-delegated-fleet-wide-to.md`
- `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md`
- `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md`
- `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md`
- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
- `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md`
- `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md`
- `[2026-09-22]` **Uptime Kuma rebuilt from scratch on 2.5.5, moved esh-docker-vm → ana-docker.** — ⚠ `:latest` is a TRAP — it tracks 1.x, so an Aug-2026 pull gave a Dec-2024 build. → `persistent-memory.d/2026-09-22-uptime-kuma-rebuilt-from-scratch-on-2-5-5-moved-esh-docker.md`
- `[2026-09-22]` **Beszel and Uptime Kuma are DISJOINT, not redundant** — Beszel's alerts bind to a *system* with a threshold; there is no URL column, so it is structurally incapable… → `persistent-memory.d/2026-09-22-beszel-and-uptime-kuma-are-disjoint-not-redundant.md`
- `[2026-09-22]` **Failed-START alarms on 23 nh3-dev units** — (`services/althing-notify-failure/`). → `persistent-memory.d/2026-09-22-failed-start-alarms-on-23-nh3-dev-units.md`
- `[2026-09-22]` **Backup coverage is a property of the SYSTEM, never one job's scope** — establish it by querying the repo for the path in a real snapshot, never by reading a job's `SRC=`. → `persistent-memory.d/2026-09-22-backup-coverage-is-a-property-of-the-system-never-one-job-s.md`
- `[2026-09-22]` **restic checks now assert CONTENT and are DISCOVERED not enumerated** — conjunctive (recent AND present AND restores non-zero bytes); repos discovered from each NAS, hand list… → `persistent-memory.d/2026-09-22-restic-checks-now-assert-content-and-are-discovered-not.md`
- `[2026-09-22]` **irv-ml1: Irvine is a TENANCY behind a Fortinet PFI does not control** — its TLS inspection breaks Tailscale relay/control intermittently (41 cert warnings/week). → `persistent-memory.d/2026-09-22-irv-ml1-irvine-is-a-tenancy-behind-a-fortinet-pfi-does-not.md`
- `[2026-09-21]` ⭐⭐⭐ **lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone.** 720 generations, 3 arms. Voice +0.152 at **2.9×** floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and **~3/4 of the gain survives stripping every punctuation mark**, so it is not the cheap win. ⭐ Memorisation: **ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12**, all 96 matches READ and every one stock grammar (`he looked at the wolf and he looked at him`); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: **20% / 28% of generations overshoot the 90–140 band against base's 1%**, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribed `score_beats.py`'s **v1** criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by **0.01 against a 0.200 floor**, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. → `persistent-memory.d/2026-09-21-lv-mccarthy-gate.md`
- `[2026-09-21]` **The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong.** — On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one… → `persistent-memory.d/2026-09-21-the-two-epoch-recipe-is-now-0-for-3-and-this-time-the-loss.md`
- `[2026-09-21]` **The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass.** → `persistent-memory.d/2026-09-21-the-memorisation-control-that-shipped-broken-is-now.md`
- `[2026-09-21]` **My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug.** → `persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md`
- `[2026-09-21]` ⭐⭐⭐ **The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable.** Claim dropped by a sub-tool, hook behind graphify's eight `exit 0`s, no handle in the env, ssh-target written as a hostname. The general shape is **configured ≠ effective**; twelve instruments reported confidently and wrongly across three days, five of them mine. → `persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md`
- `[2026-09-21]` ⭐⭐ **The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing** — a reveal handler Jinja discarded for sitting after `{% endblock %}`, and a `×` a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. `scripts/layout-probe.py` took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → `persistent-memory.d/2026-09-21-booth-two-dead-controls.md`
- `[2026-09-21]` ⭐⭐ **claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to `pfi`; vh-token use is now standing-authorized from the vault.** ⚠ `vh` is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ `~/.config/claude-bot/gitea-token` is DEAD and had been misreporting permissions; the working one is `gitea-token-repo-create`. → `persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md`
- `[2026-09-21]` ⭐⭐ **nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from `/tmp`, which this box sweeps at 3 days.** `babyyarros` existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → `persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md`
- `[2026-09-21]` **Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested** — (34c4179, e574b91, 2e08edc). → `persistent-memory.d/2026-09-21-draupnir-geometry-engine-provisioned-on-irv-ml1-and.md`
- `[2026-09-21]` **Backup alarm verdict split: `STALE` (exit 1) vs `ERRORED-JOBS` (exit 3)** — (7fe4102). → `persistent-memory.d/2026-09-21-backup-alarm-verdict-split-stale-exit-1-vs-errored-jobs.md`
- `[2026-09-21]` **`vh/forgefirm` mirrored** — from `github.com/openglow-org/forgefirm`, following the house convention read off the existing 17: `vh/`… → `persistent-memory.d/2026-09-21-vh-forgefirm-mirrored.md`
- `[2026-09-20]` **`ravenpen.com` REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered.** — infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a… → `persistent-memory.d/2026-09-20-ravenpen-com-registered-by-the-operator-the-09-18-hold-is.md`
- `[2026-09-19]` **FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed).** → `persistent-memory.d/2026-09-19-fv-is-a-dedicated-20-a-circuit-carrying-fv-ml1-and-the-r420.md`
- `[2026-09-19]` **althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning.** → `persistent-memory.d/2026-09-19-althing-3-7-0-rolled-out-to-both-infra-ops-surfaces-and-the.md`
- `[2026-09-19]` **Three agents commit as one git author, and closing that gap took three instruments to get right.** — An unattributable commit (`e43e262`) appeared in the push set between two of mine — unidentifiable from git… → `persistent-memory.d/2026-09-19-three-agents-commit-as-one-git-author-and-closing-that-gap.md`
- `[2026-09-19]` **The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking.** → `persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md`
- `[2026-09-19]` **elway evaluated `when:` / `creates:` / `removes:` / `changed_when:` WITHOUT the step's sudo, and it fails silently in the dangerous direction.** → `persistent-memory.d/2026-09-19-elway-evaluated-when-creates-removes-changedwhen-without.md`
- `[2026-09-19]` **ESH VM 102 (`esh-vm-workstation`) excluded from the nightly backup job — operator ruling.** — It is a Windows 11 Parsec/RDP **sandbox** (no password, no state to recover), and its vzdump had failed… → `persistent-memory.d/2026-09-19-esh-vm-102-esh-vm-workstation-excluded-from-the-nightly.md`
- `[2026-09-19]` **The ops log is BUILT — `scripts/ops-log`, automatic writers, and a detector for the path they cannot cover.** → `persistent-memory.d/2026-09-19-the-ops-log-is-built-scripts-ops-log-automatic-writers-and.md`
- `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md`
- `[2026-09-18]` ⭐⭐⭐ **NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure.** Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; `tailscale ping` 373–522 ms → **6 ms direct**, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs `central-nat`, so a policy `dstaddr` is the REAL internal address, not the VIP. No OOB access — back up with `show` to a local file and make additive changes ONLY. irv-ml1 still relayed. → `persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md`
- `[2026-09-18]` ⭐⭐⭐ **`.internal` DNS was failing ~10% of lookups fleet-wide, from two independent causes.** A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped `ratelimit: 20` shared across an entire **/24**, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ `resolv.conf` is DHCP-managed — change it at the UDM/FortiGate, not the file. → `persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md`
- `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md`
- `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md`
- `[2026-09-18]` ⭐ **FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.** `docs/fleettools/` + `~/FLEETTOOLS.md`; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → `persistent-memory.d/2026-09-18-fleettools-agent-index.md`
- `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md`
- `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md`
- `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md`
⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).
- `[2026-09-17]` **`dragonfireacoustics.com` IS configured on `pfi-ana-webhost`, and the whole thing is dead — a forgotten public-facing VM.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md`
- `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md`
- `[2026-09-17]` **The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev** — `dev_launch.py` has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → `persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md`
- `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md`
- `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
- `[2026-09-15]` **breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, `move-the-tts-stack-breeze-tts-bragi-tts-gateway`), triggered on evacuating embed/rerank/reward.** ⚠ Trigger as stated says "gpu0" but those three are on **GPU 1** (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only `breeze-tts` is GPU-resident (~10.3 GiB, **growing**) while `bragi` and `tts-gateway` are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. **breeze-tts sizing — original recommendation NOT to move it.** ~**10.3 GiB** measured under load at 53 min uptime, **up from 9.2 GiB** shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a **1.7 GB margin and shrinking**, on the live chat serving path. ⚠ Two measurement traps: it reports **nothing at idle on the wrong card** (`BREEZE_GPU_DEVICES=0` = the **3090**, not the A6000), and an early reading understates it. ⭐ The real objection is **topology**: `tts-gateway` is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
- `[2026-09-15]` **Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos `7165272`)** → `persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md`
- `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md`
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md`
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
_167 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
- `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.
- `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.
- `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.
- `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`.
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
- `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
- `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`.
- `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`.
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
- `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md`
- `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`.
- `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`.
- `[2026-09-25]` **The esh-pve NVIDIA DKMS recipe on nh3-pve** — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flights `mokutil --sb-state` (`7ddd116`).
- `[2026-09-25]` **elway against an unpinned host-key name** — every `when:` hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 in `when:` should fail.
- `[2026-09-24]` **Testing "Restore AC Power Loss" with an OS shutdown** — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button.
- `[2026-09-23]` **`booth link --help`** — there is no help flag; it posts `--help` to the operator's link board as a link. Read `booth` with no args for usage.
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md`
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`
- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md`
_118 older entries archived to archival-memory.md._