memory: snapshot — the ops log, and a day spent on instruments that report without looking

Archived 15 entries (Recent decisions 14, Tried and abandoned 1) oldest-first
to archival-memory.md; 5 held back on the open-deferred-work guard and 164 on
the 14-day guard, so the index stays over the soft cap at 477 lines. An
over-cap file that keeps live decisions beats a scannable one that lost a
belayed item.

Four new detail files cover the day: the ops log and its four self-inflicted
failure modes, the Booth's two dead controls and the four-iteration layout
probe, the Gitea org grant plus the dead claude-bot token that had been
misreporting permissions, and the disk triage that rescued a LoRA adapter from
a directory this box sweeps at three days.

lv-mccarthy's run outcome remains unverified after two days and is the first
line of the in-flight section and step 1 of the handoff.
This commit is contained in:
vh
2026-09-21 14:26:55 -07:00
parent 2e08edcaab
commit e52def115c
14 changed files with 572 additions and 428 deletions
+53 -115
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-21 ~05:55 UTC (⭐ **ravenpen.com registered** by the operator — Cloudflare registrar, expires 2028-09-20, zone active but BARE; hamr-dev answered, the 09-18 hold discharged. ⭐ Booth gained keep-both-ways + per-item cosmetic blur. ⭐ Backup alarm verdict split: STALE (exit 1) vs ERRORED-JOBS (exit 3) — it was crying STALE over a fleet whose every body was fresh. ⭐ The ops log is BUILT; commit attribution now works end-to-end, both controls measured. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED.)_
_Last updated: 2026-09-21 ~14:30 PT (⭐ **the ops log is BUILT** and then failed four ways in its first hours — every one recording something unfindable; the day's subject was instruments that report without looking. ⭐ Draupnir engine COMPLETE and acceptance-tested on irv-ml1. ⭐ Booth gained blur + a closed keep round trip after shipping two dead controls. ⭐ claude-bot is an org Owner; cicada+draupnir moved to `pfi`; vh-token use standing-authorized from the vault. ⭐ nh3-dev 84%→72%, and a LoRA adapter rescued from a 3-day-swept /tmp. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED after two days.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,112 +115,65 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-09-19 ~07:15 PT._
_As of 2026-09-21 ~14:30 PT._
### ✅ BUILT — the ops log (`scripts/ops-log`), 2026-09-19
### ⚠ FIRST — lv-mccarthy's run outcome is STILL UNVERIFIED
Shipped. `scripts/ops-log` + `docs/pfi/ops-log.md`, wired into `deploy-stack.sh`
(claims + records) and `elway` (records). Baseline laid: **136 stacks across 6 hosts**
marked pre-ops-log, so the detector starts from today instead of reporting the whole
fleet as unattributable forever.
Carried unchanged from the 2026-09-19 handoff and untouched for two full days.
Look at `~/r49-runs/mccarthy-4b-pairs-3ep/` before anything else. ⚠ `pfi-gx10`
does not resolve from nh3-dev by that name. Assume nothing — it was never
checked, not checked-and-found-good. Then: checkpoint selection off the loss
curve → the v2 gate (`eval-*.sh`) → ship-or-park. ⚠ Before the gate, settle how
the VOICE axis is read: D1 pre-registered a punctuation-normalised secondary
read, and `--system-from` hands the register's tics to the BASE arm too, which
changes what the primary number means.
**The four open questions, settled:**
### Draupnir engine — COMPLETE on irv-ml1, acceptance passing
1. **Where it lives — CENTRAL on nh3-dev** (`<repo>/.ops-log/`, gitignored), not
per-host and not the post office. Decider: both agents run as the *same unix user*
on nh3-dev (infra-hermes is a **user** unit under `lkraven`), so one file is shared
instantly with zero provisioning. Per-host needs a writable path on ~25 heterogeneous
boxes and puts the record of "we changed host Y" *on host Y*. syslog/journald looked
free but journald shows an unprivileged reader only their own `_UID` — the log would
have split silently between the `infra-ops` and `lkraven` halves of the fleet. The
post office is a bus, and an outage there would block ops during the incident you are
reconstructing. ⚠ **Known hole, stated not papered over:** an actor on a box *other
than nh3-dev* is uncovered. Today that is only the operator's laptop; a third agent
elsewhere is what would force a revisit.
2. **Claim = advisory, enforced in the tooling.** `deploy-stack.sh` refuses (exit 3) a
stack another agent holds, across the diff, the y/N prompt AND the apply — the whole
review window, which is where the 09-18 collision actually happened. Acquire is
`mkdir` (atomic → genuinely race-free). TTL 30m; a stale claim auto-breaks **and the
break is logged**, so an ineffective claim is visible rather than silent.
3. **Writers are AUTOMATIC.** This was the one that mattered — a log you must remember
to write is the same class of instrument as a health check that passes in both states.
4. **There is a DETECTOR, not just a rule.** `ops-log audit` asks each host what changed
on disk and compares it to the newest log line for that stack. Covers the manual
`ssh`-and-edit path the automatic writers structurally cannot.
build123d 0.12.0 + OCP, numpy 2.4.6, trimesh 5.1.0 on py3.11.2; FreeCAD 1.0.0
AppImage headless; OrcaSlicer 2.4.2 containerised at `~/bin/orca-slice`;
artifact root `/mnt/smithy/draupnir` 2775, 500 GB budget. The contrastive
control pair PASSES. **The operator has moved to a code session inside the
Draupnir repo**, so the next questions come from there rather than from
brokkr-smithy-dev, and everything is in the repo (b8db50b) rather than only in
the althing thread.
⚠ **Found and fixed a bug in my own detector mid-build:** it printed "audit clean" for a
host it never reached. Now `INCOMPLETE` + exit 5 — an unreachable host is not a clean
host. Both directions proven on a live host: after baseline, a mtime-only `touch` on
`nh3-docker/beszel-agent-nh3` fired the detector at 15 s resolution, and recording it
cleared it.
### The fleet is quiet and nothing is blocked on me
**Not covered, on purpose:** raw `ssh` (audit is the backstop), per-host claims for elway,
`corviduo-dev` (CI/CD rewrites the tree constantly → permanent false positives), the
SureFire tenant hosts. **Follow-ons:** hook `dns-sync.py` / UniFi / FortiGate helpers so
control-plane changes record themselves; run `audit` on a timer.
- **Backups:** `RESULT: all backups fresh`. The two `⏸` policy exclusions
(ana-scale CT 114, esh-vm-workstation VM 102) are correct and visible.
infra-hermes holds the watch and names CT 107 explicitly rather than trusting
the absence of red — ⚠ that lock was RELEASED BY HAND, not self-healed, and
its mechanism is still unexplained.
- **Disk:** nh3-dev at 72%, 66 GB free after the 2026-09-21 triage.
- **ops-log:** clean audit across 6 hosts; commit attribution working.
- **19 commits unpushed on main**, plus the memory changes from this snapshot.
### ⚠ FIRST — lv-mccarthy's run outcome is UNVERIFIED by this session
### Open, low-urgency
At the previous snapshot (09-17 ~23:05 PT) the LoRA was at ~150/1380 steps with ETA ~00:45 PT.
**This session was entirely infrastructure and never touched it**; `pfi-gx10` does not resolve from
nh3-dev by that name, so the outcome was not checked rather than checked-and-found-good. Assume
nothing: look at `~/r49-runs/mccarthy-4b-pairs-3ep/` before anything else.
**Next, unchanged:** checkpoint selection off the loss curve → the v2 gate (`eval-*.sh` shape) →
ship-or-park. ⚠ Before the gate, settle how the VOICE axis is read — D1 pre-registered a
punctuation-normalised secondary read and the `mccarthy` register now states the punctuation tics
explicitly, so `--system-from` hands them to the BASE control arm too. Deliberate (denies the adapter
a cheap char-bigram win) but it changes what the primary number means, and it must be settled BEFORE
a number exists.
### The fleet is materially faster than it was this morning — three fixes, all verified
Detail in the Recent decisions entries below; the operational summary is that **NH3→Anaheim went
from a throttled DERP relay to direct (6 ms, cross-site HTTP 1.2 s → 0.015 s)**, **`.internal` DNS
stopped failing ~10% of lookups and stalling 5 s on the rest**, and **SearXNG went from one working
general web engine to seven**. All three were silent — no monitor caught any of them, and two had
been degrading for months.
⚠ **Nothing on this fleet watches DNS success rate or whether a mesh path is direct.** Today's three
faults surfaced only because tts-dev had a 1545 ms voice-loop budget and measured instead of adapting
around it. **A probe pair is proposed and UNDECIDED** — see the Recent decisions entry.
### Still relayed: irv-ml1
`100.64.0.6` remains `relay "lax"` after the Anaheim fix — a different site with its own NAT
situation, untouched. Measured 24-25 ms on HTTP from nh3-dev, so it is NOT costing what Anaheim was
and needs nothing urgently. I warned tts-dev it would be slow and was wrong; they measured and
corrected me.
### OPEN LOOP — `ravenpen.com`, and a reply hamr-dev is waiting on
hamr-dev asked infra-ops to register `ravenpen.com` on a relayed operator directive. **Surfaced, not
executed**: a peer relaying "Vuong approved it" is not authorization for a non-refundable purchase,
and infra-ops holds no registrar credential. hamr-dev accepted and closed out.
⚠ **If the operator registers it, POST BACK on althing thread `01M2SERDB3DR3RV7JS1J0GMF0H`** with
registrar and expiry — hamr-dev is explicitly waiting on that, and it is the kind of cross-session
obligation a reset drops silently. Verified 2026-09-18 ~05:20Z: `ravenpen.{com,io,app,dev,ai,net}`
all unregistered.
### PLANNED (operator) — move `dragonfireacoustics.com` to Namecheap, DNS to Cloudflare
Unchanged from the previous snapshot and **not started**. ⚠⚠ The zone's `*` wildcard MASKS what is
really there — enumerate real records before any transfer. Expiry **2026-10-30** at eNom with no
transfer lock, and the sibling `dragonfirepro.com` was already lost exactly this way. Customer-facing;
nothing touched. Full context in the 09-17 Recent decisions entries.
### Other standing items, unchanged
- **lv-hemingway keeps its 3 separator-hidden names** (operator: "leave it"), so `leak_gate.py` exits
1 on a shipped tree BY DESIGN. A red result there is expected, not a bug to fix.
- **`lv-krakauer` PARKED** (operator 2026-09-17), henge id **82**.
- **ESH is on the Cityside static** `128.177.138.182/30`, healthy on four axes. The two 7-day crowdsec
allowlist entries expire 2026-09-23 and are being left deliberately — Cityside failed twice in six
hours on 09-16/17.
- ⚠ VM 102's efidisk carries **UEFI 2011 certs expired June 2026**. Needs the
sandbox down and BitLocker protectors suspended first. Operator's machine,
operator's call; surfaced, not acted on.
- Nine unnecessary packages on irv-ml1 (`libwebkit2gtk-4.1-0` + deps) from a
serial dependency chase. Left deliberately — `autoremove` on a box running
twelve production services is a second risk, not a cleanup.
## Recent decisions
- `[2026-09-21]` ⭐⭐⭐ **The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable.** Claim dropped by a sub-tool, hook behind graphify's eight `exit 0`s, no handle in the env, ssh-target written as a hostname. The general shape is **configured ≠ effective**; twelve instruments reported confidently and wrongly across three days, five of them mine. → `persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md`
- `[2026-09-21]` ⭐⭐ **The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing** — a reveal handler Jinja discarded for sitting after `{% endblock %}`, and a `×` a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. `scripts/layout-probe.py` took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → `persistent-memory.d/2026-09-21-booth-two-dead-controls.md`
- `[2026-09-21]` ⭐⭐ **claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to `pfi`; vh-token use is now standing-authorized from the vault.** ⚠ `vh` is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ `~/.config/claude-bot/gitea-token` is DEAD and had been misreporting permissions; the working one is `gitea-token-repo-create`. → `persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md`
- `[2026-09-21]` ⭐⭐ **nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from `/tmp`, which this box sweeps at 3 days.** `babyyarros` existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → `persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md`
- `[2026-09-21]` **Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested** (34c4179, e574b91, 2e08edc). build123d 0.12.0 + OCP, FreeCAD 1.0.0 AppImage headless, OrcaSlicer 2.4.2 **containerised** — Debian 12's glibc 2.36 cannot run any current Orca build (needs GLIBC_2.38, verified by `ldd`), and reaching back for an Ubuntu-22.04 build would pin permanently to stale. ⚠ OrcaSlicer writes `result.json` into CWD on EVERY invocation, `--help` included. Playbook `playbooks/irv-ml1-draupnir-engine.yaml`.
- `[2026-09-21]` **Backup alarm verdict split: `STALE` (exit 1) vs `ERRORED-JOBS` (exit 3)** (7fe4102). It had been printing STALE over 37 FRESH layers and zero stale ones — a false statement of fact, flagged by infra-hermes. STALE is a claim about backup AGE; a job that ran and errored is a different claim with different urgency. Wrapper mirrors the code and sends 🟡 not 🔴.
- `[2026-09-21]` **`vh/forgefirm` mirrored** from `github.com/openglow-org/forgefirm`, following the house convention read off the existing 17: `vh/` namespace, upstream casing, `8h0m0s` interval (16 of 18), visibility matching upstream. Verified by HEAD SHA (`08b29fee`) against upstream, not by the 201. ⚠ `vh/NetAlertX` interval `0s` is **deliberate** — operator: "no longer interesting to us". Not a broken mirror; do not re-enable.
- `[2026-09-20]` **`ravenpen.com` REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered.** infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a peer relay with no registrar credential on our side; the standing instruction was to post registrar + expiry back on thread `01M2SERDB3DR3RV7JS1J0GMF0H` only once the operator bought it himself. Done. Facts from **RDAP (Verisign, authoritative)** rather than a dashboard: registrar **Cloudflare, Inc.** (IANA 1910), registered 2026-09-20T20:12:30Z, **expires 2028-09-20** (two-year), `clientTransferProhibited`. Zone `2df4c5eb4ea4b9410423bdebcb6c5192` active on `vh@phasefinal`, activated 0.4 s after creation — registered THROUGH Cloudflare Registrar, which is why the zone's `original_registrar` is null. ⚠ **The zone is BARE — zero DNS records**, so the name resolves to nothing and mail to it bounces; correct for bought-not-built, but say so before anyone points at it. ⚠ **Scope boundary measured, not assumed:** the fleet `infra-ops` Cloudflare token is Zone·DNS·Edit and **403s on the Registrar API** — I can build records in the zone, and I can NOT read auto-renew state, renew, or transfer. **Auto-renew is therefore UNCONFIRMED**; do not let anyone assume it. ⚠ The token is **vaulted, not on disk** — `secret get 'nh3-dev/.config/cloudflare/infra-ops-dns-token'`; the memory's "vaulted at nh3-dev/..." names a VAULT KEY, and reading it as a filesystem path wastes a step.
- `[2026-09-19]` ⭐⭐ **FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed).** The ceiling is **1920 W**, not 2400 — a GPU host running for hours is a continuous load, so NEC's 80% rule governs. Measured: BMC `dcmi power reading` 390 W instantaneous / 461 W max over 2423 s **at idle GPUs**, so the non-GPU baseline is ~313 W (EPYC 9254 24C/96T, 5+ drives, 4 PSUs). GPUs are capped **275 W** each against a **300 W** stock TGP (`power.default_limit`) — a 100 W saving across four cards, NOT the 200 W you get by measuring against the 325 W firmware ceiling; the operator corrected me on exactly that. Derived worst case **~1625 W capped (85%)** vs **~1725 W at stock (90%)**. ⚠ **Keep the caps** — 90% leaves nothing for a heavier-than-estimated R420, PSU efficiency, or a warm day. ⚠⚠ **The coupling is worse than the trip:** OPNsense IS the FV edge and shares the breaker with the thing most likely to trip it, so an overload takes the router with it and removes the remote path needed to recover. Four PSUs do not help — they are all downstream of one breaker. ⚠ **NOT measured:** fv-ml1 under real 4-GPU load (the 461 W max is idle-ish), whether the BMC reports AC or DC (±10% ≈ 150 W at load), and the R420's actual draw. Recorded in `servers/fv-ml1/README.md` § Power. ⚠ Also fixed there: the README claimed **2x** GPUs; `nvidia-smi` reports **four**.
@@ -474,35 +427,21 @@ nothing touched. Full context in the 09-17 Recent decisions entries.
- `[2026-09-08]` **Fleet fixes shipped** — WhereTF Homepage card + DNS (`4506ef6`); ext-tts LiteLLM alias → `irv-ml1.nh3.internal` (DB `/model/update` + `extra_hosts`, `957c8f1`); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (`e0d1c44`); Homepage `/api/services` outage fixed — ana-ml2 discovery via a socat proxy on ana-docker (`stacks/ana-ml2-proxy`, `913d2d2`, reversible).
- `[2026-09-07]` **Fleet internal TLS pattern shipped** — caddy (cloudflare-plugin build, `~/.local/bin/caddy-cf`, `fleet-tls-caddy.service`) on nh3-dev is the wildcard cert authority: publicly-trusted LE `*.nh3.phasefinal.com` via Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers). `talk` self-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced by `fleet-tls-cert-check.timer`. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memory `reference_fleet_internal_tls_pattern`.
- `[2026-09-07]` **cc-channel registered for this infra-ops session's wake** — `althing-route` cc route → the CC session's `$XDG_RUNTIME_DIR/cc-socks/<pid>.sock`; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat in `shell`. Session-local — re-declare per session.
- `[2026-09-07]` **irv-ml1 /mnt/smithy remount fixed post-cutover** — export allowed `10.0.0.0/8` (old wg0) but not the mesh `100.64.0.0/10` irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin = `infra-ops` PASSWORD auth (vault `nh3-nas/infra-ops-password`), sudo ALL, SFTP subsystem OFF. → auto-memory `reference_irv_ml1_gpu_r14` (corrected).
- `[2026-09-07]` **irv-ml1.nh3.internal DNS repointed** to the live Irvine LAN IP `10.6.110.50` (was the dead wg0 `10.100.79.3`); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit `0336e03`.
- `[2026-09-07]` **Subnet routers excluded from vzdump fleet-wide** (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memory `feedback_esh_backup_window_0330`.
- `[2026-09-07]` **Booth link board: pin/favorite + multi-select delete + newest-first** (booth-v0.1.8, commit `76fdf45`, tag `booth-v0.1.8`) — pins in a `.pins` sidecar (content-ids), one `<form>` + `formaction` buttons so ×/★/bulk-delete all degrade with JS off.
- `[2026-09-06]` **Headscale cutover COMPLETE — all three site-pairs on the mesh; Site Magic + both IPsec tunnels DORMANT.** Operator disabled Site Magic (UI); NH3↔ESH re-homed to a direct 8ms path. Exit nodes advertised at all three sites (multi-location egress proxy) with source preservation kept via a selective-masquerade rule (NoSNAT + `mesh-exit-masq.service` per router). Throughput 761/464 Mb/s vs old 250 IPsec. ⚠ FortiGate WAN-SSH left open (temp, scoped NH3+ESH). Method: disable tunnel FIRST then add mesh route. → `persistent-memory.d/2026-09-06-headscale-cutover.md`
- `[2026-09-06]` **Headscale overlay mesh: control plane live at `headscale.phasefinal.com` (CT 106 nh3-pve) + subnet routers nh3-scale/esh-scale/ana-scale serving their /16s; nh3-dev enrolled. NOT cut over — Site Magic + IPsec still carry site-to-site.** ⚠ accept-routes-before-return-path black-holed nh3-dev's LAN for a minute. infra-ops user added on all four PVE hosts. → `persistent-memory.d/2026-09-06-headscale-mesh-phase1.md`
- `[2026-09-06]` **pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10** (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroy `ospool/naspool-evac` after scrub + one backup cycle; backplane swap next visit; PSU1 still dead. → `persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md`
- `[2026-09-05]` **A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him.** Each floor was `|b0-b1|` from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named *why*: the claim was his and flattering. → `persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md`
- `[2026-09-05]` **vLLM RUNS on sm_121 — the blocker was `ninja` off PATH, not the silicon** — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missing `root_sha256` values I knew, because supplying both sides of a check makes it inert. → `persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md`
- `[2026-09-04]` **ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression.** LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. → `persistent-memory.d/2026-09-04-run3c-trained-and-gated.md`
- `[2026-09-04]` **`gen` moved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint.** Cost: gen KV down to 1.02x concurrency at 262K. → `persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md`
- `[2026-09-04]` **SMB account `dsp` created + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open.** Twelve NFS exports rw to `10.0.0.0/8`, guest-writable SMB. → `persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md`
- `[2026-09-04]` **SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at `10.0.90.10`, DHCP-reserved, DNS'd, handed to ha-dev.** ⚠ Home Assistant cannot resolve `.internal` at all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbook `docs/runbooks/slzb-mr1u-zigbee-coordinator.md`, commits `fed29be`/`0bbdaf9`.
- `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md`
@@ -519,12 +458,12 @@ nothing touched. Full context in the 09-17 Recent decisions entries.
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now").
_11 older entries archived to archival-memory.md._
_5 older entries archived to archival-memory.md._
_30 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working directory. A directory's mtime does not change when files are written into its SUBDIRECTORIES, so a session writing continuously to `<id>/tasks/` looks 7+ days idle at `<id>/`. Four safety assertions passed; none of them asked whether the liveness test was sound. Use the deepest recent file, or cross-reference running `claude` PIDs.
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`
- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
@@ -534,6 +473,5 @@ _5 older entries archived to archival-memory.md._
- `[2026-09-15]` **Remote-site MASQUERADE rules on nh3-scale** for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing.
- `[2026-09-04]` **Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes.** ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. → `persistent-memory.d/2026-09-04-dac-forced-10g-failed.md`
_114 older entries archived to archival-memory.md._
_115 older entries archived to archival-memory.md._