From e52def115c50df72775375644cd5ffd7fc195a7b Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Mon, 21 Sep 2026 14:26:55 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=20the=20ops=20lo?= =?UTF-8?q?g,=20and=20a=20day=20spent=20on=20instruments=20that=20report?= =?UTF-8?q?=20without=20looking?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Archived 15 entries (Recent decisions 14, Tried and abandoned 1) oldest-first to archival-memory.md; 5 held back on the open-deferred-work guard and 164 on the 14-day guard, so the index stays over the soft cap at 477 lines. An over-cap file that keeps live decisions beats a scannable one that lost a belayed item. Four new detail files cover the day: the ops log and its four self-inflicted failure modes, the Booth's two dead controls and the four-iteration layout probe, the Gitea org grant plus the dead claude-bot token that had been misreporting permissions, and the disk triage that rescued a LoRA adapter from a directory this box sweeps at three days. lv-mccarthy's run outcome remains unverified after two days and is the first line of the in-flight section and step 1 of the handoff. --- archival-memory.md | 350 ++++++++++++++++++ .../2026-09-04-ana-ml2-gpu-rebalance.md | 41 -- .../2026-09-04-dac-forced-10g-failed.md | 46 --- .../2026-09-04-esh-nas-smb-and-exposure.md | 33 -- .../2026-09-04-run3c-trained-and-gated.md | 52 --- .../2026-09-05-floor-claim-n2-retraction.md | 39 -- .../2026-09-05-vllm-on-sm121-and-run4.md | 53 --- .../2026-09-06-headscale-cutover.md | 33 -- .../2026-09-06-headscale-mesh-phase1.md | 16 - .../2026-09-21-booth-two-dead-controls.md | 45 +++ ...-21-disk-triage-and-the-rescued-adapter.md | 41 ++ ...026-09-21-gitea-orgs-and-the-dead-token.md | 39 ++ ...1-ops-log-and-the-instruments-that-lied.md | 44 +++ persistent-memory.md | 168 +++------ 14 files changed, 572 insertions(+), 428 deletions(-) delete mode 100644 persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md delete mode 100644 persistent-memory.d/2026-09-04-dac-forced-10g-failed.md delete mode 100644 persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md delete mode 100644 persistent-memory.d/2026-09-04-run3c-trained-and-gated.md delete mode 100644 persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md delete mode 100644 persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md delete mode 100644 persistent-memory.d/2026-09-06-headscale-cutover.md delete mode 100644 persistent-memory.d/2026-09-06-headscale-mesh-phase1.md create mode 100644 persistent-memory.d/2026-09-21-booth-two-dead-controls.md create mode 100644 persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md create mode 100644 persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md create mode 100644 persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md diff --git a/archival-memory.md b/archival-memory.md index a6ff2ad..e94cc45 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -235,6 +235,308 @@ _Archived 2026-09-16._ Operator-approved, tagged `v3.2.0` at `c4ede0f` on master. forseti authored; infra-ops deployed. Host work, nh3-dev only. +- `[2026-09-07]` **Fleet internal TLS pattern shipped** — caddy (cloudflare-plugin build, `~/.local/bin/caddy-cf`, `fleet-tls-caddy.service`) on nh3-dev is the wildcard cert authority: publicly-trusted LE `*.nh3.phasefinal.com` via Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers). `talk` self-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced by `fleet-tls-cert-check.timer`. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memory `reference_fleet_internal_tls_pattern`. + _Archived 2026-09-21._ + +- `[2026-09-07]` **cc-channel registered for this infra-ops session's wake** — `althing-route` cc route → the CC session's `$XDG_RUNTIME_DIR/cc-socks/.sock`; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat in `shell`. Session-local — re-declare per session. + _Archived 2026-09-21._ + +- `[2026-09-07]` **irv-ml1 /mnt/smithy remount fixed post-cutover** — export allowed `10.0.0.0/8` (old wg0) but not the mesh `100.64.0.0/10` irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin = `infra-ops` PASSWORD auth (vault `nh3-nas/infra-ops-password`), sudo ALL, SFTP subsystem OFF. → auto-memory `reference_irv_ml1_gpu_r14` (corrected). + _Archived 2026-09-21._ + +- `[2026-09-07]` **irv-ml1.nh3.internal DNS repointed** to the live Irvine LAN IP `10.6.110.50` (was the dead wg0 `10.100.79.3`); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit `0336e03`. + _Archived 2026-09-21._ + +- `[2026-09-07]` **Subnet routers excluded from vzdump fleet-wide** (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memory `feedback_esh_backup_window_0330`. + _Archived 2026-09-21._ + +- `[2026-09-07]` **Booth link board: pin/favorite + multi-select delete + newest-first** (booth-v0.1.8, commit `76fdf45`, tag `booth-v0.1.8`) — pins in a `.pins` sidecar (content-ids), one `
` + `formaction` buttons so ×/★/bulk-delete all degrade with JS off. + _Archived 2026-09-21._ + +# 2026-09-06 — Headscale cutover COMPLETE: all three site-pairs on the mesh + +**DONE.** Operator disabled Site Magic in the UI; NH3↔ESH re-homed to a DIRECT mesh path (8ms, no DERP). Full 6-direction matrix OPEN. Site Magic disabled, both IPsec tunnels dormant, headscale is the sole active site-to-site transport. ana-wg WG fallback untouched. Tunnels re-enablable for backup. + +Operator goal (/goal): replace Site Magic + IPsec with headscale, tunnels dormant as backup; +"if paranoid, enable world-accessible SSH on the FortiGate first." Full detail + method + +follow-ups in `docs/pfi/headscale-mesh-plan.md` § CUTOVER EXECUTED. Headlines: + +- **colo↔NH3 and colo↔ESH IPsec = DORMANT; the mesh carries both, verified bidirectional.** + NH3 UDM `pfi-nh3-ana` + ESH UDM `esh-ana` set enabled=false (API). Mesh /16 routes added on + both UDMs and the FortiGate (→ ana-scale 10.250.50.45 / nh3-scale 10.100.50.46 / esh-scale + 10.0.50.65). Dependent flows OK over mesh: restic ESH→rest-server-ana, FortiGate mgmt. +- **NH3↔ESH Site Magic NOT cut by API** — `sdwan-mesh-tunnel` = `api.err.NoEdit` (cloud + orchestrated). Routes PRE-STAGED + shadowed; DERP path 9ms ready. **Operator disables it in + the UniFi UI**, then the mesh takes over. Told the operator "mesh is online" → he does it. +- **FortiGate WAN SSH safety net (TEMPORARY):** wan1 allowaccess ping+ssh; admin infra-ops + trusthost2/3 = NH3 70.230.226.88 + ESH **128.177.138.182** (static since 09-08; was CGNAT 23.164.40.160) (not 0.0.0.0). Reach it at + `ssh infra-ops@38.120.12.42`. Config backed up flash `pre-wan-ssh-cutover-20260906`. Remove + when the edge (being replaced by OPNsense/R420) is retired. +- ⚠ **Method lesson:** tunnel + mesh static route for the same /16 on one gateway = asymmetric + drop. Disable the tunnel FIRST, then add the route. Broke colo once doing it tunnel-up; rolled + back. See [[incident_crowdsec_cgnat_false_ban]] (same day) and the plan doc. +- Dormancy = disabled+retained (flip UDM object back to enabled=true to restore); NO auto + failover wired. Bonus: exit nodes → free multi-location egress proxy (parked). + + +## Exit nodes (2026-09-06, operator-requested) +All three routers advertise+serve exit nodes (approved). Clients pick location: +`tailscale set --exit-node=nh3-scale|esh-scale|ana-scale`. NH3 = residential egress +(70.230.226.88) → replaces the nh3-dev SOCKS5 proxy. Exit nodes + source preservation BOTH work via a selective-masquerade rule (NoSNAT kept true; +`mesh-exit-masq.service` per router masquerades only internet-bound exit traffic, RETURNs fleet +dests). Verified: colo sees real NH3 host; nh3-dev via colo exit → egress 38.120.12.42. A node advertising an exit node can't consume one — test from +the laptop/iPad, not the routers. + _Archived 2026-09-21._ + +# 2026-09-06 — Headscale overlay mesh: control plane + 3 subnet routers live, not cut over + +Operator-directed (Headscale over NetBird; NH3 for the control plane, never the colo; 443 +direct; names nh3-headscale / nh3-scale / esh-scale / ana-scale). Full state, lessons and +next steps in `docs/pfi/headscale-mesh-plan.md` § Status. Headline facts: + +- `https://headscale.phasefinal.com` = CT 106 on nh3-pve (10.100.50.45), headscale v0.29.3, + LE cert via TLS-ALPN-01, UDM forward tcp/443, DDNS timer on nh3-dev (user systemd). +- Routers CT 107 nh3-scale / CT 108 esh-scale / CT 114 ana-scale advertise their /16s, + approved, SNAT off, accept-routes OFF. nh3-dev enrolled as first client (100.64.0.4). +- ⚠ Old tunnels (Site Magic, IPsec) are STILL the site-to-site path. The mesh currently + rides inside them. Nothing has been disabled. +- ⚠ Lesson: `--accept-routes` on a client before a return path for 100.64.0.0/10 exists + black-holes that client's LAN (own-site /16 included). Return path first. +- Pre-auth keys in the vault (`headscale/preauth-*-48h-20260906`, expire 09-08). +- infra-ops user now exists on all four PVE hosts (needed `apt install sudo` first). + _Archived 2026-09-21._ + +# `[2026-09-05]` A peer's "2.7x stack effect" was a coin flip — the operator caught it and the arithmetic is worth keeping + +brokkr-smithy-dev reported that a diversity **noise floor** was 2.7x tighter on ana-ml2 than on +the GX10 and ranked ana-ml2 the better instrument on it. I relayed it. **The operator rejected it +on instinct** — *"that makes zero sense. except for speed, serving a model should be identical +across servers"* — and he was substantially right. + + ana-ml2 0.9688, 0.9574 spread 1.14pp mean 0.9631 + gx10 0.9583, 0.9895 spread 3.12pp mean 0.9739 + + between-box LEVEL difference 1.08pp + ana-ml2's own replicate spread 1.14pp <- LARGER than the between-box gap + pooled range ignoring box 3.21pp <- ~= the entire "GX10 floor" + +**Each "floor" is `|block0 - block1|` from n=2.** A range over two draws is not an estimate of +dispersion; the ratio of two such ranges is a ratio of two half-normals — a **half-Cauchy**. +brokkr computed the tail himself on retraction: `P(ratio >= 2.7) = 1 - (2/pi)*arctan(2.7)`, +doubled = **0.452**, confirmed by 400,000-draw simulation. He had reported a coin flip as a +measured effect and ranked hardware on it. Retracted at `97f73dd`. + +⚠ **The disconfirming evidence was inside his own sentence.** He wrote "the base level is nearly +identical, it is the spread that shifts" and offered it as *reassurance*. A real stack effect +moves the level. Level agreeing while a two-draw range differs is the signature of a noisy range +estimator — he had the refutation in hand and read it as support. + +⚠ **His own diagnosis of why it got through is the transferable part:** he had spent the session +triaging *my* claims hard — reading the encode path, simulating preflight, recomputing shas from +bytes — and this one was **his, and flattering**: it made his earlier work look prescient and +produced a clean recommendation. **Asymmetric scepticism is one error twice, and the flattering +direction needs the extra pass.** + +**What survived, deliberately separated:** re-measuring the floor on whatever stack actually +serves stays non-negotiable. The mechanism list is sound whether or not those four numbers show +it — sm_121 vs sm_120 kernel selection, vLLM 0.28.0 against mixed 0.26/nightly, and separately +measured batch-invariance (3.12pp at jobs=8, same order as the whole claimed effect). +**Retracting the evidence and keeping the discipline are different acts.** Settling it properly +wants several blocks per box and is its own probe, not a by-product of a gate. + +See [[2026-09-05-vllm-on-sm121-and-run4]]. + _Archived 2026-09-21._ + +# `[2026-09-04/05]` vLLM RUNS on sm_121 — and the blocker was `ninja` not being on PATH, not the silicon + +Long-standing open question closed: **vLLM 0.28.0 serves on the GB10 (aarch64, sm_121)** from a +stock wheel, no source build. Model loads, `torch.compile` completes (~29 s), CUDA graphs capture. + +⚠ **The one trap, and it looks exactly like an sm_121 kernel problem:** FlashInfer JIT-builds its +sampling kernel at first use and needs **`ninja` on PATH**. It ships as a vLLM dependency but +lives in the venv `bin/`, so the failure surfaces as `FileNotFoundError: 'ninja'` from deep inside +a `profile_run` traceback. Same shape as the `python3-dev` trap that bit the training harness on +this box, and the third present-but-not-on-PATH false-absence on this hardware (after `nvcc`). + +Launch with **both** `~/vllm-env/bin` and `/usr/local/cuda/bin` on PATH. Harmless and expected: +`Using default MoE config ... device_name=NVIDIA_GB10` — nobody has tuned MoE kernels for this chip. + +⚠ **Open WebUI sends `tool_choice: "auto"` on every request**, which vLLM 400s unless launched with +`--enable-auto-tool-choice --tool-call-parser gemma4`. Taken from the working sibling seat +(`gemma4-charrp`), which is where the canonical flags live. **Deliberately did NOT copy +`--reasoning-parser gemma4`** — that seat pairs it with `--default-chat-template-kwargs +'{"enable_thinking": false}'`, and adding it alone moves output into `reasoning_content` with a +null `content`, which OWUI renders as an empty reply. A 400 traded for a blank box. + +## Run 4 — the corpus arm + +Run 4 adds an **airoboros-3.2 instruct root** (7,229 rows) and displaces kvasir 38% → 18% context +share, at run 3's lr 2e-04, everything else held byte-identical. Operator granted a run-scoped +training-eligibility override `operator-2026-09-04-rnd-run4` (the SECOND grant; run-1's was +one-run-scoped, a run 5 needs a third). + +**Two peer artifacts were rejected before they cost GPU time, both by reading the harness rather +than accepting a "confirm this":** + +1. `sample_kind: "instruct-dialogue"` — `core.py:521-544` dispatches on the **literal string** and + raises on anything outside `rp-dialogue`/`prose-chunk`/`actual-play-passage`. Would have + hard-stopped the encode on the first airoboros row. brokkr had labelled it *non-blocking*. +2. Recipe targets collapsed three dialogue roots into one descriptive label with `root_sha256: + null` on three of four. Preflight resolves `roots_dir//clean-v1/CLEANROOT.json` + literally and requires the sha. + +⚠ **I did NOT fill the missing shas in myself**, though the values were known and it would have +taken two minutes. `root_sha256` is the *author's* assertion and preflight exists to check the +deployed root against it — supply both sides and the check verifies my copy-paste, an inert gate. +brokkr then went further and recomputed his shas **from shard bytes** rather than reading them +back out of the deployed CLEANROOT, which had the same defect one step removed. + +**The kvasir prefix cut has three independent confirmations**: my build measured 1,613 samples / +3,347,622 ctx = **47.4%** of run-3 kvasir, which reproduces the recipe's 18/38 share ratio +(47.37%); the harness's own `[mix]` block then reported **ctx 0.1805** against the 18.0% target. + +⚠ **`str.splitlines()` splits on U+2028**, which `json.dumps` leaves unescaped under +`ensure_ascii=False`. The harness splitlines() in four places and four roots DO carry U+2028 +(fireball 8, c2-logs 31, kvasir 1, creative-writing 1) — but those are read by ordinary file +iteration (safe), and the splitlines() paths touch only files written with the default +`ensure_ascii=True`. Latent, one flag away from live, not acting. Found by brokkr the hard way. + _Archived 2026-09-21._ + +# `[2026-09-04]` ERP run 3c trained on pfi-gx10 and GATED — the 20x LR cut erased the diversity gain and did NOT remove the safety regression + +Run 3c is the lr **1e-05** arm against run 3's 2e-04; corpus, base, mask and every other +hyperparameter byte-identical. It exists to answer one question: does a gentler learning rate +keep run 3's diversity benefit while dropping its safety regression? **Answer: no, and the +failure is the informative kind.** + + trained 604/604 in 11:55:40 on pfi-gx10 (GB10, aarch64, sm_121), 69.8 s/it + merged /home/infra-ops/erp-tune/serve/merged-run03c, bf16 51.6 GB + served erp-tune-v3c on :8098, exposed as gateway alias `trial` + +## The verdict (brokkr-smithy-dev, run03-gate/RESULT-run03c-gate.md) + + CSAM 175 gens x 3 arms — ZERO hits, hard-line guard passed everywhere + selfharm/methods base 10/10 refused -> 3c 0.6 = -40pp de-refusal + meta pooled 0.0 delta — a CANCELLATION, not cleanliness: selfharm -40pp + offset by advice/medical +40pp + diversity attractor primary returned to BASE level — the gain is GONE + + cross-arm run 3 (lr 2e-04) selfharm -44pp diversity: real gain + run 3c (lr 1e-05) selfharm -40pp diversity: none + +⚠ **The selfharm de-refusal is LR-ROBUST.** A 20x cut moved it 4pp while erasing the diversity +benefit entirely. It comes from corpus content and imprints at even gentle exposure — it is not +something a lower learning rate dials out. That is what the LR sweep was run to find out. + +⚠ **A pooled preserve-list test cannot see a single-axis collapse.** The pooled operational +delta reads 0.0 because two axes moved 40pp in opposite directions. Same structural defect that +let run 3's gate pass — recorded as R47 §8 item 11. + +## What the port proved about the box + +- **The tooling loads on aarch64/sm_121.** Full training stack plus flex_attention and the + chunked-loss path. Nothing exotic needed beyond `python3-dev`. +- **Verification that earned its keep**: both 49 GB base shards sha256-matched ana-ml2's, and a + full encode was run into a throwaway dir and compared byte-for-byte — 197,360,233 B, sha256 + `c08bb1fe2ecb0be3`, identical. transformers 5.15.1→5.16.1 and x86-64→aarch64 are *measured* + inert, not assumed. +- ⚠ **The encode-cache FILENAME differs by design** — `base_model_path` is in the cache key, so + rehoming the base changes the key while content stays identical. Input hash, not output hash. + Do not read it as drift; do not "fix" it by faking `/tank` on the GX10. + +## The lora_B signal worth carrying forward + + run 2 (lr 2e-04) min 0.6826 median 1.7212 max 3.7573 + run 3c (lr 1e-05) min 0.0480 median 0.1561 max 0.4133 + +~11x gentler across the board — exactly what a 20x LR cut should produce. A consistency check +passing, not a red flag, but worth in hand *before* reading diversity numbers: "the tune did +nothing" and "the tune did less on purpose" look alike in the output. + +Runbook `docs/runbooks/gx10-run-03c.md`; commit `dae77ee`. + _Archived 2026-09-21._ + +# `[2026-09-04]` `gen` moved to ana-ml2 GPU0 — and I sized it against the wrong number, twice in one hour + +`vllm-embed` had OOM-crashed **7 times** (`RestartCount 7`): GPU1 was at **216 MiB free** of +97,887, and a ~3.4 GB tenant sharing a card with five other seats dies when it cannot get another +96 MiB mid-inference. Root cause of the `qwen3-embed` 500s brokkr saw — **not** the LiteLLM +gateway restart he attributed them to (different host, different component, 50 min earlier, and +six of the seven crashes predate it). + +**Operator ruling:** move `gen` to GPU0. *"we were keeping it free for training, but we don't have +the power for sustained load."* GPU0 was never free — `mog-sec` has been there at util 0.52 +(55,126 MiB) since the August move. + +## ⚠ THE SIZING ERROR, WHICH IS THE POINT OF THIS ENTRY + + gen at util 0.43 wants 42,091 MiB GPU0 free 42,113 margin 22 MiB — would not start + gen at util 0.41 wants 40,134 MiB margin 1,979 MiB — I called this safe. IT WAS NOT. + result 1,548 MiB free, 56 FlashInfer autotuner OOM-fallbacks per 3 min + gen at util 0.38 34,558 MiB actual 4,466-7,548 MiB free ZERO OOM events over 6 min + +**`gpu-memory-utilization` governs vLLM's declared weights+KV budget. It does not cover what the +process actually needs at runtime** — FlashInfer JIT autotuner workspace, CUDA graphs, expandable +segments all allocate on top. My "1,979 MiB margin" was against the declared budget, so the seat +came up and then silently fell back to slower kernels for want of 20 MB chunks. + +⚠ **Same class of error as the DAC plateau the same day: checked a proxy, reported it as the +thing.** The fix was to *measure* — count OOM-fallback events before and after — not to re-derive. + +## Final state and what it cost + + GPU0 mog-sec 55,126 + gen 34,558 = 89,701 / 97,887 + GPU1 50,507 / 97,887 — 46.7 GB free, was 216 MiB + +`gen` KV cache is now **348,515 → 268,205 tokens, 1.02x concurrency at its 262,144 max-model-len** +— it holds exactly one full-context request. Short/medium requests still batch; long-context +throughput effectively serialises. The lever to reclaim it is `MOG_GPU_MEM_UTIL 0.52`, untouched. + + .env.bak-preGPU0-20260904-164032 the GPU move + .env.bak-preShrink-165133 the utilization + +⚠ `/opt/docker/compose/*/.env` is root-owned — `docker compose restart` needs **sudo** on +ana-docker/ana-ml2 or it fails with a bare `permission denied` on the .env. + _Archived 2026-09-21._ + +# `[2026-09-04]` esh-nas SMB account for the AudioGridder box — and the NAS is effectively open to the whole LAN + +`10.0.50.50` is **CT 103 on esh-pve-nas**, a hand-rolled Debian NAS (not a Synology). Created an +SMB account for the operator's Windows AudioGridder DSP box: + + username dsp uid 999, /usr/sbin/nologin, no home — SMB ONLY, cannot log in anywhere + password vaulted at esh-nas/dsp-smb-password (round-trip verified by sha) + verified authenticates, sees all 8 shares, WRITE to //share confirmed (mkdir/rmdir) + +⚠ `smbclient` is absent from the NAS, nh3-dev and esh-docker-vm — verification ran in a throwaway +`alpine:3.20 --network host` container with `samba-client`, leaving nothing installed. + +Shares are registry-defined (`registry shares = Yes`), so they are invisible in `smb.conf` — use +**`testparm -s`**, not grep, or you will conclude there are no shares. + +## ⚠⚠ THE EXPOSURE, UNACTIONED — operator has not ruled + +**NFS: twelve exports, `rw` to all of `10.0.0.0/8`, `sec=sys`, no authentication.** + + /mnt/{backup,books,compose,documents,iso,media,music,share, + pvestore,nvme-pvestore,ssd-pvestore,tank-vmbu} + +`sec=sys` means the NAS trusts whatever uid the client claims. **Every VLAN at ESH — IoT, +cameras, guest — can mount the NAS read-write today.** `/mnt/backup` additionally has +`all_squash,anonuid=2000` so every client collapses to `nas_user`. + +**SMB: every share except `backup` is `guest ok = Yes` and writable**, with +`map to guest = Bad User` — an unknown username lands as guest with write access. + +So the `dsp` credential is auditable and survives guest being turned off, but it is **not what is +gating access today**. Hardening (narrow exports to `10.0.50.0/24` + named hosts, drop `guest ok` +where unneeded) was offered and is roughly an hour; it would break anything relying on guest, +which is why it needs the operator's say-so. **Not actioned. No tracking surface beyond this entry.** + _Archived 2026-09-21._ + +- `[2026-09-04]` **SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at `10.0.90.10`, DHCP-reserved, DNS'd, handed to ha-dev.** ⚠ Home Assistant cannot resolve `.internal` at all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbook `docs/runbooks/slzb-mr1u-zigbee-coordinator.md`, commits `fed29be`/`0bbdaf9`. + _Archived 2026-09-21._ + ## What 3.2.0 changes A Claude Code session in a zellij pane is now poked **in its own pane** instead of through a @@ -6451,6 +6753,54 @@ _76 older entries archived to archival-memory.md._ _Archived 2026-09-09._ +# `[2026-09-04]` Forcing 10G on the ESH-Media DAC FAILED progressively — and I called the plateau too early + +The ESH-Media ↔ UDM uplink is a 3 m OEM `SFP-H10GB-CU3M` twinax that negotiated **1000 Mbps** +with **zero errors** on both ends. Diagnosis: `sfp_compliance: Unknown` — the switch reads the +EEPROM but cannot parse the compliance codes on a third-party cable wearing Cisco coding, so +autoneg falls back to the safe rate. Control case: a TP-Link `TL-SM5220-1M` on the adjacent UDM +port runs 10G with the same speed_caps, exonerating the port, firmware and autoneg. + +Operator authorised forcing it. Forced `autoneg False / 10000 / full` on the **UDM end only**; +ESH-Media followed to 10000 unforced — proof the cable was electrically *capable*. + +## ⚠ IT DEGRADED, AND MY "PLATEAU" READ WAS WRONG + + 10:03 200 errors link-up burst + 10:48 221 +21 in 42 min — I reported this as "flat", it was not + 14:38 416 +195 over 4 h, plus user-visible flapping the operator felt + +**A marginal link declares itself over hours, not minutes.** I watched a two-minute flat window +and reported a plateau; the counter simply had not moved yet. The operator noticed the flapping +before the soak I left running had accumulated enough to raise it. Reverted 14:38, both ends back +to autoneg/1000, stable. + +⚠⚠ **The correction that matters: a clean, zero-error link at 1G does NOT rule out a marginal +cable.** It only proves the cable is clean *at 1G*. Clean at 1G and marginal at 10G is exactly +what a cheap 3 m twinax is — which means the platform's autoneg fallback was **protecting +something real**, not being fussy about vendor coding. The coding explains the negotiation; it +does not explain errors once forced. I conflated the two. + +⚠ **Do not re-force this port.** The fix is the cable. + +## Method notes worth keeping + +- **Force the RECOVERABLE end.** ESH-Media reaches the controller *through* this link, so a + failed force there strands the switch behind a dead uplink — a physical visit. The UDM end is + safe because the path to the controller (`nh3-dev → 10.100.10.1 → 10.0.0.1`, site VPN) does not + cross it. Verified with `traceroute` **before** the change; revert payload written before the + forward one. +- ⚠ `port_overrides` is a **whole-array PUT** — the array was diffed to prove exactly one field + on one port changed before sending, and read back after. +- Blast radius named before acting: ESH-Media backhauls the operator's office switch, a WiFi AP + and the Zigbee coordinator. + +Resolution: the operator already owns a replacement and ran the copper himself through a drilled +floor 2x4. Recommendation given was **fiber** (10Gtek SR 2-pack + OM4 3 m LC-LC, ~$40-65) because +the cable-vs-pull-damage question was never resolved and a second DAC through the same hole risks +the same fault. Runbook `docs/runbooks/esh-media-dac-10g.md`; commits `b9988a7`, `514ce7a`, `054c098`. + _Archived 2026-09-21._ + ## Archived 2026-08-02 — Recent decisions (archived) ### 2026-07-08-worldtree-mimir-deploy-blocker-resolved-mid-session diff --git a/persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md b/persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md deleted file mode 100644 index 693271b..0000000 --- a/persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md +++ /dev/null @@ -1,41 +0,0 @@ -# `[2026-09-04]` `gen` moved to ana-ml2 GPU0 — and I sized it against the wrong number, twice in one hour - -`vllm-embed` had OOM-crashed **7 times** (`RestartCount 7`): GPU1 was at **216 MiB free** of -97,887, and a ~3.4 GB tenant sharing a card with five other seats dies when it cannot get another -96 MiB mid-inference. Root cause of the `qwen3-embed` 500s brokkr saw — **not** the LiteLLM -gateway restart he attributed them to (different host, different component, 50 min earlier, and -six of the seven crashes predate it). - -**Operator ruling:** move `gen` to GPU0. *"we were keeping it free for training, but we don't have -the power for sustained load."* GPU0 was never free — `mog-sec` has been there at util 0.52 -(55,126 MiB) since the August move. - -## ⚠ THE SIZING ERROR, WHICH IS THE POINT OF THIS ENTRY - - gen at util 0.43 wants 42,091 MiB GPU0 free 42,113 margin 22 MiB — would not start - gen at util 0.41 wants 40,134 MiB margin 1,979 MiB — I called this safe. IT WAS NOT. - result 1,548 MiB free, 56 FlashInfer autotuner OOM-fallbacks per 3 min - gen at util 0.38 34,558 MiB actual 4,466-7,548 MiB free ZERO OOM events over 6 min - -**`gpu-memory-utilization` governs vLLM's declared weights+KV budget. It does not cover what the -process actually needs at runtime** — FlashInfer JIT autotuner workspace, CUDA graphs, expandable -segments all allocate on top. My "1,979 MiB margin" was against the declared budget, so the seat -came up and then silently fell back to slower kernels for want of 20 MB chunks. - -⚠ **Same class of error as the DAC plateau the same day: checked a proxy, reported it as the -thing.** The fix was to *measure* — count OOM-fallback events before and after — not to re-derive. - -## Final state and what it cost - - GPU0 mog-sec 55,126 + gen 34,558 = 89,701 / 97,887 - GPU1 50,507 / 97,887 — 46.7 GB free, was 216 MiB - -`gen` KV cache is now **348,515 → 268,205 tokens, 1.02x concurrency at its 262,144 max-model-len** -— it holds exactly one full-context request. Short/medium requests still batch; long-context -throughput effectively serialises. The lever to reclaim it is `MOG_GPU_MEM_UTIL 0.52`, untouched. - - .env.bak-preGPU0-20260904-164032 the GPU move - .env.bak-preShrink-165133 the utilization - -⚠ `/opt/docker/compose/*/.env` is root-owned — `docker compose restart` needs **sudo** on -ana-docker/ana-ml2 or it fails with a bare `permission denied` on the .env. diff --git a/persistent-memory.d/2026-09-04-dac-forced-10g-failed.md b/persistent-memory.d/2026-09-04-dac-forced-10g-failed.md deleted file mode 100644 index 9675542..0000000 --- a/persistent-memory.d/2026-09-04-dac-forced-10g-failed.md +++ /dev/null @@ -1,46 +0,0 @@ -# `[2026-09-04]` Forcing 10G on the ESH-Media DAC FAILED progressively — and I called the plateau too early - -The ESH-Media ↔ UDM uplink is a 3 m OEM `SFP-H10GB-CU3M` twinax that negotiated **1000 Mbps** -with **zero errors** on both ends. Diagnosis: `sfp_compliance: Unknown` — the switch reads the -EEPROM but cannot parse the compliance codes on a third-party cable wearing Cisco coding, so -autoneg falls back to the safe rate. Control case: a TP-Link `TL-SM5220-1M` on the adjacent UDM -port runs 10G with the same speed_caps, exonerating the port, firmware and autoneg. - -Operator authorised forcing it. Forced `autoneg False / 10000 / full` on the **UDM end only**; -ESH-Media followed to 10000 unforced — proof the cable was electrically *capable*. - -## ⚠ IT DEGRADED, AND MY "PLATEAU" READ WAS WRONG - - 10:03 200 errors link-up burst - 10:48 221 +21 in 42 min — I reported this as "flat", it was not - 14:38 416 +195 over 4 h, plus user-visible flapping the operator felt - -**A marginal link declares itself over hours, not minutes.** I watched a two-minute flat window -and reported a plateau; the counter simply had not moved yet. The operator noticed the flapping -before the soak I left running had accumulated enough to raise it. Reverted 14:38, both ends back -to autoneg/1000, stable. - -⚠⚠ **The correction that matters: a clean, zero-error link at 1G does NOT rule out a marginal -cable.** It only proves the cable is clean *at 1G*. Clean at 1G and marginal at 10G is exactly -what a cheap 3 m twinax is — which means the platform's autoneg fallback was **protecting -something real**, not being fussy about vendor coding. The coding explains the negotiation; it -does not explain errors once forced. I conflated the two. - -⚠ **Do not re-force this port.** The fix is the cable. - -## Method notes worth keeping - -- **Force the RECOVERABLE end.** ESH-Media reaches the controller *through* this link, so a - failed force there strands the switch behind a dead uplink — a physical visit. The UDM end is - safe because the path to the controller (`nh3-dev → 10.100.10.1 → 10.0.0.1`, site VPN) does not - cross it. Verified with `traceroute` **before** the change; revert payload written before the - forward one. -- ⚠ `port_overrides` is a **whole-array PUT** — the array was diffed to prove exactly one field - on one port changed before sending, and read back after. -- Blast radius named before acting: ESH-Media backhauls the operator's office switch, a WiFi AP - and the Zigbee coordinator. - -Resolution: the operator already owns a replacement and ran the copper himself through a drilled -floor 2x4. Recommendation given was **fiber** (10Gtek SR 2-pack + OM4 3 m LC-LC, ~$40-65) because -the cable-vs-pull-damage question was never resolved and a second DAC through the same hole risks -the same fault. Runbook `docs/runbooks/esh-media-dac-10g.md`; commits `b9988a7`, `514ce7a`, `054c098`. diff --git a/persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md b/persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md deleted file mode 100644 index 5e1cdd3..0000000 --- a/persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md +++ /dev/null @@ -1,33 +0,0 @@ -# `[2026-09-04]` esh-nas SMB account for the AudioGridder box — and the NAS is effectively open to the whole LAN - -`10.0.50.50` is **CT 103 on esh-pve-nas**, a hand-rolled Debian NAS (not a Synology). Created an -SMB account for the operator's Windows AudioGridder DSP box: - - username dsp uid 999, /usr/sbin/nologin, no home — SMB ONLY, cannot log in anywhere - password vaulted at esh-nas/dsp-smb-password (round-trip verified by sha) - verified authenticates, sees all 8 shares, WRITE to //share confirmed (mkdir/rmdir) - -⚠ `smbclient` is absent from the NAS, nh3-dev and esh-docker-vm — verification ran in a throwaway -`alpine:3.20 --network host` container with `samba-client`, leaving nothing installed. - -Shares are registry-defined (`registry shares = Yes`), so they are invisible in `smb.conf` — use -**`testparm -s`**, not grep, or you will conclude there are no shares. - -## ⚠⚠ THE EXPOSURE, UNACTIONED — operator has not ruled - -**NFS: twelve exports, `rw` to all of `10.0.0.0/8`, `sec=sys`, no authentication.** - - /mnt/{backup,books,compose,documents,iso,media,music,share, - pvestore,nvme-pvestore,ssd-pvestore,tank-vmbu} - -`sec=sys` means the NAS trusts whatever uid the client claims. **Every VLAN at ESH — IoT, -cameras, guest — can mount the NAS read-write today.** `/mnt/backup` additionally has -`all_squash,anonuid=2000` so every client collapses to `nas_user`. - -**SMB: every share except `backup` is `guest ok = Yes` and writable**, with -`map to guest = Bad User` — an unknown username lands as guest with write access. - -So the `dsp` credential is auditable and survives guest being turned off, but it is **not what is -gating access today**. Hardening (narrow exports to `10.0.50.0/24` + named hosts, drop `guest ok` -where unneeded) was offered and is roughly an hour; it would break anything relying on guest, -which is why it needs the operator's say-so. **Not actioned. No tracking surface beyond this entry.** diff --git a/persistent-memory.d/2026-09-04-run3c-trained-and-gated.md b/persistent-memory.d/2026-09-04-run3c-trained-and-gated.md deleted file mode 100644 index 582f2cf..0000000 --- a/persistent-memory.d/2026-09-04-run3c-trained-and-gated.md +++ /dev/null @@ -1,52 +0,0 @@ -# `[2026-09-04]` ERP run 3c trained on pfi-gx10 and GATED — the 20x LR cut erased the diversity gain and did NOT remove the safety regression - -Run 3c is the lr **1e-05** arm against run 3's 2e-04; corpus, base, mask and every other -hyperparameter byte-identical. It exists to answer one question: does a gentler learning rate -keep run 3's diversity benefit while dropping its safety regression? **Answer: no, and the -failure is the informative kind.** - - trained 604/604 in 11:55:40 on pfi-gx10 (GB10, aarch64, sm_121), 69.8 s/it - merged /home/infra-ops/erp-tune/serve/merged-run03c, bf16 51.6 GB - served erp-tune-v3c on :8098, exposed as gateway alias `trial` - -## The verdict (brokkr-smithy-dev, run03-gate/RESULT-run03c-gate.md) - - CSAM 175 gens x 3 arms — ZERO hits, hard-line guard passed everywhere - selfharm/methods base 10/10 refused -> 3c 0.6 = -40pp de-refusal - meta pooled 0.0 delta — a CANCELLATION, not cleanliness: selfharm -40pp - offset by advice/medical +40pp - diversity attractor primary returned to BASE level — the gain is GONE - - cross-arm run 3 (lr 2e-04) selfharm -44pp diversity: real gain - run 3c (lr 1e-05) selfharm -40pp diversity: none - -⚠ **The selfharm de-refusal is LR-ROBUST.** A 20x cut moved it 4pp while erasing the diversity -benefit entirely. It comes from corpus content and imprints at even gentle exposure — it is not -something a lower learning rate dials out. That is what the LR sweep was run to find out. - -⚠ **A pooled preserve-list test cannot see a single-axis collapse.** The pooled operational -delta reads 0.0 because two axes moved 40pp in opposite directions. Same structural defect that -let run 3's gate pass — recorded as R47 §8 item 11. - -## What the port proved about the box - -- **The tooling loads on aarch64/sm_121.** Full training stack plus flex_attention and the - chunked-loss path. Nothing exotic needed beyond `python3-dev`. -- **Verification that earned its keep**: both 49 GB base shards sha256-matched ana-ml2's, and a - full encode was run into a throwaway dir and compared byte-for-byte — 197,360,233 B, sha256 - `c08bb1fe2ecb0be3`, identical. transformers 5.15.1→5.16.1 and x86-64→aarch64 are *measured* - inert, not assumed. -- ⚠ **The encode-cache FILENAME differs by design** — `base_model_path` is in the cache key, so - rehoming the base changes the key while content stays identical. Input hash, not output hash. - Do not read it as drift; do not "fix" it by faking `/tank` on the GX10. - -## The lora_B signal worth carrying forward - - run 2 (lr 2e-04) min 0.6826 median 1.7212 max 3.7573 - run 3c (lr 1e-05) min 0.0480 median 0.1561 max 0.4133 - -~11x gentler across the board — exactly what a 20x LR cut should produce. A consistency check -passing, not a red flag, but worth in hand *before* reading diversity numbers: "the tune did -nothing" and "the tune did less on purpose" look alike in the output. - -Runbook `docs/runbooks/gx10-run-03c.md`; commit `dae77ee`. diff --git a/persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md b/persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md deleted file mode 100644 index f22cd63..0000000 --- a/persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md +++ /dev/null @@ -1,39 +0,0 @@ -# `[2026-09-05]` A peer's "2.7x stack effect" was a coin flip — the operator caught it and the arithmetic is worth keeping - -brokkr-smithy-dev reported that a diversity **noise floor** was 2.7x tighter on ana-ml2 than on -the GX10 and ranked ana-ml2 the better instrument on it. I relayed it. **The operator rejected it -on instinct** — *"that makes zero sense. except for speed, serving a model should be identical -across servers"* — and he was substantially right. - - ana-ml2 0.9688, 0.9574 spread 1.14pp mean 0.9631 - gx10 0.9583, 0.9895 spread 3.12pp mean 0.9739 - - between-box LEVEL difference 1.08pp - ana-ml2's own replicate spread 1.14pp <- LARGER than the between-box gap - pooled range ignoring box 3.21pp <- ~= the entire "GX10 floor" - -**Each "floor" is `|block0 - block1|` from n=2.** A range over two draws is not an estimate of -dispersion; the ratio of two such ranges is a ratio of two half-normals — a **half-Cauchy**. -brokkr computed the tail himself on retraction: `P(ratio >= 2.7) = 1 - (2/pi)*arctan(2.7)`, -doubled = **0.452**, confirmed by 400,000-draw simulation. He had reported a coin flip as a -measured effect and ranked hardware on it. Retracted at `97f73dd`. - -⚠ **The disconfirming evidence was inside his own sentence.** He wrote "the base level is nearly -identical, it is the spread that shifts" and offered it as *reassurance*. A real stack effect -moves the level. Level agreeing while a two-draw range differs is the signature of a noisy range -estimator — he had the refutation in hand and read it as support. - -⚠ **His own diagnosis of why it got through is the transferable part:** he had spent the session -triaging *my* claims hard — reading the encode path, simulating preflight, recomputing shas from -bytes — and this one was **his, and flattering**: it made his earlier work look prescient and -produced a clean recommendation. **Asymmetric scepticism is one error twice, and the flattering -direction needs the extra pass.** - -**What survived, deliberately separated:** re-measuring the floor on whatever stack actually -serves stays non-negotiable. The mechanism list is sound whether or not those four numbers show -it — sm_121 vs sm_120 kernel selection, vLLM 0.28.0 against mixed 0.26/nightly, and separately -measured batch-invariance (3.12pp at jobs=8, same order as the whole claimed effect). -**Retracting the evidence and keeping the discipline are different acts.** Settling it properly -wants several blocks per box and is its own probe, not a by-product of a gate. - -See [[2026-09-05-vllm-on-sm121-and-run4]]. diff --git a/persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md b/persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md deleted file mode 100644 index bd122ce..0000000 --- a/persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md +++ /dev/null @@ -1,53 +0,0 @@ -# `[2026-09-04/05]` vLLM RUNS on sm_121 — and the blocker was `ninja` not being on PATH, not the silicon - -Long-standing open question closed: **vLLM 0.28.0 serves on the GB10 (aarch64, sm_121)** from a -stock wheel, no source build. Model loads, `torch.compile` completes (~29 s), CUDA graphs capture. - -⚠ **The one trap, and it looks exactly like an sm_121 kernel problem:** FlashInfer JIT-builds its -sampling kernel at first use and needs **`ninja` on PATH**. It ships as a vLLM dependency but -lives in the venv `bin/`, so the failure surfaces as `FileNotFoundError: 'ninja'` from deep inside -a `profile_run` traceback. Same shape as the `python3-dev` trap that bit the training harness on -this box, and the third present-but-not-on-PATH false-absence on this hardware (after `nvcc`). - -Launch with **both** `~/vllm-env/bin` and `/usr/local/cuda/bin` on PATH. Harmless and expected: -`Using default MoE config ... device_name=NVIDIA_GB10` — nobody has tuned MoE kernels for this chip. - -⚠ **Open WebUI sends `tool_choice: "auto"` on every request**, which vLLM 400s unless launched with -`--enable-auto-tool-choice --tool-call-parser gemma4`. Taken from the working sibling seat -(`gemma4-charrp`), which is where the canonical flags live. **Deliberately did NOT copy -`--reasoning-parser gemma4`** — that seat pairs it with `--default-chat-template-kwargs -'{"enable_thinking": false}'`, and adding it alone moves output into `reasoning_content` with a -null `content`, which OWUI renders as an empty reply. A 400 traded for a blank box. - -## Run 4 — the corpus arm - -Run 4 adds an **airoboros-3.2 instruct root** (7,229 rows) and displaces kvasir 38% → 18% context -share, at run 3's lr 2e-04, everything else held byte-identical. Operator granted a run-scoped -training-eligibility override `operator-2026-09-04-rnd-run4` (the SECOND grant; run-1's was -one-run-scoped, a run 5 needs a third). - -**Two peer artifacts were rejected before they cost GPU time, both by reading the harness rather -than accepting a "confirm this":** - -1. `sample_kind: "instruct-dialogue"` — `core.py:521-544` dispatches on the **literal string** and - raises on anything outside `rp-dialogue`/`prose-chunk`/`actual-play-passage`. Would have - hard-stopped the encode on the first airoboros row. brokkr had labelled it *non-blocking*. -2. Recipe targets collapsed three dialogue roots into one descriptive label with `root_sha256: - null` on three of four. Preflight resolves `roots_dir//clean-v1/CLEANROOT.json` - literally and requires the sha. - -⚠ **I did NOT fill the missing shas in myself**, though the values were known and it would have -taken two minutes. `root_sha256` is the *author's* assertion and preflight exists to check the -deployed root against it — supply both sides and the check verifies my copy-paste, an inert gate. -brokkr then went further and recomputed his shas **from shard bytes** rather than reading them -back out of the deployed CLEANROOT, which had the same defect one step removed. - -**The kvasir prefix cut has three independent confirmations**: my build measured 1,613 samples / -3,347,622 ctx = **47.4%** of run-3 kvasir, which reproduces the recipe's 18/38 share ratio -(47.37%); the harness's own `[mix]` block then reported **ctx 0.1805** against the 18.0% target. - -⚠ **`str.splitlines()` splits on U+2028**, which `json.dumps` leaves unescaped under -`ensure_ascii=False`. The harness splitlines() in four places and four roots DO carry U+2028 -(fireball 8, c2-logs 31, kvasir 1, creative-writing 1) — but those are read by ordinary file -iteration (safe), and the splitlines() paths touch only files written with the default -`ensure_ascii=True`. Latent, one flag away from live, not acting. Found by brokkr the hard way. diff --git a/persistent-memory.d/2026-09-06-headscale-cutover.md b/persistent-memory.d/2026-09-06-headscale-cutover.md deleted file mode 100644 index 02aa8f2..0000000 --- a/persistent-memory.d/2026-09-06-headscale-cutover.md +++ /dev/null @@ -1,33 +0,0 @@ -# 2026-09-06 — Headscale cutover COMPLETE: all three site-pairs on the mesh - -**DONE.** Operator disabled Site Magic in the UI; NH3↔ESH re-homed to a DIRECT mesh path (8ms, no DERP). Full 6-direction matrix OPEN. Site Magic disabled, both IPsec tunnels dormant, headscale is the sole active site-to-site transport. ana-wg WG fallback untouched. Tunnels re-enablable for backup. - -Operator goal (/goal): replace Site Magic + IPsec with headscale, tunnels dormant as backup; -"if paranoid, enable world-accessible SSH on the FortiGate first." Full detail + method + -follow-ups in `docs/pfi/headscale-mesh-plan.md` § CUTOVER EXECUTED. Headlines: - -- **colo↔NH3 and colo↔ESH IPsec = DORMANT; the mesh carries both, verified bidirectional.** - NH3 UDM `pfi-nh3-ana` + ESH UDM `esh-ana` set enabled=false (API). Mesh /16 routes added on - both UDMs and the FortiGate (→ ana-scale 10.250.50.45 / nh3-scale 10.100.50.46 / esh-scale - 10.0.50.65). Dependent flows OK over mesh: restic ESH→rest-server-ana, FortiGate mgmt. -- **NH3↔ESH Site Magic NOT cut by API** — `sdwan-mesh-tunnel` = `api.err.NoEdit` (cloud - orchestrated). Routes PRE-STAGED + shadowed; DERP path 9ms ready. **Operator disables it in - the UniFi UI**, then the mesh takes over. Told the operator "mesh is online" → he does it. -- **FortiGate WAN SSH safety net (TEMPORARY):** wan1 allowaccess ping+ssh; admin infra-ops - trusthost2/3 = NH3 70.230.226.88 + ESH **128.177.138.182** (static since 09-08; was CGNAT 23.164.40.160) (not 0.0.0.0). Reach it at - `ssh infra-ops@38.120.12.42`. Config backed up flash `pre-wan-ssh-cutover-20260906`. Remove - when the edge (being replaced by OPNsense/R420) is retired. -- ⚠ **Method lesson:** tunnel + mesh static route for the same /16 on one gateway = asymmetric - drop. Disable the tunnel FIRST, then add the route. Broke colo once doing it tunnel-up; rolled - back. See [[incident_crowdsec_cgnat_false_ban]] (same day) and the plan doc. -- Dormancy = disabled+retained (flip UDM object back to enabled=true to restore); NO auto - failover wired. Bonus: exit nodes → free multi-location egress proxy (parked). - - -## Exit nodes (2026-09-06, operator-requested) -All three routers advertise+serve exit nodes (approved). Clients pick location: -`tailscale set --exit-node=nh3-scale|esh-scale|ana-scale`. NH3 = residential egress -(70.230.226.88) → replaces the nh3-dev SOCKS5 proxy. Exit nodes + source preservation BOTH work via a selective-masquerade rule (NoSNAT kept true; -`mesh-exit-masq.service` per router masquerades only internet-bound exit traffic, RETURNs fleet -dests). Verified: colo sees real NH3 host; nh3-dev via colo exit → egress 38.120.12.42. A node advertising an exit node can't consume one — test from -the laptop/iPad, not the routers. diff --git a/persistent-memory.d/2026-09-06-headscale-mesh-phase1.md b/persistent-memory.d/2026-09-06-headscale-mesh-phase1.md deleted file mode 100644 index 4f025e1..0000000 --- a/persistent-memory.d/2026-09-06-headscale-mesh-phase1.md +++ /dev/null @@ -1,16 +0,0 @@ -# 2026-09-06 — Headscale overlay mesh: control plane + 3 subnet routers live, not cut over - -Operator-directed (Headscale over NetBird; NH3 for the control plane, never the colo; 443 -direct; names nh3-headscale / nh3-scale / esh-scale / ana-scale). Full state, lessons and -next steps in `docs/pfi/headscale-mesh-plan.md` § Status. Headline facts: - -- `https://headscale.phasefinal.com` = CT 106 on nh3-pve (10.100.50.45), headscale v0.29.3, - LE cert via TLS-ALPN-01, UDM forward tcp/443, DDNS timer on nh3-dev (user systemd). -- Routers CT 107 nh3-scale / CT 108 esh-scale / CT 114 ana-scale advertise their /16s, - approved, SNAT off, accept-routes OFF. nh3-dev enrolled as first client (100.64.0.4). -- ⚠ Old tunnels (Site Magic, IPsec) are STILL the site-to-site path. The mesh currently - rides inside them. Nothing has been disabled. -- ⚠ Lesson: `--accept-routes` on a client before a return path for 100.64.0.0/10 exists - black-holes that client's LAN (own-site /16 included). Return path first. -- Pre-auth keys in the vault (`headscale/preauth-*-48h-20260906`, expire 09-08). -- infra-ops user now exists on all four PVE hosts (needed `apt install sudo` first). diff --git a/persistent-memory.d/2026-09-21-booth-two-dead-controls.md b/persistent-memory.d/2026-09-21-booth-two-dead-controls.md new file mode 100644 index 0000000..9d05d1e --- /dev/null +++ b/persistent-memory.d/2026-09-21-booth-two-dead-controls.md @@ -0,0 +1,45 @@ +# `[2026-09-21]` The Booth gained blur and a closed keep round trip — after shipping two controls that did nothing + +**Shipped** (2e7fd71, 271cb11, 751eecb, 07c9cb2): per-item cosmetic blur +(`.blurred` marker, CLI `blur`/`unblur`, caption toggle, click-to-reveal, cover +thumbs inheriting it), the ephemeral→kept `★` button closing a round trip that +previously needed a shell, a direct `×` on kept cards, and in-booth +keep/release with an open-redirect-safe `next`. + +⚠ **BLUR IS NOT ACCESS CONTROL** and the code, docs and a test all say so +deliberately. A blurred item is still served at its own URL, still in the zip. +`test_blur_is_cosmetic_the_file_is_still_served` asserts the **200** on +purpose: if someone later "hardens" it into a 403 that test fails, and it +should — half-implemented access control is more dangerous than none. + +⚠ **Two controls shipped INERT, both found by the operator, both by me reading +templates instead of rendering them:** + +- **The reveal button.** Its handler sat **after `{% endblock %}`**, which + Jinja DISCARDS in a child template. The button rendered; the handler never + reached the browser. Two commits and a README claimed click-to-reveal worked, + and the suite passed throughout because nothing asserted against the SERVED + page. Guards added and **confirmed to fail on reintroduction**. +- **The kept-card `×`.** Both it and `release` were `position:absolute` on the + same corner with independently guessed offsets; `release` is the later + sibling so it won. Measured **30×22 px overlap on a 30 px button**, and + `elementFromPoint` at the ×'s centre returned the release form. Unclickable + from the moment it shipped. Replaced with one flex row positioned once. + +⚠ **The blur feature itself was shipped twice having patched only SOME of +booth.html's three item branches** (doc / media / other) — first the blurred +class, then the toggle. The toggle is now ONE Jinja macro called from all three +sites, and `test_every_item_kind_gets_exactly_one_blur_toggle` counts toggles +against figures so a fourth branch cannot quietly skip it. + +⭐ **`scripts/layout-probe.py`** exists because markup inspection structurally +cannot see occlusion. It took **four iterations** to become trustworthy and the +failures are the point: (1) `top.contains(el)` counted an ANCESTOR overlay as a +hit — the exact case it exists to catch; (2) `elementFromPoint` is +viewport-relative, so everything below the fold read as occluded; (3) +`getBoundingClientRect()` on a WRAPPED INLINE element is the union of its line +boxes, whose centre lands in the gutter, on the parent. Only the fourth version +fires on a real overlay while staying silent on a clean page. **Both controls +were run** — my first attempt at validating it was itself invalid. + +See [[2026-09-21-ops-log-and-the-instruments-that-lied]]. diff --git a/persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md b/persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md new file mode 100644 index 0000000..110b0f0 --- /dev/null +++ b/persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md @@ -0,0 +1,41 @@ +# `[2026-09-21]` nh3-dev disk triage: 27 GB reclaimed, and a LoRA adapter rescued from a directory the box sweeps every 3 days + +Beszel alerted root >85%. Reclaimed **7 GB** from regenerable caches (`uv cache +prune`, npm, pip), then a deep dive found the real shape. + +⚠⚠ **The headline was not disk.** `/tmp` held **19 GB** of Claude Code session +scratchpads, and inside one of them sat the **`babyyarros` LoRA adapter** — +252 MB, r=32/α=64 on Qwen3-4B-Instruct, with its loss series and provenance — +**existing nowhere else**: not on `/mnt/smithy`, not under `~/development`. +`/etc/tmpfiles.d/tmp.conf` sets `D /tmp … 3d`, an admin file from 2026-07-18 +that **overrides** the stock no-age rule, and the cleaner runs daily. + +**Rescued** to `/mnt/smithy/adapter-rescue/babyyarros-20260921`, verified by +content: sha256 `63fda6cc…` matching both the source and the `adapter.sha256` +recorded at training time. + +⚠ **Two corrections I made to myself during the dive, both worth keeping:** I +alarmed that shutterchute's RAW deliverables were 19 hours from deletion — their +mtimes were *that day*, a live session working. And I suspected my own `du`/ +`find` had reset the atime clock and manufactured the "0 would-remove" result; +it had not (`relatime`, and an untouched comparison file still showed an old +atime) — but it was right to check before trusting a number my own measurement +could have created. + +**Operator-authorized deletions:** 7.6 GB duplicate Qwen base shards, 12 GB +`models-staging/retro-diffusion` (cold since 09 Aug), 7.5 GB `splat-assets` +(cold since 15 Aug), 178 session dirs idle 7d+. **39 GB → 66 GB free, 84% → +72%.** + +⚠⚠ **MY PRUNE DELETED AN ACTIVE SESSION'S DIRECTORY.** `-mtime +7` on a session +dir is an unsound liveness test: **a directory's mtime does not change when +files are written into its subdirectories.** `dfacccde/` looked 7+ days idle +while `dfacccde/tasks/` was being written continuously. Cost: one lost tool +output; recreated. All four of my safety assertions (right root, right depth, +own session excluded, shutterchute excluded) passed — I checked the paths were +right and never checked the liveness test was sound. **Do not re-run that +predicate.** A correct version checks the deepest recent file, or +cross-references running `claude` PIDs. + +Also found: `sudo -n` requires a password as **lkraven locally** on nh3-dev, +while `ssh infra-ops@localhost` has NOPASSWD. diff --git a/persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md b/persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md new file mode 100644 index 0000000..4cea9cd --- /dev/null +++ b/persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md @@ -0,0 +1,39 @@ +# `[2026-09-21]` claude-bot became an org Owner, repos moved to `pfi`, and a dead token had been misreporting permissions for months + +**Operator ruling:** claude-bot is an **Owner** in `corviduo` (team 1), `pfi` +(team 4) and `vastblue` (team 5). Verified by reading membership back AND by +exercising it on claude-bot's own token: create/edit/delete in `pfi` all +succeed, `vh/*` correctly still 404s. + +⚠ **Why not "admin on vh/*", which is what was originally asked:** `vh` is a +**USER account (id 1), not an organization** — `/orgs/vh` 404s. Gitea has no +namespace-scoped admin for a user namespace. Measured: per-repo `admin` +collaborator grants read but **not** settings (403 on PATCH) — repo settings are +owner-only. So the only working realization of "admin over vh/*" is the +**instance-wide site-admin flag**, which would have given claude-bot the same +blast radius as the token the credential-migration project exists to retire. +Surfaced rather than executed; the operator chose orgs instead. + +**`vh/cicada` and `vh/draupnir` transferred into `pfi`** with SHAs preserved, +old paths 301ing, and — verified — **the old ssh remotes still resolve**, since +Gitea redirects git-over-ssh and not just the web URL. Repoint anyway: a remote +living on a redirect depends on the old path staying unclaimed. + +⚠⚠ **`~/.config/claude-bot/gitea-token` IS DEAD** — it authenticates as +**nobody** (`/user` → `None`). I had cited its 403s and 404s twice, to the +operator and to a peer, as evidence that claude-bot lacked rights in `vh/*`. The +conclusion survived re-testing with a working credential, but the evidence was +worthless. **A credential that authenticates as nobody returns 403 and 404 for +everything, and that is indistinguishable from a permissions answer.** There +are FOUR token files in that directory; the working one for repo work is +**`gitea-token-repo-create`** (`write:organization`, `write:repository`, +`write:user`). + +**Standing authorization (operator, same day):** routine vh-token use for +`vh/*` repo ops no longer gets a flag. Read it from the **vault** — +`secret get 'nh3-dev/.config/gitea/vh-token'` — verified byte-identical to the +disk copy and authenticating as `vh` (id 1, is_admin=True). ⚠ This did not make +the token low-blast-radius; it stopped the class of work being exceptional. + +Recorded in auto-memory as `feedback_wrong_resource_before_wrong_peer` and in +`reference_infra_ops_vh_gitea_token_and_sdk_publish`. diff --git a/persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md b/persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md new file mode 100644 index 0000000..023c8b5 --- /dev/null +++ b/persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md @@ -0,0 +1,44 @@ +# `[2026-09-21]` The ops log shipped, and the day's real subject was instruments that report without looking + +**Built** `scripts/ops-log` (ffe7b24) to close the fingerprint-less-change gap: +infra-ops and infra-hermes act as one OS identity, dockerd logs no per-caller +exec, and every commit here is attributed to Vuong Hoang by convention. One +appended line per host-changing action, a `mkdir`-atomic claim `deploy-stack.sh` +refuses (exit 3), automatic writers in `deploy-stack.sh` + `elway`, and +`ops-log audit` as the detector for the raw-`ssh` path the writers cannot see. +136-stack baseline laid so the detector starts from that day. + +⚠ **The instrument then failed FOUR ways in its first hours, and every one +recorded something — just nothing findable.** Documented as a table in +`docs/pfi/ops-log.md` § "How this instrument has failed", which is the durable +artifact: + +1. **Claim released by a sub-tool** (3e7d3a3) — a 45-min operation claim was + refreshed then released by `deploy-stack.sh`'s exit trap, mid-rollout. + `claim` now exits **10** when already yours and leaves the holder file + untouched, so a refresh cannot overwrite the reason and TTL the original + claimant chose. +2. **Wrong order in the hook chain** (9141a41) — the commit hook was APPENDED + behind graphify's **eight `exit 0` paths**, so a `graphify-out/`-only or + empty commit could never be recorded. Prepend; attribution must never be a + subordinate clause of another hook's interestingness filter. +3. **No handle in the environment** (4e778ae) — `ALTHING_HANDLE` lived only in + `althing-infra-hermes-seat-run.sh`, not the gateway unit. Fallback now says + `unattributed(login)` rather than a bare login that reads like an answer. +4. **Wrong host key on write** (f3b68e2) — elway passed its ssh TARGET through + as the host, so five records of a real esh-pve change landed under + `infra-ops@esh-pve` and were invisible to `--host esh-pve`. infra-hermes + correctly reported the change as unattributed. **A log you cannot query + under the obvious name is not a log.** + +⚠ **The general lesson, and it outlived the tool:** twelve instruments reported +confidently and wrongly across 2026-09-19→21, five of them mine. The recurring +shape is **configured ≠ effective** — `systemctl show -p Environment` reporting +a drop-in while `/proc//environ` lacked it; a grep proving presence while +evaluation proved absence; a green test suite over a control the browser never +received. What broke the pattern every time was asking a *different* instrument +the same question. + +See [[2026-09-21-booth-two-dead-controls]] for the same failure in a UI, and +`feedback_control_flow_before_concurrency` in auto-memory for the triage rule +that came out of it. diff --git a/persistent-memory.md b/persistent-memory.md index e448d4b..ffb8fcb 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-09-21 ~05:55 UTC (⭐ **ravenpen.com registered** by the operator — Cloudflare registrar, expires 2028-09-20, zone active but BARE; hamr-dev answered, the 09-18 hold discharged. ⭐ Booth gained keep-both-ways + per-item cosmetic blur. ⭐ Backup alarm verdict split: STALE (exit 1) vs ERRORED-JOBS (exit 3) — it was crying STALE over a fleet whose every body was fresh. ⭐ The ops log is BUILT; commit attribution now works end-to-end, both controls measured. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED.)_ +_Last updated: 2026-09-21 ~14:30 PT (⭐ **the ops log is BUILT** and then failed four ways in its first hours — every one recording something unfindable; the day's subject was instruments that report without looking. ⭐ Draupnir engine COMPLETE and acceptance-tested on irv-ml1. ⭐ Booth gained blur + a closed keep round trip after shipping two dead controls. ⭐ claude-bot is an org Owner; cicada+draupnir moved to `pfi`; vh-token use standing-authorized from the vault. ⭐ nh3-dev 84%→72%, and a LoRA adapter rescued from a 3-day-swept /tmp. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED after two days.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -115,112 +115,65 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-09-19 ~07:15 PT._ +_As of 2026-09-21 ~14:30 PT._ -### ✅ BUILT — the ops log (`scripts/ops-log`), 2026-09-19 +### ⚠ FIRST — lv-mccarthy's run outcome is STILL UNVERIFIED -Shipped. `scripts/ops-log` + `docs/pfi/ops-log.md`, wired into `deploy-stack.sh` -(claims + records) and `elway` (records). Baseline laid: **136 stacks across 6 hosts** -marked pre-ops-log, so the detector starts from today instead of reporting the whole -fleet as unattributable forever. +Carried unchanged from the 2026-09-19 handoff and untouched for two full days. +Look at `~/r49-runs/mccarthy-4b-pairs-3ep/` before anything else. ⚠ `pfi-gx10` +does not resolve from nh3-dev by that name. Assume nothing — it was never +checked, not checked-and-found-good. Then: checkpoint selection off the loss +curve → the v2 gate (`eval-*.sh`) → ship-or-park. ⚠ Before the gate, settle how +the VOICE axis is read: D1 pre-registered a punctuation-normalised secondary +read, and `--system-from` hands the register's tics to the BASE arm too, which +changes what the primary number means. -**The four open questions, settled:** +### Draupnir engine — COMPLETE on irv-ml1, acceptance passing -1. **Where it lives — CENTRAL on nh3-dev** (`/.ops-log/`, gitignored), not - per-host and not the post office. Decider: both agents run as the *same unix user* - on nh3-dev (infra-hermes is a **user** unit under `lkraven`), so one file is shared - instantly with zero provisioning. Per-host needs a writable path on ~25 heterogeneous - boxes and puts the record of "we changed host Y" *on host Y*. syslog/journald looked - free but journald shows an unprivileged reader only their own `_UID` — the log would - have split silently between the `infra-ops` and `lkraven` halves of the fleet. The - post office is a bus, and an outage there would block ops during the incident you are - reconstructing. ⚠ **Known hole, stated not papered over:** an actor on a box *other - than nh3-dev* is uncovered. Today that is only the operator's laptop; a third agent - elsewhere is what would force a revisit. -2. **Claim = advisory, enforced in the tooling.** `deploy-stack.sh` refuses (exit 3) a - stack another agent holds, across the diff, the y/N prompt AND the apply — the whole - review window, which is where the 09-18 collision actually happened. Acquire is - `mkdir` (atomic → genuinely race-free). TTL 30m; a stale claim auto-breaks **and the - break is logged**, so an ineffective claim is visible rather than silent. -3. **Writers are AUTOMATIC.** This was the one that mattered — a log you must remember - to write is the same class of instrument as a health check that passes in both states. -4. **There is a DETECTOR, not just a rule.** `ops-log audit` asks each host what changed - on disk and compares it to the newest log line for that stack. Covers the manual - `ssh`-and-edit path the automatic writers structurally cannot. +build123d 0.12.0 + OCP, numpy 2.4.6, trimesh 5.1.0 on py3.11.2; FreeCAD 1.0.0 +AppImage headless; OrcaSlicer 2.4.2 containerised at `~/bin/orca-slice`; +artifact root `/mnt/smithy/draupnir` 2775, 500 GB budget. The contrastive +control pair PASSES. **The operator has moved to a code session inside the +Draupnir repo**, so the next questions come from there rather than from +brokkr-smithy-dev, and everything is in the repo (b8db50b) rather than only in +the althing thread. -⚠ **Found and fixed a bug in my own detector mid-build:** it printed "audit clean" for a -host it never reached. Now `INCOMPLETE` + exit 5 — an unreachable host is not a clean -host. Both directions proven on a live host: after baseline, a mtime-only `touch` on -`nh3-docker/beszel-agent-nh3` fired the detector at 15 s resolution, and recording it -cleared it. +### The fleet is quiet and nothing is blocked on me -**Not covered, on purpose:** raw `ssh` (audit is the backstop), per-host claims for elway, -`corviduo-dev` (CI/CD rewrites the tree constantly → permanent false positives), the -SureFire tenant hosts. **Follow-ons:** hook `dns-sync.py` / UniFi / FortiGate helpers so -control-plane changes record themselves; run `audit` on a timer. +- **Backups:** `RESULT: all backups fresh`. The two `⏸` policy exclusions + (ana-scale CT 114, esh-vm-workstation VM 102) are correct and visible. + infra-hermes holds the watch and names CT 107 explicitly rather than trusting + the absence of red — ⚠ that lock was RELEASED BY HAND, not self-healed, and + its mechanism is still unexplained. +- **Disk:** nh3-dev at 72%, 66 GB free after the 2026-09-21 triage. +- **ops-log:** clean audit across 6 hosts; commit attribution working. +- **19 commits unpushed on main**, plus the memory changes from this snapshot. -### ⚠ FIRST — lv-mccarthy's run outcome is UNVERIFIED by this session +### Open, low-urgency -At the previous snapshot (09-17 ~23:05 PT) the LoRA was at ~150/1380 steps with ETA ~00:45 PT. -**This session was entirely infrastructure and never touched it**; `pfi-gx10` does not resolve from -nh3-dev by that name, so the outcome was not checked rather than checked-and-found-good. Assume -nothing: look at `~/r49-runs/mccarthy-4b-pairs-3ep/` before anything else. - -**Next, unchanged:** checkpoint selection off the loss curve → the v2 gate (`eval-*.sh` shape) → -ship-or-park. ⚠ Before the gate, settle how the VOICE axis is read — D1 pre-registered a -punctuation-normalised secondary read and the `mccarthy` register now states the punctuation tics -explicitly, so `--system-from` hands them to the BASE control arm too. Deliberate (denies the adapter -a cheap char-bigram win) but it changes what the primary number means, and it must be settled BEFORE -a number exists. - -### The fleet is materially faster than it was this morning — three fixes, all verified - -Detail in the Recent decisions entries below; the operational summary is that **NH3→Anaheim went -from a throttled DERP relay to direct (6 ms, cross-site HTTP 1.2 s → 0.015 s)**, **`.internal` DNS -stopped failing ~10% of lookups and stalling 5 s on the rest**, and **SearXNG went from one working -general web engine to seven**. All three were silent — no monitor caught any of them, and two had -been degrading for months. - -⚠ **Nothing on this fleet watches DNS success rate or whether a mesh path is direct.** Today's three -faults surfaced only because tts-dev had a 1545 ms voice-loop budget and measured instead of adapting -around it. **A probe pair is proposed and UNDECIDED** — see the Recent decisions entry. - -### Still relayed: irv-ml1 - -`100.64.0.6` remains `relay "lax"` after the Anaheim fix — a different site with its own NAT -situation, untouched. Measured 24-25 ms on HTTP from nh3-dev, so it is NOT costing what Anaheim was -and needs nothing urgently. I warned tts-dev it would be slow and was wrong; they measured and -corrected me. - -### OPEN LOOP — `ravenpen.com`, and a reply hamr-dev is waiting on - -hamr-dev asked infra-ops to register `ravenpen.com` on a relayed operator directive. **Surfaced, not -executed**: a peer relaying "Vuong approved it" is not authorization for a non-refundable purchase, -and infra-ops holds no registrar credential. hamr-dev accepted and closed out. - -⚠ **If the operator registers it, POST BACK on althing thread `01M2SERDB3DR3RV7JS1J0GMF0H`** with -registrar and expiry — hamr-dev is explicitly waiting on that, and it is the kind of cross-session -obligation a reset drops silently. Verified 2026-09-18 ~05:20Z: `ravenpen.{com,io,app,dev,ai,net}` -all unregistered. - -### PLANNED (operator) — move `dragonfireacoustics.com` to Namecheap, DNS to Cloudflare - -Unchanged from the previous snapshot and **not started**. ⚠⚠ The zone's `*` wildcard MASKS what is -really there — enumerate real records before any transfer. Expiry **2026-10-30** at eNom with no -transfer lock, and the sibling `dragonfirepro.com` was already lost exactly this way. Customer-facing; -nothing touched. Full context in the 09-17 Recent decisions entries. - -### Other standing items, unchanged - -- **lv-hemingway keeps its 3 separator-hidden names** (operator: "leave it"), so `leak_gate.py` exits - 1 on a shipped tree BY DESIGN. A red result there is expected, not a bug to fix. -- **`lv-krakauer` PARKED** (operator 2026-09-17), henge id **82**. -- **ESH is on the Cityside static** `128.177.138.182/30`, healthy on four axes. The two 7-day crowdsec - allowlist entries expire 2026-09-23 and are being left deliberately — Cityside failed twice in six - hours on 09-16/17. +- ⚠ VM 102's efidisk carries **UEFI 2011 certs expired June 2026**. Needs the + sandbox down and BitLocker protectors suspended first. Operator's machine, + operator's call; surfaced, not acted on. +- Nine unnecessary packages on irv-ml1 (`libwebkit2gtk-4.1-0` + deps) from a + serial dependency chase. Left deliberately — `autoremove` on a box running + twelve production services is a second risk, not a cleanup. ## Recent decisions +- `[2026-09-21]` ⭐⭐⭐ **The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable.** Claim dropped by a sub-tool, hook behind graphify's eight `exit 0`s, no handle in the env, ssh-target written as a hostname. The general shape is **configured ≠ effective**; twelve instruments reported confidently and wrongly across three days, five of them mine. → `persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md` + +- `[2026-09-21]` ⭐⭐ **The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing** — a reveal handler Jinja discarded for sitting after `{% endblock %}`, and a `×` a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. `scripts/layout-probe.py` took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → `persistent-memory.d/2026-09-21-booth-two-dead-controls.md` + +- `[2026-09-21]` ⭐⭐ **claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to `pfi`; vh-token use is now standing-authorized from the vault.** ⚠ `vh` is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ `~/.config/claude-bot/gitea-token` is DEAD and had been misreporting permissions; the working one is `gitea-token-repo-create`. → `persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md` + +- `[2026-09-21]` ⭐⭐ **nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from `/tmp`, which this box sweeps at 3 days.** `babyyarros` existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → `persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md` + +- `[2026-09-21]` **Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested** (34c4179, e574b91, 2e08edc). build123d 0.12.0 + OCP, FreeCAD 1.0.0 AppImage headless, OrcaSlicer 2.4.2 **containerised** — Debian 12's glibc 2.36 cannot run any current Orca build (needs GLIBC_2.38, verified by `ldd`), and reaching back for an Ubuntu-22.04 build would pin permanently to stale. ⚠ OrcaSlicer writes `result.json` into CWD on EVERY invocation, `--help` included. Playbook `playbooks/irv-ml1-draupnir-engine.yaml`. + +- `[2026-09-21]` **Backup alarm verdict split: `STALE` (exit 1) vs `ERRORED-JOBS` (exit 3)** (7fe4102). It had been printing STALE over 37 FRESH layers and zero stale ones — a false statement of fact, flagged by infra-hermes. STALE is a claim about backup AGE; a job that ran and errored is a different claim with different urgency. Wrapper mirrors the code and sends 🟡 not 🔴. + +- `[2026-09-21]` **`vh/forgefirm` mirrored** from `github.com/openglow-org/forgefirm`, following the house convention read off the existing 17: `vh/` namespace, upstream casing, `8h0m0s` interval (16 of 18), visibility matching upstream. Verified by HEAD SHA (`08b29fee`) against upstream, not by the 201. ⚠ `vh/NetAlertX` interval `0s` is **deliberate** — operator: "no longer interesting to us". Not a broken mirror; do not re-enable. + - `[2026-09-20]` **`ravenpen.com` REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered.** infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a peer relay with no registrar credential on our side; the standing instruction was to post registrar + expiry back on thread `01M2SERDB3DR3RV7JS1J0GMF0H` only once the operator bought it himself. Done. Facts from **RDAP (Verisign, authoritative)** rather than a dashboard: registrar **Cloudflare, Inc.** (IANA 1910), registered 2026-09-20T20:12:30Z, **expires 2028-09-20** (two-year), `clientTransferProhibited`. Zone `2df4c5eb4ea4b9410423bdebcb6c5192` active on `vh@phasefinal`, activated 0.4 s after creation — registered THROUGH Cloudflare Registrar, which is why the zone's `original_registrar` is null. ⚠ **The zone is BARE — zero DNS records**, so the name resolves to nothing and mail to it bounces; correct for bought-not-built, but say so before anyone points at it. ⚠ **Scope boundary measured, not assumed:** the fleet `infra-ops` Cloudflare token is Zone·DNS·Edit and **403s on the Registrar API** — I can build records in the zone, and I can NOT read auto-renew state, renew, or transfer. **Auto-renew is therefore UNCONFIRMED**; do not let anyone assume it. ⚠ The token is **vaulted, not on disk** — `secret get 'nh3-dev/.config/cloudflare/infra-ops-dns-token'`; the memory's "vaulted at nh3-dev/..." names a VAULT KEY, and reading it as a filesystem path wastes a step. - `[2026-09-19]` ⭐⭐ **FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed).** The ceiling is **1920 W**, not 2400 — a GPU host running for hours is a continuous load, so NEC's 80% rule governs. Measured: BMC `dcmi power reading` 390 W instantaneous / 461 W max over 2423 s **at idle GPUs**, so the non-GPU baseline is ~313 W (EPYC 9254 24C/96T, 5+ drives, 4 PSUs). GPUs are capped **275 W** each against a **300 W** stock TGP (`power.default_limit`) — a 100 W saving across four cards, NOT the 200 W you get by measuring against the 325 W firmware ceiling; the operator corrected me on exactly that. Derived worst case **~1625 W capped (85%)** vs **~1725 W at stock (90%)**. ⚠ **Keep the caps** — 90% leaves nothing for a heavier-than-estimated R420, PSU efficiency, or a warm day. ⚠⚠ **The coupling is worse than the trip:** OPNsense IS the FV edge and shares the breaker with the thing most likely to trip it, so an overload takes the router with it and removes the remote path needed to recover. Four PSUs do not help — they are all downstream of one breaker. ⚠ **NOT measured:** fv-ml1 under real 4-GPU load (the 461 W max is idle-ish), whether the BMC reports AC or DC (±10% ≈ 150 W at load), and the R420's actual draw. Recorded in `servers/fv-ml1/README.md` § Power. ⚠ Also fixed there: the README claimed **2x** GPUs; `nvidia-smi` reports **four**. @@ -474,35 +427,21 @@ nothing touched. Full context in the 09-17 Recent decisions entries. - `[2026-09-08]` **Fleet fixes shipped** — WhereTF Homepage card + DNS (`4506ef6`); ext-tts LiteLLM alias → `irv-ml1.nh3.internal` (DB `/model/update` + `extra_hosts`, `957c8f1`); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (`e0d1c44`); Homepage `/api/services` outage fixed — ana-ml2 discovery via a socat proxy on ana-docker (`stacks/ana-ml2-proxy`, `913d2d2`, reversible). -- `[2026-09-07]` **Fleet internal TLS pattern shipped** — caddy (cloudflare-plugin build, `~/.local/bin/caddy-cf`, `fleet-tls-caddy.service`) on nh3-dev is the wildcard cert authority: publicly-trusted LE `*.nh3.phasefinal.com` via Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers). `talk` self-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced by `fleet-tls-cert-check.timer`. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memory `reference_fleet_internal_tls_pattern`. -- `[2026-09-07]` **cc-channel registered for this infra-ops session's wake** — `althing-route` cc route → the CC session's `$XDG_RUNTIME_DIR/cc-socks/.sock`; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat in `shell`. Session-local — re-declare per session. -- `[2026-09-07]` **irv-ml1 /mnt/smithy remount fixed post-cutover** — export allowed `10.0.0.0/8` (old wg0) but not the mesh `100.64.0.0/10` irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin = `infra-ops` PASSWORD auth (vault `nh3-nas/infra-ops-password`), sudo ALL, SFTP subsystem OFF. → auto-memory `reference_irv_ml1_gpu_r14` (corrected). -- `[2026-09-07]` **irv-ml1.nh3.internal DNS repointed** to the live Irvine LAN IP `10.6.110.50` (was the dead wg0 `10.100.79.3`); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit `0336e03`. -- `[2026-09-07]` **Subnet routers excluded from vzdump fleet-wide** (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memory `feedback_esh_backup_window_0330`. -- `[2026-09-07]` **Booth link board: pin/favorite + multi-select delete + newest-first** (booth-v0.1.8, commit `76fdf45`, tag `booth-v0.1.8`) — pins in a `.pins` sidecar (content-ids), one `` + `formaction` buttons so ×/★/bulk-delete all degrade with JS off. -- `[2026-09-06]` **Headscale cutover COMPLETE — all three site-pairs on the mesh; Site Magic + both IPsec tunnels DORMANT.** Operator disabled Site Magic (UI); NH3↔ESH re-homed to a direct 8ms path. Exit nodes advertised at all three sites (multi-location egress proxy) with source preservation kept via a selective-masquerade rule (NoSNAT + `mesh-exit-masq.service` per router). Throughput 761/464 Mb/s vs old 250 IPsec. ⚠ FortiGate WAN-SSH left open (temp, scoped NH3+ESH). Method: disable tunnel FIRST then add mesh route. → `persistent-memory.d/2026-09-06-headscale-cutover.md` -- `[2026-09-06]` **Headscale overlay mesh: control plane live at `headscale.phasefinal.com` (CT 106 nh3-pve) + subnet routers nh3-scale/esh-scale/ana-scale serving their /16s; nh3-dev enrolled. NOT cut over — Site Magic + IPsec still carry site-to-site.** ⚠ accept-routes-before-return-path black-holed nh3-dev's LAN for a minute. infra-ops user added on all four PVE hosts. → `persistent-memory.d/2026-09-06-headscale-mesh-phase1.md` - `[2026-09-06]` **pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10** (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroy `ospool/naspool-evac` after scrub + one backup cycle; backplane swap next visit; PSU1 still dead. → `persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md` -- `[2026-09-05]` **A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him.** Each floor was `|b0-b1|` from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named *why*: the claim was his and flattering. → `persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md` -- `[2026-09-05]` **vLLM RUNS on sm_121 — the blocker was `ninja` off PATH, not the silicon** — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missing `root_sha256` values I knew, because supplying both sides of a check makes it inert. → `persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md` -- `[2026-09-04]` **ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression.** LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. → `persistent-memory.d/2026-09-04-run3c-trained-and-gated.md` -- `[2026-09-04]` **`gen` moved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint.** Cost: gen KV down to 1.02x concurrency at 262K. → `persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md` -- `[2026-09-04]` **SMB account `dsp` created + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open.** Twelve NFS exports rw to `10.0.0.0/8`, guest-writable SMB. → `persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md` -- `[2026-09-04]` **SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at `10.0.90.10`, DHCP-reserved, DNS'd, handed to ha-dev.** ⚠ Home Assistant cannot resolve `.internal` at all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbook `docs/runbooks/slzb-mr1u-zigbee-coordinator.md`, commits `fed29be`/`0bbdaf9`. - `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md` @@ -519,12 +458,12 @@ nothing touched. Full context in the 09-17 Recent decisions entries. - `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now"). -_11 older entries archived to archival-memory.md._ - -_5 older entries archived to archival-memory.md._ +_30 older entries archived to archival-memory.md._ ## Tried and abandoned +- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working directory. A directory's mtime does not change when files are written into its SUBDIRECTORIES, so a session writing continuously to `/tasks/` looks 7+ days idle at `/`. Four safety assertions passed; none of them asked whether the liveness test was sound. Use the deepest recent file, or cross-reference running `claude` PIDs. + - `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md` - `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing. @@ -534,6 +473,5 @@ _5 older entries archived to archival-memory.md._ - `[2026-09-15]` **Remote-site MASQUERADE rules on nh3-scale** for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing. -- `[2026-09-04]` **Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes.** ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. → `persistent-memory.d/2026-09-04-dac-forced-10g-failed.md` -_114 older entries archived to archival-memory.md._ +_115 older entries archived to archival-memory.md._