From 5abd43e1c8150d8f5174e8a40ef6d7b3a8031fed Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Sat, 3 Oct 2026 13:36:42 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=202026-10-03=20i?= =?UTF-8?q?n-flight=20(nh3-pve=208=E2=86=929=20under=20way),=20two-tier=20?= =?UTF-8?q?split=20of=2027=20entries,=207=20archived?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- archival-memory.md | 316 ++++++++++++++++++ ...-on-mnt-smithy-approved-but-deferred-at.md | 3 + ...-searxngbraveapikey-in-searxng-settings.md | 3 - ...2026-09-18-fleet-dns-ring-and-ratelimit.md | 94 ------ .../2026-09-18-fleettools-agent-index.md | 57 ---- ...-plugin-install-is-now-a-symlink-to-the.md | 3 - .../2026-09-18-nh3-ana-derp-relay.md | 91 ----- .../2026-09-18-searxng-esh-egress-reverted.md | 47 --- ...-worldtree-s-env-sh-secrets-are-vaulted.md | 3 - ...ted-1-silently-and-the-alarm-carried-no.md | 3 + ...p-connect-as-the-is-blender-ready-probe.md | 3 + ...gateway-restarted-0401-for-highseat-dev.md | 3 + ...on-the-build-cache-surviving-on-esh-ml1.md | 3 + ...p-server-on-nh3-dev-reaching-the-add-on.md | 3 + ...t-requirements-txt-uv-pip-install-for-a.md | 3 + ...arakeet-slices-as-scriberr-s-memory-fix.md | 3 + ...webhook-allowed-host-list-external-10-0.md | 3 + ...ble-v2-auto-rigger-live-on-fv-ml1-gpu-3.md | 3 + ...1700-backups-dev-backup-retention-fixed.md | 3 + ...ench-leftovers-no-upstream-for-scriberr.md | 3 + ...-scans-to-pick-reclaimable-staging-on-a.md | 3 + ...-webhook-repointed-from-the-retired-wg0.md | 3 + ...-0-1-0-live-on-nh3-docker-albok-dev-ask.md | 3 + ...n-rolled-back-worldtree-dev-urgent-6684.md | 3 + ...future-esh-dev-host-onboarded-infra-ops.md | 3 + ...esh-nas-pve-175-t-raidz2-2-was-degraded.md | 3 + ...ot-grown-250-378-gb-prime-resized-scsi0.md | 3 + ...2-status-at-end-of-day-amt-done-proxmox.md | 3 + ...t-is-in-tacticalrmm-s-meshcentral-group.md | 3 + ...e-dev-backup-gets-dailies-weeklies-done.md | 3 + ...03s-get-proxmox-dummy-plugs-in-hand-one.md | 3 + ...-0-1-2-live-on-nh3-docker-albok-dev-ask.md | 3 + ...g-live-on-pfi-tacticalrmm-prime-go-1232.md | 3 + ...-finished-prime-on-site-host-10-100-250.md | 3 + ...ty-reframed-regression-check-not-a-gate.md | 3 + persistent-memory.md | 106 +++--- 36 files changed, 465 insertions(+), 336 deletions(-) create mode 100644 persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md delete mode 100644 persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md delete mode 100644 persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md delete mode 100644 persistent-memory.d/2026-09-18-fleettools-agent-index.md delete mode 100644 persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md delete mode 100644 persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md delete mode 100644 persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md delete mode 100644 persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md create mode 100644 persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md create mode 100644 persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md create mode 100644 persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md create mode 100644 persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md create mode 100644 persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md create mode 100644 persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md create mode 100644 persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md create mode 100644 persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md create mode 100644 persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md create mode 100644 persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md create mode 100644 persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md create mode 100644 persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md create mode 100644 persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md create mode 100644 persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md create mode 100644 persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md create mode 100644 persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md create mode 100644 persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md create mode 100644 persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md create mode 100644 persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md create mode 100644 persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md create mode 100644 persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md create mode 100644 persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md create mode 100644 persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md create mode 100644 persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md create mode 100644 persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md create mode 100644 persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md diff --git a/archival-memory.md b/archival-memory.md index a0afb95..c195933 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -13243,3 +13243,319 @@ ruled. heid also found and killed two live instructions in their own persistent- fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment after a context reset. _Archived 2026-10-01._ + +## Recent decisions (archived 2026-10-03 batch) + +# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it + +Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a +direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP +is a deliberately throttled fallback — a reachability mechanism, not a data plane — +so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM, +Beszel, task-board, vor and the Henge alike. + +## How it surfaced, and how it did NOT + +tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway +than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost. +They had already cleared GPU contention, the ASR itself and output quality with +real controls. + +**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks +whether a mesh path is direct. + +## The diagnosis, and two wrong causes published on the way + +Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt` +has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like +a great candidate for a fixed floor — env vars still set, container gone — but it was +retired from the callback list 2026-06-20 and nothing calls it per-request. + +The decisive arm was one neither of us had run: **measure from ana-docker itself**, so +both arms share the identical onward path to FV. + + ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s + ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s + +The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own +host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the +path AND which destination the packets crossed the network to reach — and the entire +delta was attributed to the variable of interest. + + nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s + nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s + + nh3-scale$ tailscale ping 100.64.0.3 + pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493 + direct connection not established + +## Why hole-punching failed + +ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the +Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or +NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT; +simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct; +only this pair failed. + +## The fix — four ADDITIVE objects on ana-gw (10.250.0.1) + +⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal +address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house +template; writing what a FortiGate normally wants would have failed. + +| object | value | +|---|---| +| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` | +| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` | +| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 | +| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept | + +## Measured before → after + +| path | before | after | +|---|---|---| +| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** | +| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** | +| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** | + +tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 → +548 ms, converged with their direct arm. + +## Operational notes for next time + +⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1 +'show' > file` captures the full non-default config (11,320 lines) without needing a +tftp server; kept at `~/backups/ana-gw-config-BEFORE-.txt`. Credentials vaulted +at `fortigate/ana-gw-infra-ops-password`, reached with +`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an +existing policy or interface over a 327 ms link you cannot recover. + +⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed +by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was. +I warned tts-dev their TTS stack was paying for it; they measured and I was wrong. + +Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See +[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable +cross-site DNS backup because of this fix. + _Archived 2026-10-03._ + +# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide + +Two independent faults, fixed in order. Together they were costing roughly one +`.internal` lookup in ten either a hard failure or a five-second stall, on every +DHCP client at every site. + +## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone + +DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal +resolver missed a packet, the resolver waited out its timeout, fell through to +Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a +definitive "no such host" rather than a retry. Binary failure: instant, or five +seconds then `gaierror`. + +**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and +forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at +all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own +`resolv.conf`. + +⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand +editing it survives exactly until the next lease renewal — tts-dev's original +framing was "fix nh3-dev's resolv.conf", which would have worked, been verified, +and then silently reverted. Change it at the source: + +- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object. +- **FortiGate**: `config system dhcp server`, `set dns-service specify` + + `dns-server1` / `dns-server2`. +- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease. + +**Operator's ruling was broader than the proposal**: each site's backup resolver +should be *another site's* resolver. ESH already worked this way; the other two +were set, and the ring closes. + + ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing] + NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed] + ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed] + +All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP +server (`sfsrv`) and the fortilink switch-management server were deliberately left +alone — tenant property. + +**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.** + +## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three + + ratelimit: 20 + ratelimit_subnet_len_ipv4: 24 + ratelimit_whitelist: [] + +**20 queries per second shared across an entire /24 — per subnet, not per client.** +Every host on a VLAN draws from one bucket, so one busy container starves the rest, +and anything over the line is **silently dropped**, costing the client its full 5 s +resolver timeout. + +Demonstrated rather than inferred — 60 concurrent queries at one resolver: + +| | before | after | +|---|---|---| +| median | 5044 ms | 72 ms | +| timed out (>4 s) | **40 of 60** | **0 of 60** | + +Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three +firewalls; that setting exists to blunt internet-side DNS amplification, which is +not this. + +⚠ **Restarting AdGuard is required for a config change and briefly drops that site's +primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a +live member. That the ring made this rollout safe is the point of having built it +first, an hour earlier. + +## End-to-end, the original symptom + +40 TCP connects to a service by name vs by IP; the by-IP arm is the control. + +| | original | after ring only | after ratelimit 0 | +|---|---|---|---| +| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** | +| by NAME >1 s | 3/40 | 5/40 | **0/40** | +| by NAME failed | 2/40 | 0/40 | **0/40** | +| by IP p90 (control) | 25 ms | — | 11.7 ms | + +tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the +median**. Resolution has stopped being a cost rather than become a smaller one. + +## What this says about monitoring + +Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor +today was a voice loop with a 1545 ms budget whose owner chose to measure instead +of adapting around the problem. Both faults were invisible to Beszel and Uptime +Kuma — every host was up, every service answered. + +Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible +cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same +afternoon. The two changes compose. + _Archived 2026-10-03._ + +# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families + +Every agent on nh3-dev needs the same answers — what runs here, how do I call it, +what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only +Claude sessions load, interleaved with operator preferences other families have no +use for. + +## Shape + +**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line +index an agent reads whole; each entry links to a detail file it opens only when it +needs that tool. Reading the index costs about a fifth of reading the tree. + +- Canonical: `docs/fleettools/` in this repo (git-tracked) +- Discovery: `~/FLEETTOOLS.md` → symlink to the index +- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent + `cat`s the path rather than following a link, and may be anywhere on the filesystem. + +Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine, +speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and +a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills). + +## Autoload — verified per family, not assumed + +- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string + `"Failed to read global AGENTS.md instructions from"`. +- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless + of name** — from its own embedded docs. + +Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than +copied — one file, three families, no drift surface. + +⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat +settings. That directory is empty today, but Grok can see Claude's config. + +## Rule zero, which is the load-bearing part + +**Query live inventories, never a written list.** The index points at +Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and +every FastAPI seat's `/openapi.json`. A copied service table would be stale within a +month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said +~30. + +## The CLAUDE.md swap + +`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39** +(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed +inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck` +before expensive work, post operator-facing links to the Booth board, and the vault is +the credential source of truth. **A rule behind a file read is a rule that stops +firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`. + +The shared all-agents LiteLLM key stays inline there too — every session needs it and a +vault round-trip measured over two minutes — but it is now ALSO vaulted at +`litellm/all-agents-shared-key`, since it had been single-copy. + +Commits `53c3e80` `21d24c5`. + _Archived 2026-10-03._ + +# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy + +**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart. + _Archived 2026-10-03._ + +# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted + +**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped. + _Archived 2026-10-03._ + +## Tried and abandoned (archived 2026-10-03 batch) + +# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing + +On 2026-09-17, at operator request, SearXNG's search egress was moved from +nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale** +(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change. +Reverted 2026-09-18. + +## Why it was reverted, and why the STATED reason was wrong + +It was reverted on a diagnosis that turned out to be false: that the move had cost +three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a +**byte-identical** result — same three engines down — which falsified it. A live +`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale +suspension timers too. + +The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]). + +**The revert still stands on its own merits**: ESH egress bought no measurable +improvement while making all fleet search depend on ESH WAN and mesh availability. +The simpler configuration is the better one. It is simply not the fix for the engines. + +## What was built, because it was built correctly + +The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran +as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed +source `10.100.50.40` alone — everything else had to supply a password regenerated at +each start and never distributed. Allowed-host egress and denied-host rejection were +both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the +service is **stopped and disabled** on esh-scale, and `tailscaled` there was not +touched (CT 108 is ESH's whole-site mesh SPOF). + +Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's +WAN address moves and nothing pins it, so the number in the stack README is historical. + +## The transferable lesson + +⚠ **I asserted causation from correlation with no baseline.** The only evidence that +residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was +no longer true of that address. Five samples of the post-change state and zero of the +working state is not a comparison. The rollback WAS the counterfactual, and it +falsified the claim I had already published in a commit message. + +This was the first of three wrong causal attributions in a single afternoon. The +common shape: measure after a change, attribute the delta to *my* change, never check +what else moved. + +Commits `156e126` (applied), `1a35181` (reverted). + _Archived 2026-10-03._ + +# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings + +**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing. + _Archived 2026-10-03._ diff --git a/persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md b/persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md new file mode 100644 index 0000000..136edfa --- /dev/null +++ b/persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md @@ -0,0 +1,3 @@ +# `nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction + +`[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`. diff --git a/persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md b/persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md deleted file mode 100644 index 8804eed..0000000 --- a/persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings - -**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing. diff --git a/persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md b/persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md deleted file mode 100644 index 64904b6..0000000 --- a/persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md +++ /dev/null @@ -1,94 +0,0 @@ -# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide - -Two independent faults, fixed in order. Together they were costing roughly one -`.internal` lookup in ten either a hard failure or a five-second stall, on every -DHCP client at every site. - -## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone - -DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal -resolver missed a packet, the resolver waited out its timeout, fell through to -Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a -definitive "no such host" rather than a retry. Binary failure: instant, or five -seconds then `gaierror`. - -**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and -forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at -all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own -`resolv.conf`. - -⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand -editing it survives exactly until the next lease renewal — tts-dev's original -framing was "fix nh3-dev's resolv.conf", which would have worked, been verified, -and then silently reverted. Change it at the source: - -- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object. -- **FortiGate**: `config system dhcp server`, `set dns-service specify` + - `dns-server1` / `dns-server2`. -- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease. - -**Operator's ruling was broader than the proposal**: each site's backup resolver -should be *another site's* resolver. ESH already worked this way; the other two -were set, and the ring closes. - - ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing] - NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed] - ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed] - -All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP -server (`sfsrv`) and the fortilink switch-management server were deliberately left -alone — tenant property. - -**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.** - -## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three - - ratelimit: 20 - ratelimit_subnet_len_ipv4: 24 - ratelimit_whitelist: [] - -**20 queries per second shared across an entire /24 — per subnet, not per client.** -Every host on a VLAN draws from one bucket, so one busy container starves the rest, -and anything over the line is **silently dropped**, costing the client its full 5 s -resolver timeout. - -Demonstrated rather than inferred — 60 concurrent queries at one resolver: - -| | before | after | -|---|---|---| -| median | 5044 ms | 72 ms | -| timed out (>4 s) | **40 of 60** | **0 of 60** | - -Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three -firewalls; that setting exists to blunt internet-side DNS amplification, which is -not this. - -⚠ **Restarting AdGuard is required for a config change and briefly drops that site's -primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a -live member. That the ring made this rollout safe is the point of having built it -first, an hour earlier. - -## End-to-end, the original symptom - -40 TCP connects to a service by name vs by IP; the by-IP arm is the control. - -| | original | after ring only | after ratelimit 0 | -|---|---|---|---| -| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** | -| by NAME >1 s | 3/40 | 5/40 | **0/40** | -| by NAME failed | 2/40 | 0/40 | **0/40** | -| by IP p90 (control) | 25 ms | — | 11.7 ms | - -tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the -median**. Resolution has stopped being a cost rather than become a smaller one. - -## What this says about monitoring - -Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor -today was a voice loop with a 1545 ms budget whose owner chose to measure instead -of adapting around the problem. Both faults were invisible to Beszel and Uptime -Kuma — every host was up, every service answered. - -Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible -cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same -afternoon. The two changes compose. diff --git a/persistent-memory.d/2026-09-18-fleettools-agent-index.md b/persistent-memory.d/2026-09-18-fleettools-agent-index.md deleted file mode 100644 index b1d345e..0000000 --- a/persistent-memory.d/2026-09-18-fleettools-agent-index.md +++ /dev/null @@ -1,57 +0,0 @@ -# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families - -Every agent on nh3-dev needs the same answers — what runs here, how do I call it, -what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only -Claude sessions load, interleaved with operator preferences other families have no -use for. - -## Shape - -**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line -index an agent reads whole; each entry links to a detail file it opens only when it -needs that tool. Reading the index costs about a fifth of reading the tree. - -- Canonical: `docs/fleettools/` in this repo (git-tracked) -- Discovery: `~/FLEETTOOLS.md` → symlink to the index -- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent - `cat`s the path rather than following a link, and may be anywhere on the filesystem. - -Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine, -speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and -a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills). - -## Autoload — verified per family, not assumed - -- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string - `"Failed to read global AGENTS.md instructions from"`. -- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless - of name** — from its own embedded docs. - -Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than -copied — one file, three families, no drift surface. - -⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat -settings. That directory is empty today, but Grok can see Claude's config. - -## Rule zero, which is the load-bearing part - -**Query live inventories, never a written list.** The index points at -Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and -every FastAPI seat's `/openapi.json`. A copied service table would be stale within a -month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said -~30. - -## The CLAUDE.md swap - -`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39** -(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed -inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck` -before expensive work, post operator-facing links to the Booth board, and the vault is -the credential source of truth. **A rule behind a file read is a rule that stops -firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`. - -The shared all-agents LiteLLM key stays inline there too — every session needs it and a -vault round-trip measured over two minutes — but it is now ALSO vaulted at -`litellm/all-agents-shared-key`, since it had been single-copy. - -Commits `53c3e80` `21d24c5`. diff --git a/persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md b/persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md deleted file mode 100644 index a502c7c..0000000 --- a/persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy - -**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart. diff --git a/persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md b/persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md deleted file mode 100644 index c2788ee..0000000 --- a/persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md +++ /dev/null @@ -1,91 +0,0 @@ -# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it - -Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a -direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP -is a deliberately throttled fallback — a reachability mechanism, not a data plane — -so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM, -Beszel, task-board, vor and the Henge alike. - -## How it surfaced, and how it did NOT - -tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway -than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost. -They had already cleared GPU contention, the ASR itself and output quality with -real controls. - -**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks -whether a mesh path is direct. - -## The diagnosis, and two wrong causes published on the way - -Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt` -has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like -a great candidate for a fixed floor — env vars still set, container gone — but it was -retired from the callback list 2026-06-20 and nothing calls it per-request. - -The decisive arm was one neither of us had run: **measure from ana-docker itself**, so -both arms share the identical onward path to FV. - - ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s - ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s - -The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own -host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the -path AND which destination the packets crossed the network to reach — and the entire -delta was attributed to the variable of interest. - - nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s - nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s - - nh3-scale$ tailscale ping 100.64.0.3 - pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493 - direct connection not established - -## Why hole-punching failed - -ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the -Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or -NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT; -simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct; -only this pair failed. - -## The fix — four ADDITIVE objects on ana-gw (10.250.0.1) - -⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal -address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house -template; writing what a FortiGate normally wants would have failed. - -| object | value | -|---|---| -| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` | -| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` | -| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 | -| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept | - -## Measured before → after - -| path | before | after | -|---|---|---| -| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** | -| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** | -| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** | - -tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 → -548 ms, converged with their direct arm. - -## Operational notes for next time - -⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1 -'show' > file` captures the full non-default config (11,320 lines) without needing a -tftp server; kept at `~/backups/ana-gw-config-BEFORE-.txt`. Credentials vaulted -at `fortigate/ana-gw-infra-ops-password`, reached with -`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an -existing policy or interface over a 327 ms link you cannot recover. - -⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed -by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was. -I warned tts-dev their TTS stack was paying for it; they measured and I was wrong. - -Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See -[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable -cross-site DNS backup because of this fix. diff --git a/persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md b/persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md deleted file mode 100644 index caed581..0000000 --- a/persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md +++ /dev/null @@ -1,47 +0,0 @@ -# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing - -On 2026-09-17, at operator request, SearXNG's search egress was moved from -nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale** -(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change. -Reverted 2026-09-18. - -## Why it was reverted, and why the STATED reason was wrong - -It was reverted on a diagnosis that turned out to be false: that the move had cost -three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a -**byte-identical** result — same three engines down — which falsified it. A live -`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale -suspension timers too. - -The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]). - -**The revert still stands on its own merits**: ESH egress bought no measurable -improvement while making all fleet search depend on ESH WAN and mesh availability. -The simpler configuration is the better one. It is simply not the fix for the engines. - -## What was built, because it was built correctly - -The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran -as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed -source `10.100.50.40` alone — everything else had to supply a password regenerated at -each start and never distributed. Allowed-host egress and denied-host rejection were -both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the -service is **stopped and disabled** on esh-scale, and `tailscaled` there was not -touched (CT 108 is ESH's whole-site mesh SPOF). - -Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's -WAN address moves and nothing pins it, so the number in the stack README is historical. - -## The transferable lesson - -⚠ **I asserted causation from correlation with no baseline.** The only evidence that -residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was -no longer true of that address. Five samples of the post-change state and zero of the -working state is not a comparison. The rollback WAS the counterfactual, and it -falsified the claim I had already published in a commit message. - -This was the first of three wrong causal attributions in a single afternoon. The -common shape: measure after a change, attribute the delta to *my* change, never check -what else moved. - -Commits `156e126` (applied), `1a35181` (reverted). diff --git a/persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md b/persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md deleted file mode 100644 index 443158e..0000000 --- a/persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted - -**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped. diff --git a/persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md b/persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md new file mode 100644 index 0000000..cb704b0 --- /dev/null +++ b/persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md @@ -0,0 +1,3 @@ +# `headscale-ddns` exited 1 silently and the alarm carried no cause + +`[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`. diff --git a/persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md b/persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md new file mode 100644 index 0000000..bd06705 --- /dev/null +++ b/persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md @@ -0,0 +1,3 @@ +# A TCP connect as the "is Blender ready" probe + +`[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port. diff --git a/persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md b/persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md new file mode 100644 index 0000000..fc6b297 --- /dev/null +++ b/persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md @@ -0,0 +1,3 @@ +# hermes-gateway restarted 0401 for highseat-dev + +`[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call. diff --git a/persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md b/persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md new file mode 100644 index 0000000..9642736 --- /dev/null +++ b/persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md @@ -0,0 +1,3 @@ +# Relying on the build cache surviving on esh-ml1 + +`[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there. diff --git a/persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md b/persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md new file mode 100644 index 0000000..a9f5a76 --- /dev/null +++ b/persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md @@ -0,0 +1,3 @@ +# The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel + +`[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md` diff --git a/persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md b/persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md new file mode 100644 index 0000000..75ab242 --- /dev/null +++ b/persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md @@ -0,0 +1,3 @@ +# `uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project + +`[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`). diff --git a/persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md b/persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md new file mode 100644 index 0000000..987edcf --- /dev/null +++ b/persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md @@ -0,0 +1,3 @@ +# Shorter Parakeet slices as Scriberr's memory fix + +`[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor. diff --git a/persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md b/persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md new file mode 100644 index 0000000..dee0f0d --- /dev/null +++ b/persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md @@ -0,0 +1,3 @@ +# ⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10) + +`[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them. diff --git a/persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md b/persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md new file mode 100644 index 0000000..8343335 --- /dev/null +++ b/persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md @@ -0,0 +1,3 @@ +# MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev) + +`[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md` diff --git a/persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md b/persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md new file mode 100644 index 0000000..d86b97e --- /dev/null +++ b/persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md @@ -0,0 +1,3 @@ +# Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed + +`[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it. diff --git a/persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md b/persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md new file mode 100644 index 0000000..d39e2a2 --- /dev/null +++ b/persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md @@ -0,0 +1,3 @@ +# Prime: delete the bench leftovers, no upstream for Scriberr, push + +`[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. diff --git a/persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md b/persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md new file mode 100644 index 0000000..4cad7be --- /dev/null +++ b/persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md @@ -0,0 +1,3 @@ +# Size + mtime scans to pick reclaimable "staging" on a model store + +`[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it. diff --git a/persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md b/persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md new file mode 100644 index 0000000..fd1d768 --- /dev/null +++ b/persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md @@ -0,0 +1,3 @@ +# vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009` + +`[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it. diff --git a/persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md b/persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md new file mode 100644 index 0000000..fdcb0ae --- /dev/null +++ b/persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md @@ -0,0 +1,3 @@ +# albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392` + +`[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md` diff --git a/persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md b/persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md new file mode 100644 index 0000000..88668df --- /dev/null +++ b/persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md @@ -0,0 +1,3 @@ +# Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684) + +`[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423. diff --git a/persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md b/persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md new file mode 100644 index 0000000..445369e --- /dev/null +++ b/persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md @@ -0,0 +1,3 @@ +# esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt` + +`[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md` diff --git a/persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md b/persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md new file mode 100644 index 0000000..a8ee646 --- /dev/null +++ b/persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md @@ -0,0 +1,3 @@ +# ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629 + +`[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package. diff --git a/persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md b/persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md new file mode 100644 index 0000000..cf9893a --- /dev/null +++ b/persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md @@ -0,0 +1,3 @@ +# nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot) + +`[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`. diff --git a/persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md b/persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md new file mode 100644 index 0000000..c563976 --- /dev/null +++ b/persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md @@ -0,0 +1,3 @@ +# nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03 + +`[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug. diff --git a/persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md b/persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md new file mode 100644 index 0000000..ed7a90b --- /dev/null +++ b/persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md @@ -0,0 +1,3 @@ +# nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`) + +`[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md` diff --git a/persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md b/persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md new file mode 100644 index 0000000..9433601 --- /dev/null +++ b/persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md @@ -0,0 +1,3 @@ +# Prime: dev-backup gets dailies + weeklies — DONE + +`[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today. diff --git a/persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md b/persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md new file mode 100644 index 0000000..4e4f7bb --- /dev/null +++ b/persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md @@ -0,0 +1,3 @@ +# Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910) + +`[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups. diff --git a/persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md b/persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md new file mode 100644 index 0000000..e309fdd --- /dev/null +++ b/persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md @@ -0,0 +1,3 @@ +# albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1 + +`[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md` diff --git a/persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md b/persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md new file mode 100644 index 0000000..57a801f --- /dev/null +++ b/persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md @@ -0,0 +1,3 @@ +# CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232) + +`[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235. diff --git a/persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md b/persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md new file mode 100644 index 0000000..a1588e3 --- /dev/null +++ b/persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md @@ -0,0 +1,3 @@ +# nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone + +`[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md` diff --git a/persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md b/persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md new file mode 100644 index 0000000..c5e3cf6 --- /dev/null +++ b/persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md @@ -0,0 +1,3 @@ +# Worldtree memory-gate duty REFRAMED: regression check, not a gate + +`[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day. diff --git a/persistent-memory.md b/persistent-memory.md index 6cb333e..77eb583 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-10-01 ~0446 PT (gen-small OOM fixed by parakeet-nemo 0.1.1, with a hard cap and cache return; Prime: leave the gen-small KV and graphs-off as is; arbo webhook fixed (Gitea allow-list); leftovers deleted, repos pushed. Prior: Parakeet seat → unified-en LIVE, U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_ +_Last updated: 2026-10-03 ~1335 PT (nh3-pve PVE 8→9 upgrade by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -115,7 +115,43 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-10-01 ~0446 PT._ +_As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections carry their own dates._ + +### 2026-10-03: live now (as of ~1335 PT) + +- **nh3-pve PVE 8.4.1 → 9 upgrade: Prime is running it at the console now.** Steps given to him: + - latest 8.4, then `pve8to9 --full`; + - `apt remove systemd-boot` (safe: it boots GRUB via proxmox-boot-tool); + - `zfs snapshot -r rpool/ROOT@pre-pve9`; + - bookworm → trixie in `/etc/apt/sources.list`, then `apt dist-upgrade`; + - **before the reboot, `dkms status` must show nvidia 580.178.04 built for the NEW kernel** (nh3-ml1's GPU); + - reboot. + NIC names are already pinned (`/etc/systemd/network/10-pin-*.link`). **nh3-dev (VM 102, this session's host) goes down with it.** +- **Rollback for nh3-pve:** the ZFS snapshot, from a rescue boot. Never run `zpool upgrade rpool`. +- **POST-CHECK owed once nh3-pve and nh3-dev are back:** + - every guest is running: VMs 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3 (104 nh3-laser and 108 opnsense-lab were already stopped); CTs 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale, 109 nh3-ml1; + - NH3 DNS (AdGuard on nh3-docker 10.100.50.40); + - the althing post office (:8390) and albok-service (:8392); + - the mesh (headscale, nh3-scale); + - Miranda's channel (svos :8770 + hermes-gateway on nh3-dev); + - nh3-ml1's GPU (`nvidia-smi` in CT 109); + - `nh3-pve-amt` connected in MeshCentral; + - Beszel; + - `pveversion` shows 9.x. +- **nh3-pve-2** (MS-03, NH3) is live at 10.100.250.62 on nh3-mgmt: + - X710 `nic3` on USW port 25; + - AMT `nic0` with `auto` + `/etc/sysctl.d/90-amt-port.conf`; + - vmstore 913 GiB; + - subscription popup patched. + Open and NOT approved yet (Prime's call): disable the ceph enterprise apt source, apt upgrade, and one reboot to prove `.62` and nic0 survive. → `servers/nh3-pve-2/README.md` +- **esh-pve-2** (MS-03, ESH) is unplugged by Prime on purpose. vmstore + AMT phone-home are done. The permanent address is deferred by Prime ("address can wait"; `servers/esh-pve-2/README.md`). Run `playbooks/pve-nag-patch.yaml` on it when it is back. +- **cira-tunnel-watchdog** on pfi-tacticalrmm has been live since 1232 and drops :4433 tunnels silent >180 s. Watch `journalctl -t cira-tunnel-watchdog`. +- **Done today, nothing open:** + - tank verified healthy 0629, disk kept, Miranda told; + - albok-service 0.1.2 + the Nemi hourly timer; + - the post-deletion Worldtree gate (infra-hermes runs it as a regression check); + - the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve. +- **ESH PVE 8→9 plan:** written, NOT executed (needs Prime's green light). `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. The tank blocker is cleared. ### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01) @@ -287,25 +323,26 @@ _As of 2026-10-01 ~0446 PT._ ## Recent decisions -- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md` -- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235. -- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day. -- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md` -- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md` -- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423. -- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md` -- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today. -- `[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package. -- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md` -- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug. -- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups. -- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`. -- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md` -- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it. +- `[2026-10-03]` **PVE subscription popup patched out on nh3-pve, nh3-pve-2, pfi-pve, esh-pve (Prime: "every host")**: an anchored one-line edit plus an apt hook. esh-nas-pve was already de-nagged; esh-pve-2, sfsrv-ana and the PBS hosts are not done. → `services/pve-nag-patch/README.md` +- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone** → `persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md` +- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** → `persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md` +- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** → `persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md` +- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1** → `persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md` +- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`** → `persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md` +- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)** → `persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md` +- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)** → `persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md` +- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE** → `persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md` +- `[2026-10-02]` **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629** → `persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md` +- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** → `persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md` +- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03** → `persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md` +- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)** → `persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md` +- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)** → `persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md` +- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)** → `persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md` +- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed** → `persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md` - `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap. -- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them. -- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it. -- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. +- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)** → `persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md` +- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`** → `persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md` +- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push** → `persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md` - `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`). - `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` - `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` @@ -316,7 +353,7 @@ _As of 2026-10-01 ~0446 PT._ - `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md` - `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3. - `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions -- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call. +- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** → `persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md` - `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md` - `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md` - `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md` @@ -370,7 +407,7 @@ _As of 2026-10-01 ~0446 PT._ - `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md` - `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md` - `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md` -- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`. +- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** → `persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md` - `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md` - `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md` @@ -418,13 +455,8 @@ _As of 2026-10-01 ~0446 PT._ - `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md` -- `[2026-09-18]` ⭐⭐⭐ **NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure.** Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; `tailscale ping` 373–522 ms → **6 ms direct**, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs `central-nat`, so a policy `dstaddr` is the REAL internal address, not the VIP. No OOB access — back up with `show` to a local file and make additive changes ONLY. irv-ml1 still relayed. → `persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md` -- `[2026-09-18]` ⭐⭐⭐ **`.internal` DNS was failing ~10% of lookups fleet-wide, from two independent causes.** A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped `ratelimit: 20` shared across an entire **/24**, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ `resolv.conf` is DHCP-managed — change it at the UDM/FortiGate, not the file. → `persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md` - `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md` - `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md` -- `[2026-09-18]` ⭐ **FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.** `docs/fleettools/` + `~/FLEETTOOLS.md`; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → `persistent-memory.d/2026-09-18-fleettools-agent-index.md` -- `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md` -- `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md` - `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md` ⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24). @@ -445,26 +477,26 @@ _As of 2026-10-01 ~0446 PT._ - `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md` -- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`. +- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction** → `persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md` -_167 older entries archived to archival-memory.md._ +_172 older entries archived to archival-memory.md._ ## Tried and abandoned -- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it. -- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor. +- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** → `persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md` +- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix** → `persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md` - `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run. - `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead. - `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest. - `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`. -- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md` -- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port. +- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel** → `persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md` +- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe** → `persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md` - `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`). - `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra. -- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`). +- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** → `persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md` - `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`. - `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`. -- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there. +- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** → `persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md` - `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md` - `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`. - `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`. @@ -475,8 +507,6 @@ _167 older entries archived to archival-memory.md._ - `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md` -- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md` -- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md` -_118 older entries archived to archival-memory.md._ +_120 older entries archived to archival-memory.md._