memory: snapshot — 2026-10-03 in-flight (nh3-pve 8→9 under way), two-tier split of 27 entries, 7 archived

This commit is contained in:
vh
2026-10-03 13:36:42 -07:00
parent 4b2a81bf39
commit 5abd43e1c8
36 changed files with 465 additions and 336 deletions
+316
View File
@@ -13243,3 +13243,319 @@ ruled. heid also found and killed two live instructions in their own persistent-
fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment
after a context reset.
_Archived 2026-10-01._
## Recent decisions (archived 2026-10-03 batch)
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
Beszel, task-board, vor and the Henge alike.
## How it surfaced, and how it did NOT
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
They had already cleared GPU contention, the ASR itself and output quality with
real controls.
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
whether a mesh path is direct.
## The diagnosis, and two wrong causes published on the way
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
a great candidate for a fixed floor — env vars still set, container gone — but it was
retired from the callback list 2026-06-20 and nothing calls it per-request.
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
both arms share the identical onward path to FV.
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
path AND which destination the packets crossed the network to reach — and the entire
delta was attributed to the variable of interest.
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
nh3-scale$ tailscale ping 100.64.0.3
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
direct connection not established
## Why hole-punching failed
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
only this pair failed.
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
template; writing what a FortiGate normally wants would have failed.
| object | value |
|---|---|
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
## Measured before → after
| path | before | after |
|---|---|---|
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
548 ms, converged with their direct arm.
## Operational notes for next time
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
'show' > file` captures the full non-default config (11,320 lines) without needing a
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
at `fortigate/ana-gw-infra-ops-password`, reached with
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
existing policy or interface over a 327 ms link you cannot recover.
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
cross-site DNS backup because of this fix.
_Archived 2026-10-03._
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
Two independent faults, fixed in order. Together they were costing roughly one
`.internal` lookup in ten either a hard failure or a five-second stall, on every
DHCP client at every site.
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
resolver missed a packet, the resolver waited out its timeout, fell through to
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
definitive "no such host" rather than a retry. Binary failure: instant, or five
seconds then `gaierror`.
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
`resolv.conf`.
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
editing it survives exactly until the next lease renewal — tts-dev's original
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
and then silently reverted. Change it at the source:
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
`dns-server1` / `dns-server2`.
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
**Operator's ruling was broader than the proposal**: each site's backup resolver
should be *another site's* resolver. ESH already worked this way; the other two
were set, and the ring closes.
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
server (`sfsrv`) and the fortilink switch-management server were deliberately left
alone — tenant property.
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
ratelimit: 20
ratelimit_subnet_len_ipv4: 24
ratelimit_whitelist: []
**20 queries per second shared across an entire /24 — per subnet, not per client.**
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
and anything over the line is **silently dropped**, costing the client its full 5 s
resolver timeout.
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
| | before | after |
|---|---|---|
| median | 5044 ms | 72 ms |
| timed out (>4 s) | **40 of 60** | **0 of 60** |
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
firewalls; that setting exists to blunt internet-side DNS amplification, which is
not this.
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
live member. That the ring made this rollout safe is the point of having built it
first, an hour earlier.
## End-to-end, the original symptom
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
| | original | after ring only | after ratelimit 0 |
|---|---|---|---|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
| by NAME failed | 2/40 | 0/40 | **0/40** |
| by IP p90 (control) | 25 ms | — | 11.7 ms |
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
median**. Resolution has stopped being a cost rather than become a smaller one.
## What this says about monitoring
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
of adapting around the problem. Both faults were invisible to Beszel and Uptime
Kuma — every host was up, every service answered.
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
afternoon. The two changes compose.
_Archived 2026-10-03._
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
Claude sessions load, interleaved with operator preferences other families have no
use for.
## Shape
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
index an agent reads whole; each entry links to a detail file it opens only when it
needs that tool. Reading the index costs about a fifth of reading the tree.
- Canonical: `docs/fleettools/` in this repo (git-tracked)
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
## Autoload — verified per family, not assumed
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
`"Failed to read global AGENTS.md instructions from"`.
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
of name** — from its own embedded docs.
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
copied — one file, three families, no drift surface.
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
settings. That directory is empty today, but Grok can see Claude's config.
## Rule zero, which is the load-bearing part
**Query live inventories, never a written list.** The index points at
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
~30.
## The CLAUDE.md swap
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
before expensive work, post operator-facing links to the Booth board, and the vault is
the credential source of truth. **A rule behind a file read is a rule that stops
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
vault round-trip measured over two minutes — but it is now ALSO vaulted at
`litellm/all-agents-shared-key`, since it had been single-copy.
Commits `53c3e80` `21d24c5`.
_Archived 2026-10-03._
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
_Archived 2026-10-03._
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
_Archived 2026-10-03._
## Tried and abandoned (archived 2026-10-03 batch)
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
On 2026-09-17, at operator request, SearXNG's search egress was moved from
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
Reverted 2026-09-18.
## Why it was reverted, and why the STATED reason was wrong
It was reverted on a diagnosis that turned out to be false: that the move had cost
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
**byte-identical** result — same three engines down — which falsified it. A live
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
suspension timers too.
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
**The revert still stands on its own merits**: ESH egress bought no measurable
improvement while making all fleet search depend on ESH WAN and mesh availability.
The simpler configuration is the better one. It is simply not the fix for the engines.
## What was built, because it was built correctly
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
source `10.100.50.40` alone — everything else had to supply a password regenerated at
each start and never distributed. Allowed-host egress and denied-host rejection were
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
touched (CT 108 is ESH's whole-site mesh SPOF).
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
WAN address moves and nothing pins it, so the number in the stack README is historical.
## The transferable lesson
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
no longer true of that address. Five samples of the post-change state and zero of the
working state is not a comparison. The rollback WAS the counterfactual, and it
falsified the claim I had already published in a commit message.
This was the first of three wrong causal attributions in a single afternoon. The
common shape: measure after a change, attribute the delta to *my* change, never check
what else moved.
Commits `156e126` (applied), `1a35181` (reverted).
_Archived 2026-10-03._
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
_Archived 2026-10-03._
@@ -0,0 +1,3 @@
# `nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction
`[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
@@ -1,3 +0,0 @@
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
@@ -1,94 +0,0 @@
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
Two independent faults, fixed in order. Together they were costing roughly one
`.internal` lookup in ten either a hard failure or a five-second stall, on every
DHCP client at every site.
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
resolver missed a packet, the resolver waited out its timeout, fell through to
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
definitive "no such host" rather than a retry. Binary failure: instant, or five
seconds then `gaierror`.
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
`resolv.conf`.
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
editing it survives exactly until the next lease renewal — tts-dev's original
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
and then silently reverted. Change it at the source:
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
`dns-server1` / `dns-server2`.
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
**Operator's ruling was broader than the proposal**: each site's backup resolver
should be *another site's* resolver. ESH already worked this way; the other two
were set, and the ring closes.
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
server (`sfsrv`) and the fortilink switch-management server were deliberately left
alone — tenant property.
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
ratelimit: 20
ratelimit_subnet_len_ipv4: 24
ratelimit_whitelist: []
**20 queries per second shared across an entire /24 — per subnet, not per client.**
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
and anything over the line is **silently dropped**, costing the client its full 5 s
resolver timeout.
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
| | before | after |
|---|---|---|
| median | 5044 ms | 72 ms |
| timed out (>4 s) | **40 of 60** | **0 of 60** |
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
firewalls; that setting exists to blunt internet-side DNS amplification, which is
not this.
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
live member. That the ring made this rollout safe is the point of having built it
first, an hour earlier.
## End-to-end, the original symptom
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
| | original | after ring only | after ratelimit 0 |
|---|---|---|---|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
| by NAME failed | 2/40 | 0/40 | **0/40** |
| by IP p90 (control) | 25 ms | — | 11.7 ms |
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
median**. Resolution has stopped being a cost rather than become a smaller one.
## What this says about monitoring
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
of adapting around the problem. Both faults were invisible to Beszel and Uptime
Kuma — every host was up, every service answered.
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
afternoon. The two changes compose.
@@ -1,57 +0,0 @@
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
Claude sessions load, interleaved with operator preferences other families have no
use for.
## Shape
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
index an agent reads whole; each entry links to a detail file it opens only when it
needs that tool. Reading the index costs about a fifth of reading the tree.
- Canonical: `docs/fleettools/` in this repo (git-tracked)
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
## Autoload — verified per family, not assumed
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
`"Failed to read global AGENTS.md instructions from"`.
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
of name** — from its own embedded docs.
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
copied — one file, three families, no drift surface.
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
settings. That directory is empty today, but Grok can see Claude's config.
## Rule zero, which is the load-bearing part
**Query live inventories, never a written list.** The index points at
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
~30.
## The CLAUDE.md swap
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
before expensive work, post operator-facing links to the Booth board, and the vault is
the credential source of truth. **A rule behind a file read is a rule that stops
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
vault round-trip measured over two minutes — but it is now ALSO vaulted at
`litellm/all-agents-shared-key`, since it had been single-copy.
Commits `53c3e80` `21d24c5`.
@@ -1,3 +0,0 @@
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
@@ -1,91 +0,0 @@
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
Beszel, task-board, vor and the Henge alike.
## How it surfaced, and how it did NOT
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
They had already cleared GPU contention, the ASR itself and output quality with
real controls.
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
whether a mesh path is direct.
## The diagnosis, and two wrong causes published on the way
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
a great candidate for a fixed floor — env vars still set, container gone — but it was
retired from the callback list 2026-06-20 and nothing calls it per-request.
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
both arms share the identical onward path to FV.
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
path AND which destination the packets crossed the network to reach — and the entire
delta was attributed to the variable of interest.
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
nh3-scale$ tailscale ping 100.64.0.3
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
direct connection not established
## Why hole-punching failed
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
only this pair failed.
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
template; writing what a FortiGate normally wants would have failed.
| object | value |
|---|---|
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
## Measured before → after
| path | before | after |
|---|---|---|
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
548 ms, converged with their direct arm.
## Operational notes for next time
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
'show' > file` captures the full non-default config (11,320 lines) without needing a
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
at `fortigate/ana-gw-infra-ops-password`, reached with
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
existing policy or interface over a 327 ms link you cannot recover.
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
cross-site DNS backup because of this fix.
@@ -1,47 +0,0 @@
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
On 2026-09-17, at operator request, SearXNG's search egress was moved from
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
Reverted 2026-09-18.
## Why it was reverted, and why the STATED reason was wrong
It was reverted on a diagnosis that turned out to be false: that the move had cost
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
**byte-identical** result — same three engines down — which falsified it. A live
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
suspension timers too.
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
**The revert still stands on its own merits**: ESH egress bought no measurable
improvement while making all fleet search depend on ESH WAN and mesh availability.
The simpler configuration is the better one. It is simply not the fix for the engines.
## What was built, because it was built correctly
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
source `10.100.50.40` alone — everything else had to supply a password regenerated at
each start and never distributed. Allowed-host egress and denied-host rejection were
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
touched (CT 108 is ESH's whole-site mesh SPOF).
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
WAN address moves and nothing pins it, so the number in the stack README is historical.
## The transferable lesson
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
no longer true of that address. Five samples of the post-change state and zero of the
working state is not a comparison. The rollback WAS the counterfactual, and it
falsified the claim I had already published in a commit message.
This was the first of three wrong causal attributions in a single afternoon. The
common shape: measure after a change, attribute the delta to *my* change, never check
what else moved.
Commits `156e126` (applied), `1a35181` (reverted).
@@ -1,3 +0,0 @@
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
@@ -0,0 +1,3 @@
# `headscale-ddns` exited 1 silently and the alarm carried no cause
`[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
@@ -0,0 +1,3 @@
# A TCP connect as the "is Blender ready" probe
`[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
@@ -0,0 +1,3 @@
# hermes-gateway restarted 0401 for highseat-dev
`[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
@@ -0,0 +1,3 @@
# Relying on the build cache surviving on esh-ml1
`[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
@@ -0,0 +1,3 @@
# The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel
`[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
@@ -0,0 +1,3 @@
# `uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project
`[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
@@ -0,0 +1,3 @@
# Shorter Parakeet slices as Scriberr's memory fix
`[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
@@ -0,0 +1,3 @@
# ⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)
`[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
@@ -0,0 +1,3 @@
# MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)
`[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
@@ -0,0 +1,3 @@
# Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed
`[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
@@ -0,0 +1,3 @@
# Prime: delete the bench leftovers, no upstream for Scriberr, push
`[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
@@ -0,0 +1,3 @@
# Size + mtime scans to pick reclaimable "staging" on a model store
`[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
@@ -0,0 +1,3 @@
# vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`
`[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
@@ -0,0 +1,3 @@
# albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`
`[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
@@ -0,0 +1,3 @@
# Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)
`[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
@@ -0,0 +1,3 @@
# esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`
`[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
@@ -0,0 +1,3 @@
# ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629
`[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
@@ -0,0 +1,3 @@
# nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)
`[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
@@ -0,0 +1,3 @@
# nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03
`[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
@@ -0,0 +1,3 @@
# nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)
`[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
@@ -0,0 +1,3 @@
# Prime: dev-backup gets dailies + weeklies — DONE
`[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
@@ -0,0 +1,3 @@
# Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)
`[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
@@ -0,0 +1,3 @@
# albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1
`[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
@@ -0,0 +1,3 @@
# CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)
`[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
@@ -0,0 +1,3 @@
# nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone
`[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
@@ -0,0 +1,3 @@
# Worldtree memory-gate duty REFRAMED: regression check, not a gate
`[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
+68 -38
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-10-01 ~0446 PT (gen-small OOM fixed by parakeet-nemo 0.1.1, with a hard cap and cache return; Prime: leave the gen-small KV and graphs-off as is; arbo webhook fixed (Gitea allow-list); leftovers deleted, repos pushed. Prior: Parakeet seat → unified-en LIVE, U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
_Last updated: 2026-10-03 ~1335 PT (nh3-pve PVE 8→9 upgrade by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,7 +115,43 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-10-01 ~0446 PT._
_As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections carry their own dates._
### 2026-10-03: live now (as of ~1335 PT)
- **nh3-pve PVE 8.4.1 → 9 upgrade: Prime is running it at the console now.** Steps given to him:
- latest 8.4, then `pve8to9 --full`;
- `apt remove systemd-boot` (safe: it boots GRUB via proxmox-boot-tool);
- `zfs snapshot -r rpool/ROOT@pre-pve9`;
- bookworm → trixie in `/etc/apt/sources.list`, then `apt dist-upgrade`;
- **before the reboot, `dkms status` must show nvidia 580.178.04 built for the NEW kernel** (nh3-ml1's GPU);
- reboot.
NIC names are already pinned (`/etc/systemd/network/10-pin-*.link`). **nh3-dev (VM 102, this session's host) goes down with it.**
- **Rollback for nh3-pve:** the ZFS snapshot, from a rescue boot. Never run `zpool upgrade rpool`.
- **POST-CHECK owed once nh3-pve and nh3-dev are back:**
- every guest is running: VMs 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3 (104 nh3-laser and 108 opnsense-lab were already stopped); CTs 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale, 109 nh3-ml1;
- NH3 DNS (AdGuard on nh3-docker 10.100.50.40);
- the althing post office (:8390) and albok-service (:8392);
- the mesh (headscale, nh3-scale);
- Miranda's channel (svos :8770 + hermes-gateway on nh3-dev);
- nh3-ml1's GPU (`nvidia-smi` in CT 109);
- `nh3-pve-amt` connected in MeshCentral;
- Beszel;
- `pveversion` shows 9.x.
- **nh3-pve-2** (MS-03, NH3) is live at 10.100.250.62 on nh3-mgmt:
- X710 `nic3` on USW port 25;
- AMT `nic0` with `auto` + `/etc/sysctl.d/90-amt-port.conf`;
- vmstore 913 GiB;
- subscription popup patched.
Open and NOT approved yet (Prime's call): disable the ceph enterprise apt source, apt upgrade, and one reboot to prove `.62` and nic0 survive. → `servers/nh3-pve-2/README.md`
- **esh-pve-2** (MS-03, ESH) is unplugged by Prime on purpose. vmstore + AMT phone-home are done. The permanent address is deferred by Prime ("address can wait"; `servers/esh-pve-2/README.md`). Run `playbooks/pve-nag-patch.yaml` on it when it is back.
- **cira-tunnel-watchdog** on pfi-tacticalrmm has been live since 1232 and drops :4433 tunnels silent >180 s. Watch `journalctl -t cira-tunnel-watchdog`.
- **Done today, nothing open:**
- tank verified healthy 0629, disk kept, Miranda told;
- albok-service 0.1.2 + the Nemi hourly timer;
- the post-deletion Worldtree gate (infra-hermes runs it as a regression check);
- the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve.
- **ESH PVE 8→9 plan:** written, NOT executed (needs Prime's green light). `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. The tank blocker is cleared.
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
@@ -287,25 +323,26 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions
- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
- `[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
- `[2026-10-03]` **PVE subscription popup patched out on nh3-pve, nh3-pve-2, pfi-pve, esh-pve (Prime: "every host")**: an anchored one-line edit plus an apt hook. esh-nas-pve was already de-nagged; esh-pve-2, sfsrv-ana and the PBS hosts are not done. → `services/pve-nag-patch/README.md`
- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone** → `persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md`
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** → `persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md`
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** → `persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md`
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1** → `persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md`
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`** → `persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md`
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)** → `persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md`
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)** → `persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md`
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE** → `persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md`
- `[2026-10-02]` **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629** → `persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md`
- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** → `persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md`
- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03** → `persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md`
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)** → `persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md`
- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)** → `persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md`
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)** → `persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md`
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed** → `persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md`
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)** → `persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md`
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`** → `persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md`
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push** → `persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md`
- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
@@ -316,7 +353,7 @@ _As of 2026-10-01 ~0446 PT._
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
- `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3.
- `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** → `persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md`
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
- `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md`
- `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md`
@@ -370,7 +407,7 @@ _As of 2026-10-01 ~0446 PT._
- `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md`
- `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md`
- `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md`
- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** → `persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md`
- `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md`
- `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md`
@@ -418,13 +455,8 @@ _As of 2026-10-01 ~0446 PT._
- `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md`
- `[2026-09-18]` ⭐⭐⭐ **NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure.** Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; `tailscale ping` 373–522 ms → **6 ms direct**, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs `central-nat`, so a policy `dstaddr` is the REAL internal address, not the VIP. No OOB access — back up with `show` to a local file and make additive changes ONLY. irv-ml1 still relayed. → `persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md`
- `[2026-09-18]` ⭐⭐⭐ **`.internal` DNS was failing ~10% of lookups fleet-wide, from two independent causes.** A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped `ratelimit: 20` shared across an entire **/24**, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ `resolv.conf` is DHCP-managed — change it at the UDM/FortiGate, not the file. → `persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md`
- `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md`
- `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md`
- `[2026-09-18]` ⭐ **FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.** `docs/fleettools/` + `~/FLEETTOOLS.md`; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → `persistent-memory.d/2026-09-18-fleettools-agent-index.md`
- `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md`
- `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md`
- `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md`
⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).
@@ -445,26 +477,26 @@ _As of 2026-10-01 ~0446 PT._
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md`
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction** → `persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md`
_167 older entries archived to archival-memory.md._
_172 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** → `persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md`
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix** → `persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md`
- `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.
- `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.
- `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.
- `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`.
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel** → `persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md`
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe** → `persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md`
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
- `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** → `persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md`
- `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`.
- `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`.
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** → `persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md`
- `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md`
- `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`.
- `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`.
@@ -475,8 +507,6 @@ _167 older entries archived to archival-memory.md._
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md`
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`
- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md`
_118 older entries archived to archival-memory.md._
_120 older entries archived to archival-memory.md._