memory: snapshot — 2026-10-03 in-flight (nh3-pve 8→9 under way), two-tier split of 27 entries, 7 archived

This commit is contained in:
vh
2026-10-03 13:36:42 -07:00
parent 4b2a81bf39
commit 5abd43e1c8
36 changed files with 465 additions and 336 deletions
@@ -0,0 +1,3 @@
# `nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction
`[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
@@ -1,3 +0,0 @@
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
@@ -1,94 +0,0 @@
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
Two independent faults, fixed in order. Together they were costing roughly one
`.internal` lookup in ten either a hard failure or a five-second stall, on every
DHCP client at every site.
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
resolver missed a packet, the resolver waited out its timeout, fell through to
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
definitive "no such host" rather than a retry. Binary failure: instant, or five
seconds then `gaierror`.
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
`resolv.conf`.
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
editing it survives exactly until the next lease renewal — tts-dev's original
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
and then silently reverted. Change it at the source:
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
`dns-server1` / `dns-server2`.
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
**Operator's ruling was broader than the proposal**: each site's backup resolver
should be *another site's* resolver. ESH already worked this way; the other two
were set, and the ring closes.
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
server (`sfsrv`) and the fortilink switch-management server were deliberately left
alone — tenant property.
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
ratelimit: 20
ratelimit_subnet_len_ipv4: 24
ratelimit_whitelist: []
**20 queries per second shared across an entire /24 — per subnet, not per client.**
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
and anything over the line is **silently dropped**, costing the client its full 5 s
resolver timeout.
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
| | before | after |
|---|---|---|
| median | 5044 ms | 72 ms |
| timed out (>4 s) | **40 of 60** | **0 of 60** |
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
firewalls; that setting exists to blunt internet-side DNS amplification, which is
not this.
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
live member. That the ring made this rollout safe is the point of having built it
first, an hour earlier.
## End-to-end, the original symptom
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
| | original | after ring only | after ratelimit 0 |
|---|---|---|---|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
| by NAME failed | 2/40 | 0/40 | **0/40** |
| by IP p90 (control) | 25 ms | — | 11.7 ms |
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
median**. Resolution has stopped being a cost rather than become a smaller one.
## What this says about monitoring
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
of adapting around the problem. Both faults were invisible to Beszel and Uptime
Kuma — every host was up, every service answered.
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
afternoon. The two changes compose.
@@ -1,57 +0,0 @@
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
Claude sessions load, interleaved with operator preferences other families have no
use for.
## Shape
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
index an agent reads whole; each entry links to a detail file it opens only when it
needs that tool. Reading the index costs about a fifth of reading the tree.
- Canonical: `docs/fleettools/` in this repo (git-tracked)
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
## Autoload — verified per family, not assumed
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
`"Failed to read global AGENTS.md instructions from"`.
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
of name** — from its own embedded docs.
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
copied — one file, three families, no drift surface.
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
settings. That directory is empty today, but Grok can see Claude's config.
## Rule zero, which is the load-bearing part
**Query live inventories, never a written list.** The index points at
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
~30.
## The CLAUDE.md swap
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
before expensive work, post operator-facing links to the Booth board, and the vault is
the credential source of truth. **A rule behind a file read is a rule that stops
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
vault round-trip measured over two minutes — but it is now ALSO vaulted at
`litellm/all-agents-shared-key`, since it had been single-copy.
Commits `53c3e80` `21d24c5`.
@@ -1,3 +0,0 @@
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
@@ -1,91 +0,0 @@
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
Beszel, task-board, vor and the Henge alike.
## How it surfaced, and how it did NOT
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
They had already cleared GPU contention, the ASR itself and output quality with
real controls.
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
whether a mesh path is direct.
## The diagnosis, and two wrong causes published on the way
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
a great candidate for a fixed floor — env vars still set, container gone — but it was
retired from the callback list 2026-06-20 and nothing calls it per-request.
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
both arms share the identical onward path to FV.
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
path AND which destination the packets crossed the network to reach — and the entire
delta was attributed to the variable of interest.
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
nh3-scale$ tailscale ping 100.64.0.3
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
direct connection not established
## Why hole-punching failed
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
only this pair failed.
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
template; writing what a FortiGate normally wants would have failed.
| object | value |
|---|---|
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
## Measured before → after
| path | before | after |
|---|---|---|
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
548 ms, converged with their direct arm.
## Operational notes for next time
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
'show' > file` captures the full non-default config (11,320 lines) without needing a
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
at `fortigate/ana-gw-infra-ops-password`, reached with
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
existing policy or interface over a 327 ms link you cannot recover.
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
cross-site DNS backup because of this fix.
@@ -1,47 +0,0 @@
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
On 2026-09-17, at operator request, SearXNG's search egress was moved from
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
Reverted 2026-09-18.
## Why it was reverted, and why the STATED reason was wrong
It was reverted on a diagnosis that turned out to be false: that the move had cost
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
**byte-identical** result — same three engines down — which falsified it. A live
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
suspension timers too.
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
**The revert still stands on its own merits**: ESH egress bought no measurable
improvement while making all fleet search depend on ESH WAN and mesh availability.
The simpler configuration is the better one. It is simply not the fix for the engines.
## What was built, because it was built correctly
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
source `10.100.50.40` alone — everything else had to supply a password regenerated at
each start and never distributed. Allowed-host egress and denied-host rejection were
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
touched (CT 108 is ESH's whole-site mesh SPOF).
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
WAN address moves and nothing pins it, so the number in the stack README is historical.
## The transferable lesson
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
no longer true of that address. Five samples of the post-change state and zero of the
working state is not a comparison. The rollback WAS the counterfactual, and it
falsified the claim I had already published in a commit message.
This was the first of three wrong causal attributions in a single afternoon. The
common shape: measure after a change, attribute the delta to *my* change, never check
what else moved.
Commits `156e126` (applied), `1a35181` (reverted).
@@ -1,3 +0,0 @@
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
@@ -0,0 +1,3 @@
# `headscale-ddns` exited 1 silently and the alarm carried no cause
`[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
@@ -0,0 +1,3 @@
# A TCP connect as the "is Blender ready" probe
`[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
@@ -0,0 +1,3 @@
# hermes-gateway restarted 0401 for highseat-dev
`[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
@@ -0,0 +1,3 @@
# Relying on the build cache surviving on esh-ml1
`[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
@@ -0,0 +1,3 @@
# The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel
`[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
@@ -0,0 +1,3 @@
# `uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project
`[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
@@ -0,0 +1,3 @@
# Shorter Parakeet slices as Scriberr's memory fix
`[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
@@ -0,0 +1,3 @@
# ⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)
`[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
@@ -0,0 +1,3 @@
# MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)
`[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
@@ -0,0 +1,3 @@
# Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed
`[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
@@ -0,0 +1,3 @@
# Prime: delete the bench leftovers, no upstream for Scriberr, push
`[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
@@ -0,0 +1,3 @@
# Size + mtime scans to pick reclaimable "staging" on a model store
`[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
@@ -0,0 +1,3 @@
# vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`
`[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
@@ -0,0 +1,3 @@
# albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`
`[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
@@ -0,0 +1,3 @@
# Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)
`[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
@@ -0,0 +1,3 @@
# esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`
`[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
@@ -0,0 +1,3 @@
# ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629
`[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
@@ -0,0 +1,3 @@
# nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)
`[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
@@ -0,0 +1,3 @@
# nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03
`[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
@@ -0,0 +1,3 @@
# nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)
`[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
@@ -0,0 +1,3 @@
# Prime: dev-backup gets dailies + weeklies — DONE
`[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
@@ -0,0 +1,3 @@
# Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)
`[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
@@ -0,0 +1,3 @@
# albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1
`[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
@@ -0,0 +1,3 @@
# CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)
`[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
@@ -0,0 +1,3 @@
# nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone
`[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
@@ -0,0 +1,3 @@
# Worldtree memory-gate duty REFRAMED: regression check, not a gate
`[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.