memory: snapshot — 2026-10-03 in-flight (nh3-pve 8→9 under way), two-tier split of 27 entries, 7 archived
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
# `nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction
|
||||
|
||||
`[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
|
||||
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
|
||||
|
||||
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
|
||||
@@ -1,94 +0,0 @@
|
||||
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
|
||||
|
||||
Two independent faults, fixed in order. Together they were costing roughly one
|
||||
`.internal` lookup in ten either a hard failure or a five-second stall, on every
|
||||
DHCP client at every site.
|
||||
|
||||
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
|
||||
|
||||
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
|
||||
resolver missed a packet, the resolver waited out its timeout, fell through to
|
||||
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
|
||||
definitive "no such host" rather than a retry. Binary failure: instant, or five
|
||||
seconds then `gaierror`.
|
||||
|
||||
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
|
||||
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
|
||||
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
|
||||
`resolv.conf`.
|
||||
|
||||
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
|
||||
editing it survives exactly until the next lease renewal — tts-dev's original
|
||||
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
|
||||
and then silently reverted. Change it at the source:
|
||||
|
||||
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
|
||||
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
|
||||
`dns-server1` / `dns-server2`.
|
||||
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
|
||||
|
||||
**Operator's ruling was broader than the proposal**: each site's backup resolver
|
||||
should be *another site's* resolver. ESH already worked this way; the other two
|
||||
were set, and the ring closes.
|
||||
|
||||
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
|
||||
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
|
||||
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
|
||||
|
||||
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
|
||||
server (`sfsrv`) and the fortilink switch-management server were deliberately left
|
||||
alone — tenant property.
|
||||
|
||||
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
|
||||
|
||||
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
|
||||
|
||||
ratelimit: 20
|
||||
ratelimit_subnet_len_ipv4: 24
|
||||
ratelimit_whitelist: []
|
||||
|
||||
**20 queries per second shared across an entire /24 — per subnet, not per client.**
|
||||
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
|
||||
and anything over the line is **silently dropped**, costing the client its full 5 s
|
||||
resolver timeout.
|
||||
|
||||
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| median | 5044 ms | 72 ms |
|
||||
| timed out (>4 s) | **40 of 60** | **0 of 60** |
|
||||
|
||||
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
|
||||
firewalls; that setting exists to blunt internet-side DNS amplification, which is
|
||||
not this.
|
||||
|
||||
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
|
||||
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
|
||||
live member. That the ring made this rollout safe is the point of having built it
|
||||
first, an hour earlier.
|
||||
|
||||
## End-to-end, the original symptom
|
||||
|
||||
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
|
||||
|
||||
| | original | after ring only | after ratelimit 0 |
|
||||
|---|---|---|---|
|
||||
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
|
||||
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
|
||||
| by NAME failed | 2/40 | 0/40 | **0/40** |
|
||||
| by IP p90 (control) | 25 ms | — | 11.7 ms |
|
||||
|
||||
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
|
||||
median**. Resolution has stopped being a cost rather than become a smaller one.
|
||||
|
||||
## What this says about monitoring
|
||||
|
||||
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
|
||||
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
|
||||
of adapting around the problem. Both faults were invisible to Beszel and Uptime
|
||||
Kuma — every host was up, every service answered.
|
||||
|
||||
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
|
||||
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
|
||||
afternoon. The two changes compose.
|
||||
@@ -1,57 +0,0 @@
|
||||
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
|
||||
|
||||
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
|
||||
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
|
||||
Claude sessions load, interleaved with operator preferences other families have no
|
||||
use for.
|
||||
|
||||
## Shape
|
||||
|
||||
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
|
||||
index an agent reads whole; each entry links to a detail file it opens only when it
|
||||
needs that tool. Reading the index costs about a fifth of reading the tree.
|
||||
|
||||
- Canonical: `docs/fleettools/` in this repo (git-tracked)
|
||||
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
|
||||
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
|
||||
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
|
||||
|
||||
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
|
||||
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
|
||||
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
|
||||
|
||||
## Autoload — verified per family, not assumed
|
||||
|
||||
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
|
||||
`"Failed to read global AGENTS.md instructions from"`.
|
||||
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
|
||||
of name** — from its own embedded docs.
|
||||
|
||||
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
|
||||
copied — one file, three families, no drift surface.
|
||||
|
||||
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
|
||||
settings. That directory is empty today, but Grok can see Claude's config.
|
||||
|
||||
## Rule zero, which is the load-bearing part
|
||||
|
||||
**Query live inventories, never a written list.** The index points at
|
||||
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
|
||||
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
|
||||
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
|
||||
~30.
|
||||
|
||||
## The CLAUDE.md swap
|
||||
|
||||
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
|
||||
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
|
||||
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
|
||||
before expensive work, post operator-facing links to the Booth board, and the vault is
|
||||
the credential source of truth. **A rule behind a file read is a rule that stops
|
||||
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
|
||||
|
||||
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
|
||||
vault round-trip measured over two minutes — but it is now ALSO vaulted at
|
||||
`litellm/all-agents-shared-key`, since it had been single-copy.
|
||||
|
||||
Commits `53c3e80` `21d24c5`.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
|
||||
|
||||
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
|
||||
@@ -1,91 +0,0 @@
|
||||
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
|
||||
|
||||
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
|
||||
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
|
||||
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
|
||||
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
|
||||
Beszel, task-board, vor and the Henge alike.
|
||||
|
||||
## How it surfaced, and how it did NOT
|
||||
|
||||
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
|
||||
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
|
||||
They had already cleared GPU contention, the ASR itself and output quality with
|
||||
real controls.
|
||||
|
||||
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
|
||||
whether a mesh path is direct.
|
||||
|
||||
## The diagnosis, and two wrong causes published on the way
|
||||
|
||||
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
|
||||
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
|
||||
a great candidate for a fixed floor — env vars still set, container gone — but it was
|
||||
retired from the callback list 2026-06-20 and nothing calls it per-request.
|
||||
|
||||
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
|
||||
both arms share the identical onward path to FV.
|
||||
|
||||
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
|
||||
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
|
||||
|
||||
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
|
||||
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
|
||||
path AND which destination the packets crossed the network to reach — and the entire
|
||||
delta was attributed to the variable of interest.
|
||||
|
||||
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
|
||||
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
|
||||
|
||||
nh3-scale$ tailscale ping 100.64.0.3
|
||||
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
|
||||
direct connection not established
|
||||
|
||||
## Why hole-punching failed
|
||||
|
||||
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
|
||||
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
|
||||
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
|
||||
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
|
||||
only this pair failed.
|
||||
|
||||
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
|
||||
|
||||
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
|
||||
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
|
||||
template; writing what a FortiGate normally wants would have failed.
|
||||
|
||||
| object | value |
|
||||
|---|---|
|
||||
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
|
||||
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
|
||||
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
|
||||
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
|
||||
|
||||
## Measured before → after
|
||||
|
||||
| path | before | after |
|
||||
|---|---|---|
|
||||
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
|
||||
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
|
||||
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
|
||||
|
||||
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
|
||||
548 ms, converged with their direct arm.
|
||||
|
||||
## Operational notes for next time
|
||||
|
||||
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
|
||||
'show' > file` captures the full non-default config (11,320 lines) without needing a
|
||||
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
|
||||
at `fortigate/ana-gw-infra-ops-password`, reached with
|
||||
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
|
||||
existing policy or interface over a 327 ms link you cannot recover.
|
||||
|
||||
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
|
||||
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
|
||||
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
|
||||
|
||||
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
|
||||
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
|
||||
cross-site DNS backup because of this fix.
|
||||
@@ -1,47 +0,0 @@
|
||||
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
|
||||
|
||||
On 2026-09-17, at operator request, SearXNG's search egress was moved from
|
||||
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
|
||||
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
|
||||
Reverted 2026-09-18.
|
||||
|
||||
## Why it was reverted, and why the STATED reason was wrong
|
||||
|
||||
It was reverted on a diagnosis that turned out to be false: that the move had cost
|
||||
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
|
||||
**byte-identical** result — same three engines down — which falsified it. A live
|
||||
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
|
||||
suspension timers too.
|
||||
|
||||
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
|
||||
|
||||
**The revert still stands on its own merits**: ESH egress bought no measurable
|
||||
improvement while making all fleet search depend on ESH WAN and mesh availability.
|
||||
The simpler configuration is the better one. It is simply not the fix for the engines.
|
||||
|
||||
## What was built, because it was built correctly
|
||||
|
||||
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
|
||||
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
|
||||
source `10.100.50.40` alone — everything else had to supply a password regenerated at
|
||||
each start and never distributed. Allowed-host egress and denied-host rejection were
|
||||
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
|
||||
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
|
||||
touched (CT 108 is ESH's whole-site mesh SPOF).
|
||||
|
||||
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
|
||||
WAN address moves and nothing pins it, so the number in the stack README is historical.
|
||||
|
||||
## The transferable lesson
|
||||
|
||||
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
|
||||
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
|
||||
no longer true of that address. Five samples of the post-change state and zero of the
|
||||
working state is not a comparison. The rollback WAS the counterfactual, and it
|
||||
falsified the claim I had already published in a commit message.
|
||||
|
||||
This was the first of three wrong causal attributions in a single afternoon. The
|
||||
common shape: measure after a change, attribute the delta to *my* change, never check
|
||||
what else moved.
|
||||
|
||||
Commits `156e126` (applied), `1a35181` (reverted).
|
||||
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
|
||||
|
||||
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# `headscale-ddns` exited 1 silently and the alarm carried no cause
|
||||
|
||||
`[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
|
||||
@@ -0,0 +1,3 @@
|
||||
# A TCP connect as the "is Blender ready" probe
|
||||
|
||||
`[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
|
||||
@@ -0,0 +1,3 @@
|
||||
# hermes-gateway restarted 0401 for highseat-dev
|
||||
|
||||
`[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
|
||||
@@ -0,0 +1,3 @@
|
||||
# Relying on the build cache surviving on esh-ml1
|
||||
|
||||
`[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel
|
||||
|
||||
`[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
|
||||
@@ -0,0 +1,3 @@
|
||||
# `uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project
|
||||
|
||||
`[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
|
||||
@@ -0,0 +1,3 @@
|
||||
# Shorter Parakeet slices as Scriberr's memory fix
|
||||
|
||||
`[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
|
||||
@@ -0,0 +1,3 @@
|
||||
# ⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)
|
||||
|
||||
`[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)
|
||||
|
||||
`[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed
|
||||
|
||||
`[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# Prime: delete the bench leftovers, no upstream for Scriberr, push
|
||||
|
||||
`[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
|
||||
@@ -0,0 +1,3 @@
|
||||
# Size + mtime scans to pick reclaimable "staging" on a model store
|
||||
|
||||
`[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
|
||||
@@ -0,0 +1,3 @@
|
||||
# vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`
|
||||
|
||||
`[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
|
||||
@@ -0,0 +1,3 @@
|
||||
# albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`
|
||||
|
||||
`[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)
|
||||
|
||||
`[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`
|
||||
|
||||
`[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
|
||||
@@ -0,0 +1,3 @@
|
||||
# ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629
|
||||
|
||||
`[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
|
||||
@@ -0,0 +1,3 @@
|
||||
# nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)
|
||||
|
||||
`[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
|
||||
@@ -0,0 +1,3 @@
|
||||
# nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03
|
||||
|
||||
`[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
|
||||
@@ -0,0 +1,3 @@
|
||||
# nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)
|
||||
|
||||
`[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
|
||||
@@ -0,0 +1,3 @@
|
||||
# Prime: dev-backup gets dailies + weeklies — DONE
|
||||
|
||||
`[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
|
||||
@@ -0,0 +1,3 @@
|
||||
# Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)
|
||||
|
||||
`[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
|
||||
@@ -0,0 +1,3 @@
|
||||
# albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1
|
||||
|
||||
`[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)
|
||||
|
||||
`[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
|
||||
@@ -0,0 +1,3 @@
|
||||
# nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone
|
||||
|
||||
`[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# Worldtree memory-gate duty REFRAMED: regression check, not a gate
|
||||
|
||||
`[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
|
||||
Reference in New Issue
Block a user