memory: snapshot — 2026-10-03 in-flight (nh3-pve 8→9 under way), two-tier split of 27 entries, 7 archived
This commit is contained in:
@@ -13243,3 +13243,319 @@ ruled. heid also found and killed two live instructions in their own persistent-
|
|||||||
fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment
|
fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment
|
||||||
after a context reset.
|
after a context reset.
|
||||||
_Archived 2026-10-01._
|
_Archived 2026-10-01._
|
||||||
|
|
||||||
|
## Recent decisions (archived 2026-10-03 batch)
|
||||||
|
|
||||||
|
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
|
||||||
|
|
||||||
|
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
|
||||||
|
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
|
||||||
|
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
|
||||||
|
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
|
||||||
|
Beszel, task-board, vor and the Henge alike.
|
||||||
|
|
||||||
|
## How it surfaced, and how it did NOT
|
||||||
|
|
||||||
|
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
|
||||||
|
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
|
||||||
|
They had already cleared GPU contention, the ASR itself and output quality with
|
||||||
|
real controls.
|
||||||
|
|
||||||
|
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
|
||||||
|
whether a mesh path is direct.
|
||||||
|
|
||||||
|
## The diagnosis, and two wrong causes published on the way
|
||||||
|
|
||||||
|
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
|
||||||
|
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
|
||||||
|
a great candidate for a fixed floor — env vars still set, container gone — but it was
|
||||||
|
retired from the callback list 2026-06-20 and nothing calls it per-request.
|
||||||
|
|
||||||
|
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
|
||||||
|
both arms share the identical onward path to FV.
|
||||||
|
|
||||||
|
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
|
||||||
|
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
|
||||||
|
|
||||||
|
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
|
||||||
|
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
|
||||||
|
path AND which destination the packets crossed the network to reach — and the entire
|
||||||
|
delta was attributed to the variable of interest.
|
||||||
|
|
||||||
|
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
|
||||||
|
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
|
||||||
|
|
||||||
|
nh3-scale$ tailscale ping 100.64.0.3
|
||||||
|
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
|
||||||
|
direct connection not established
|
||||||
|
|
||||||
|
## Why hole-punching failed
|
||||||
|
|
||||||
|
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
|
||||||
|
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
|
||||||
|
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
|
||||||
|
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
|
||||||
|
only this pair failed.
|
||||||
|
|
||||||
|
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
|
||||||
|
|
||||||
|
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
|
||||||
|
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
|
||||||
|
template; writing what a FortiGate normally wants would have failed.
|
||||||
|
|
||||||
|
| object | value |
|
||||||
|
|---|---|
|
||||||
|
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
|
||||||
|
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
|
||||||
|
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
|
||||||
|
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
|
||||||
|
|
||||||
|
## Measured before → after
|
||||||
|
|
||||||
|
| path | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
|
||||||
|
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
|
||||||
|
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
|
||||||
|
|
||||||
|
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
|
||||||
|
548 ms, converged with their direct arm.
|
||||||
|
|
||||||
|
## Operational notes for next time
|
||||||
|
|
||||||
|
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
|
||||||
|
'show' > file` captures the full non-default config (11,320 lines) without needing a
|
||||||
|
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
|
||||||
|
at `fortigate/ana-gw-infra-ops-password`, reached with
|
||||||
|
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
|
||||||
|
existing policy or interface over a 327 ms link you cannot recover.
|
||||||
|
|
||||||
|
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
|
||||||
|
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
|
||||||
|
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
|
||||||
|
|
||||||
|
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
|
||||||
|
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
|
||||||
|
cross-site DNS backup because of this fix.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
|
||||||
|
|
||||||
|
Two independent faults, fixed in order. Together they were costing roughly one
|
||||||
|
`.internal` lookup in ten either a hard failure or a five-second stall, on every
|
||||||
|
DHCP client at every site.
|
||||||
|
|
||||||
|
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
|
||||||
|
|
||||||
|
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
|
||||||
|
resolver missed a packet, the resolver waited out its timeout, fell through to
|
||||||
|
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
|
||||||
|
definitive "no such host" rather than a retry. Binary failure: instant, or five
|
||||||
|
seconds then `gaierror`.
|
||||||
|
|
||||||
|
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
|
||||||
|
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
|
||||||
|
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
|
||||||
|
`resolv.conf`.
|
||||||
|
|
||||||
|
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
|
||||||
|
editing it survives exactly until the next lease renewal — tts-dev's original
|
||||||
|
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
|
||||||
|
and then silently reverted. Change it at the source:
|
||||||
|
|
||||||
|
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
|
||||||
|
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
|
||||||
|
`dns-server1` / `dns-server2`.
|
||||||
|
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
|
||||||
|
|
||||||
|
**Operator's ruling was broader than the proposal**: each site's backup resolver
|
||||||
|
should be *another site's* resolver. ESH already worked this way; the other two
|
||||||
|
were set, and the ring closes.
|
||||||
|
|
||||||
|
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
|
||||||
|
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
|
||||||
|
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
|
||||||
|
|
||||||
|
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
|
||||||
|
server (`sfsrv`) and the fortilink switch-management server were deliberately left
|
||||||
|
alone — tenant property.
|
||||||
|
|
||||||
|
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
|
||||||
|
|
||||||
|
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
|
||||||
|
|
||||||
|
ratelimit: 20
|
||||||
|
ratelimit_subnet_len_ipv4: 24
|
||||||
|
ratelimit_whitelist: []
|
||||||
|
|
||||||
|
**20 queries per second shared across an entire /24 — per subnet, not per client.**
|
||||||
|
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
|
||||||
|
and anything over the line is **silently dropped**, costing the client its full 5 s
|
||||||
|
resolver timeout.
|
||||||
|
|
||||||
|
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
|
||||||
|
|
||||||
|
| | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| median | 5044 ms | 72 ms |
|
||||||
|
| timed out (>4 s) | **40 of 60** | **0 of 60** |
|
||||||
|
|
||||||
|
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
|
||||||
|
firewalls; that setting exists to blunt internet-side DNS amplification, which is
|
||||||
|
not this.
|
||||||
|
|
||||||
|
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
|
||||||
|
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
|
||||||
|
live member. That the ring made this rollout safe is the point of having built it
|
||||||
|
first, an hour earlier.
|
||||||
|
|
||||||
|
## End-to-end, the original symptom
|
||||||
|
|
||||||
|
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
|
||||||
|
|
||||||
|
| | original | after ring only | after ratelimit 0 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
|
||||||
|
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
|
||||||
|
| by NAME failed | 2/40 | 0/40 | **0/40** |
|
||||||
|
| by IP p90 (control) | 25 ms | — | 11.7 ms |
|
||||||
|
|
||||||
|
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
|
||||||
|
median**. Resolution has stopped being a cost rather than become a smaller one.
|
||||||
|
|
||||||
|
## What this says about monitoring
|
||||||
|
|
||||||
|
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
|
||||||
|
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
|
||||||
|
of adapting around the problem. Both faults were invisible to Beszel and Uptime
|
||||||
|
Kuma — every host was up, every service answered.
|
||||||
|
|
||||||
|
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
|
||||||
|
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
|
||||||
|
afternoon. The two changes compose.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
|
||||||
|
|
||||||
|
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
|
||||||
|
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
|
||||||
|
Claude sessions load, interleaved with operator preferences other families have no
|
||||||
|
use for.
|
||||||
|
|
||||||
|
## Shape
|
||||||
|
|
||||||
|
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
|
||||||
|
index an agent reads whole; each entry links to a detail file it opens only when it
|
||||||
|
needs that tool. Reading the index costs about a fifth of reading the tree.
|
||||||
|
|
||||||
|
- Canonical: `docs/fleettools/` in this repo (git-tracked)
|
||||||
|
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
|
||||||
|
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
|
||||||
|
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
|
||||||
|
|
||||||
|
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
|
||||||
|
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
|
||||||
|
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
|
||||||
|
|
||||||
|
## Autoload — verified per family, not assumed
|
||||||
|
|
||||||
|
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
|
||||||
|
`"Failed to read global AGENTS.md instructions from"`.
|
||||||
|
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
|
||||||
|
of name** — from its own embedded docs.
|
||||||
|
|
||||||
|
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
|
||||||
|
copied — one file, three families, no drift surface.
|
||||||
|
|
||||||
|
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
|
||||||
|
settings. That directory is empty today, but Grok can see Claude's config.
|
||||||
|
|
||||||
|
## Rule zero, which is the load-bearing part
|
||||||
|
|
||||||
|
**Query live inventories, never a written list.** The index points at
|
||||||
|
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
|
||||||
|
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
|
||||||
|
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
|
||||||
|
~30.
|
||||||
|
|
||||||
|
## The CLAUDE.md swap
|
||||||
|
|
||||||
|
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
|
||||||
|
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
|
||||||
|
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
|
||||||
|
before expensive work, post operator-facing links to the Booth board, and the vault is
|
||||||
|
the credential source of truth. **A rule behind a file read is a rule that stops
|
||||||
|
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
|
||||||
|
|
||||||
|
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
|
||||||
|
vault round-trip measured over two minutes — but it is now ALSO vaulted at
|
||||||
|
`litellm/all-agents-shared-key`, since it had been single-copy.
|
||||||
|
|
||||||
|
Commits `53c3e80` `21d24c5`.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
|
||||||
|
|
||||||
|
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
|
||||||
|
|
||||||
|
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
## Tried and abandoned (archived 2026-10-03 batch)
|
||||||
|
|
||||||
|
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
|
||||||
|
|
||||||
|
On 2026-09-17, at operator request, SearXNG's search egress was moved from
|
||||||
|
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
|
||||||
|
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
|
||||||
|
Reverted 2026-09-18.
|
||||||
|
|
||||||
|
## Why it was reverted, and why the STATED reason was wrong
|
||||||
|
|
||||||
|
It was reverted on a diagnosis that turned out to be false: that the move had cost
|
||||||
|
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
|
||||||
|
**byte-identical** result — same three engines down — which falsified it. A live
|
||||||
|
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
|
||||||
|
suspension timers too.
|
||||||
|
|
||||||
|
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
|
||||||
|
|
||||||
|
**The revert still stands on its own merits**: ESH egress bought no measurable
|
||||||
|
improvement while making all fleet search depend on ESH WAN and mesh availability.
|
||||||
|
The simpler configuration is the better one. It is simply not the fix for the engines.
|
||||||
|
|
||||||
|
## What was built, because it was built correctly
|
||||||
|
|
||||||
|
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
|
||||||
|
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
|
||||||
|
source `10.100.50.40` alone — everything else had to supply a password regenerated at
|
||||||
|
each start and never distributed. Allowed-host egress and denied-host rejection were
|
||||||
|
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
|
||||||
|
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
|
||||||
|
touched (CT 108 is ESH's whole-site mesh SPOF).
|
||||||
|
|
||||||
|
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
|
||||||
|
WAN address moves and nothing pins it, so the number in the stack README is historical.
|
||||||
|
|
||||||
|
## The transferable lesson
|
||||||
|
|
||||||
|
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
|
||||||
|
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
|
||||||
|
no longer true of that address. Five samples of the post-change state and zero of the
|
||||||
|
working state is not a comparison. The rollback WAS the counterfactual, and it
|
||||||
|
falsified the claim I had already published in a commit message.
|
||||||
|
|
||||||
|
This was the first of three wrong causal attributions in a single afternoon. The
|
||||||
|
common shape: measure after a change, attribute the delta to *my* change, never check
|
||||||
|
what else moved.
|
||||||
|
|
||||||
|
Commits `156e126` (applied), `1a35181` (reverted).
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|
||||||
|
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
|
||||||
|
|
||||||
|
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
|
||||||
|
_Archived 2026-10-03._
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# `nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction
|
||||||
|
|
||||||
|
`[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
# `[2026-09-18]` `api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings
|
|
||||||
|
|
||||||
**`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten times with fleet search down. There is no env-var path into that file at all: the loader reads only `SEARXNG_SETTINGS_PATH` and the entrypoint substitutes only `ultrasecretkey`. Literal or nothing.
|
|
||||||
@@ -1,94 +0,0 @@
|
|||||||
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
|
|
||||||
|
|
||||||
Two independent faults, fixed in order. Together they were costing roughly one
|
|
||||||
`.internal` lookup in ten either a hard failure or a five-second stall, on every
|
|
||||||
DHCP client at every site.
|
|
||||||
|
|
||||||
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
|
|
||||||
|
|
||||||
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
|
|
||||||
resolver missed a packet, the resolver waited out its timeout, fell through to
|
|
||||||
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
|
|
||||||
definitive "no such host" rather than a retry. Binary failure: instant, or five
|
|
||||||
seconds then `gaierror`.
|
|
||||||
|
|
||||||
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
|
|
||||||
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
|
|
||||||
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
|
|
||||||
`resolv.conf`.
|
|
||||||
|
|
||||||
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
|
|
||||||
editing it survives exactly until the next lease renewal — tts-dev's original
|
|
||||||
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
|
|
||||||
and then silently reverted. Change it at the source:
|
|
||||||
|
|
||||||
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
|
|
||||||
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
|
|
||||||
`dns-server1` / `dns-server2`.
|
|
||||||
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
|
|
||||||
|
|
||||||
**Operator's ruling was broader than the proposal**: each site's backup resolver
|
|
||||||
should be *another site's* resolver. ESH already worked this way; the other two
|
|
||||||
were set, and the ring closes.
|
|
||||||
|
|
||||||
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
|
|
||||||
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
|
|
||||||
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
|
|
||||||
|
|
||||||
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
|
|
||||||
server (`sfsrv`) and the fortilink switch-management server were deliberately left
|
|
||||||
alone — tenant property.
|
|
||||||
|
|
||||||
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
|
|
||||||
|
|
||||||
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
|
|
||||||
|
|
||||||
ratelimit: 20
|
|
||||||
ratelimit_subnet_len_ipv4: 24
|
|
||||||
ratelimit_whitelist: []
|
|
||||||
|
|
||||||
**20 queries per second shared across an entire /24 — per subnet, not per client.**
|
|
||||||
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
|
|
||||||
and anything over the line is **silently dropped**, costing the client its full 5 s
|
|
||||||
resolver timeout.
|
|
||||||
|
|
||||||
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
|
|
||||||
|
|
||||||
| | before | after |
|
|
||||||
|---|---|---|
|
|
||||||
| median | 5044 ms | 72 ms |
|
|
||||||
| timed out (>4 s) | **40 of 60** | **0 of 60** |
|
|
||||||
|
|
||||||
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
|
|
||||||
firewalls; that setting exists to blunt internet-side DNS amplification, which is
|
|
||||||
not this.
|
|
||||||
|
|
||||||
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
|
|
||||||
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
|
|
||||||
live member. That the ring made this rollout safe is the point of having built it
|
|
||||||
first, an hour earlier.
|
|
||||||
|
|
||||||
## End-to-end, the original symptom
|
|
||||||
|
|
||||||
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
|
|
||||||
|
|
||||||
| | original | after ring only | after ratelimit 0 |
|
|
||||||
|---|---|---|---|
|
|
||||||
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
|
|
||||||
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
|
|
||||||
| by NAME failed | 2/40 | 0/40 | **0/40** |
|
|
||||||
| by IP p90 (control) | 25 ms | — | 11.7 ms |
|
|
||||||
|
|
||||||
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
|
|
||||||
median**. Resolution has stopped being a cost rather than become a smaller one.
|
|
||||||
|
|
||||||
## What this says about monitoring
|
|
||||||
|
|
||||||
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
|
|
||||||
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
|
|
||||||
of adapting around the problem. Both faults were invisible to Beszel and Uptime
|
|
||||||
Kuma — every host was up, every service answered.
|
|
||||||
|
|
||||||
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
|
|
||||||
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
|
|
||||||
afternoon. The two changes compose.
|
|
||||||
@@ -1,57 +0,0 @@
|
|||||||
# [2026-09-18] FleetTools — one capability index, autoloaded by three agent families
|
|
||||||
|
|
||||||
Every agent on nh3-dev needs the same answers — what runs here, how do I call it,
|
|
||||||
what will bite me — but that knowledge lived in `~/.claude/CLAUDE.md`, which only
|
|
||||||
Claude sessions load, interleaved with operator preferences other families have no
|
|
||||||
use for.
|
|
||||||
|
|
||||||
## Shape
|
|
||||||
|
|
||||||
**Two-tier, matching the persistent-memory split.** `FLEETTOOLS.md` is a ~135-line
|
|
||||||
index an agent reads whole; each entry links to a detail file it opens only when it
|
|
||||||
needs that tool. Reading the index costs about a fifth of reading the tree.
|
|
||||||
|
|
||||||
- Canonical: `docs/fleettools/` in this repo (git-tracked)
|
|
||||||
- Discovery: `~/FLEETTOOLS.md` → symlink to the index
|
|
||||||
- **Detail links are ABSOLUTE paths**, not relative markdown — a non-Claude agent
|
|
||||||
`cat`s the path rather than following a link, and may be anywhere on the filesystem.
|
|
||||||
|
|
||||||
Covers: althing, the Booth, `secret`, LiteLLM, inference seats + Asset Engine,
|
|
||||||
speech, Arbo, elway, fleet SSH, observability, graphify, Playwright, the Henge, and
|
|
||||||
a segregated Claude-only file (ratecheck, remote-ssh MCP, task-board, skills).
|
|
||||||
|
|
||||||
## Autoload — verified per family, not assumed
|
|
||||||
|
|
||||||
- **Codex** reads `$CODEX_HOME/AGENTS.md` — the binary carries the string
|
|
||||||
`"Failed to read global AGENTS.md instructions from"`.
|
|
||||||
- **Grok** always scans `$GROK_HOME/rules/` and loads **every `*.md` in it regardless
|
|
||||||
of name** — from its own embedded docs.
|
|
||||||
|
|
||||||
Both locations were empty. `AGENT-BOOTSTRAP.md` is **symlinked** into each rather than
|
|
||||||
copied — one file, three families, no drift surface.
|
|
||||||
|
|
||||||
⚠ Grok also scans `~/.claude/rules/` and recognises `CLAUDE.md` via harness-compat
|
|
||||||
settings. That directory is empty today, but Grok can see Claude's config.
|
|
||||||
|
|
||||||
## Rule zero, which is the load-bearing part
|
|
||||||
|
|
||||||
**Query live inventories, never a written list.** The index points at
|
|
||||||
Homepage `/api/services`, asset-engine `/api/v1/services`, LiteLLM `/v1/models`, and
|
|
||||||
every FastAPI seat's `/openapi.json`. A copied service table would be stale within a
|
|
||||||
month — the LiteLLM roster was already at 40 models where the old CLAUDE.md note said
|
|
||||||
~30.
|
|
||||||
|
|
||||||
## The CLAUDE.md swap
|
|
||||||
|
|
||||||
`~/.claude/CLAUDE.md`'s "Global tools available" section went from **231 lines to 39**
|
|
||||||
(file 1218 → 1026, a 16% cut to what every Claude session loads). Three things stayed
|
|
||||||
inline deliberately because they govern BEHAVIOUR rather than lookup: run `ratecheck`
|
|
||||||
before expensive work, post operator-facing links to the Booth board, and the vault is
|
|
||||||
the credential source of truth. **A rule behind a file read is a rule that stops
|
|
||||||
firing.** Backup at `~/.claude/CLAUDE.md.bak-20260918-074048`.
|
|
||||||
|
|
||||||
The shared all-agents LiteLLM key stays inline there too — every session needs it and a
|
|
||||||
vault round-trip measured over two minutes — but it is now ALSO vaulted at
|
|
||||||
`litellm/all-agents-shared-key`, since it had been single-copy.
|
|
||||||
|
|
||||||
Commits `53c3e80` `21d24c5`.
|
|
||||||
-3
@@ -1,3 +0,0 @@
|
|||||||
# `[2026-09-18]` Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy
|
|
||||||
|
|
||||||
**Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** (operator-approved). Verified empirically — `is_dir()`/`exists()` are stat-based and follow symlinks — and the copy had ALREADY drifted (README, since 09-15). Pre-restart gate is `uv run python scripts/sync_plugin.py --check`; `hermes plugins list` is confirmation-after, never permission-before. Copy kept at `~/backups/svos_miranda-copy-20260918-1444`. ⚠ The install now tracks a working tree — an uncommitted edit in svos is what loads at the next restart.
|
|
||||||
@@ -1,91 +0,0 @@
|
|||||||
# [2026-09-18] NH3↔Anaheim was DERP-relayed, not direct — a UDP 41641 port-forward on ana-gw fixed it
|
|
||||||
|
|
||||||
Every NH3→Anaheim flow had been crossing Tailscale's LA DERP relay rather than a
|
|
||||||
direct path, long enough to have carried **78 GB tx on the NH3 side alone**. DERP
|
|
||||||
is a deliberately throttled fallback — a reachability mechanism, not a data plane —
|
|
||||||
so this imposed both a fixed round-trip penalty and a bandwidth ceiling on LiteLLM,
|
|
||||||
Beszel, task-board, vor and the Henge alike.
|
|
||||||
|
|
||||||
## How it surfaced, and how it did NOT
|
|
||||||
|
|
||||||
tts-dev reported ext-stt transcriptions 4–18× slower through the LiteLLM gateway
|
|
||||||
than straight at the Parakeet host, with a ~1.9 s fixed floor plus a per-byte cost.
|
|
||||||
They had already cleared GPU contention, the ASR itself and output quality with
|
|
||||||
real controls.
|
|
||||||
|
|
||||||
**No monitor caught it.** Beszel had every host green. Nothing on this fleet checks
|
|
||||||
whether a mesh path is direct.
|
|
||||||
|
|
||||||
## The diagnosis, and two wrong causes published on the way
|
|
||||||
|
|
||||||
Their hypothesis was a dead first target in a LiteLLM fallback list. Checked: `ext-stt`
|
|
||||||
has exactly ONE deployment, no fallback, nothing to retry against. Langfuse looked like
|
|
||||||
a great candidate for a fixed floor — env vars still set, container gone — but it was
|
|
||||||
retired from the callback list 2026-06-20 and nothing calls it per-request.
|
|
||||||
|
|
||||||
The decisive arm was one neither of us had run: **measure from ana-docker itself**, so
|
|
||||||
both arms share the identical onward path to FV.
|
|
||||||
|
|
||||||
ana-docker -> fv-ml1 DIRECT 0.289 / 0.218 / 0.273 / 0.220 s
|
|
||||||
ana-docker -> gateway -> fv-ml1 0.463 / 0.211 / 0.222 / 0.222 s
|
|
||||||
|
|
||||||
The gateway adds ~0. tts-dev measured 2618 ms for that clip; from the gateway's own
|
|
||||||
host it is 220 ms. Their two arms differed in TWO things — whether LiteLLM was in the
|
|
||||||
path AND which destination the packets crossed the network to reach — and the entire
|
|
||||||
delta was attributed to the variable of interest.
|
|
||||||
|
|
||||||
nh3-dev -> fv-ml1 direct 0.269 / 0.219 / 0.216 s
|
|
||||||
nh3-dev -> ANA gateway 1.449 / 1.399 / 1.404 s
|
|
||||||
|
|
||||||
nh3-scale$ tailscale ping 100.64.0.3
|
|
||||||
pong via DERP(lax) in 507ms, 518, 521, 522, 373, 423, 464, 493
|
|
||||||
direct connection not established
|
|
||||||
|
|
||||||
## Why hole-punching failed
|
|
||||||
|
|
||||||
ana-scale advertised `38.120.12.42:41641`, but its netcheck mapped to `:60798` — the
|
|
||||||
Anaheim NAT was not preserving the port — and `PortMapping` was empty, so no UPnP or
|
|
||||||
NAT-PMP was establishing one. `MappingVariesByDestIP: false`, so not symmetric NAT;
|
|
||||||
simply no reachable inbound endpoint. ESH↔Anaheim and NH3↔ESH were already direct;
|
|
||||||
only this pair failed.
|
|
||||||
|
|
||||||
## The fix — four ADDITIVE objects on ana-gw (10.250.0.1)
|
|
||||||
|
|
||||||
⚠ **This box runs `central-nat enable`, so a policy's `dstaddr` is the REAL internal
|
|
||||||
address and not the VIP.** The existing `wg-to-ana-wg` VIP+policy pair is the house
|
|
||||||
template; writing what a FortiGate normally wants would have failed.
|
|
||||||
|
|
||||||
| object | value |
|
|
||||||
|---|---|
|
|
||||||
| `firewall address` | `ana-scale-ip` → 10.250.50.45/32, iface `servers` |
|
|
||||||
| `firewall service custom` | `Tailscale-41641` → `udp-portrange 41641` |
|
|
||||||
| `firewall vip` | `tailscale-to-ana-scale` → 38.120.12.42:41641 udp → 10.250.50.45:41641, extintf wan1 |
|
|
||||||
| `firewall policy` id 75 | wan1→servers, all→`ana-scale-ip`, `Tailscale-41641`, accept |
|
|
||||||
|
|
||||||
## Measured before → after
|
|
||||||
|
|
||||||
| path | before | after |
|
|
||||||
|---|---|---|
|
|
||||||
| `tailscale ping` nh3-scale→ana-scale | 373–522 ms via DERP(lax) | **6 ms direct** |
|
|
||||||
| STT via the ANA gateway (96 kB clip) | 1.399–1.449 s | **0.237–0.270 s** |
|
|
||||||
| Beszel HTTP nh3-dev→ana-docker | 0.94–1.29 s | **0.014–0.016 s** |
|
|
||||||
|
|
||||||
tts-dev confirmed independently from talk's own vantage: 2618 → 249 ms and 4119 →
|
|
||||||
548 ms, converged with their direct arm.
|
|
||||||
|
|
||||||
## Operational notes for next time
|
|
||||||
|
|
||||||
⚠ **This edge has NO out-of-band access.** Back up first — `ssh infra-ops@10.250.0.1
|
|
||||||
'show' > file` captures the full non-default config (11,320 lines) without needing a
|
|
||||||
tftp server; kept at `~/backups/ana-gw-config-BEFORE-<stamp>.txt`. Credentials vaulted
|
|
||||||
at `fortigate/ana-gw-infra-ops-password`, reached with
|
|
||||||
`sshpass -e ssh -o PubkeyAuthentication=no`. **Additive objects only** — never edit an
|
|
||||||
existing policy or interface over a 327 ms link you cannot recover.
|
|
||||||
|
|
||||||
⚠ **irv-ml1 (100.64.0.6) remains `relay "lax"`.** Same class, different site, not fixed
|
|
||||||
by this. Measured 24-25 ms on HTTP from nh3-dev, so it is not costing what Anaheim was.
|
|
||||||
I warned tts-dev their TTS stack was paying for it; they measured and I was wrong.
|
|
||||||
|
|
||||||
Documented in `docs/pfi/headscale-mesh-plan.md`; commit `5a9fad8`. See
|
|
||||||
[[2026-09-18-fleet-dns-ring-and-ratelimit]] — the ANA AdGuard only became a viable
|
|
||||||
cross-site DNS backup because of this fix.
|
|
||||||
@@ -1,47 +0,0 @@
|
|||||||
# [2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
|
|
||||||
|
|
||||||
On 2026-09-17, at operator request, SearXNG's search egress was moved from
|
|
||||||
nh3-docker's direct residential path to `socks5h://10.0.50.65:1080` on **esh-scale**
|
|
||||||
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
|
|
||||||
Reverted 2026-09-18.
|
|
||||||
|
|
||||||
## Why it was reverted, and why the STATED reason was wrong
|
|
||||||
|
|
||||||
It was reverted on a diagnosis that turned out to be false: that the move had cost
|
|
||||||
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
|
|
||||||
**byte-identical** result — same three engines down — which falsified it. A live
|
|
||||||
`!ddg` bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
|
|
||||||
suspension timers too.
|
|
||||||
|
|
||||||
The real cause was elsewhere (see [[2026-09-18-searxng-one-engine-to-seven]]).
|
|
||||||
|
|
||||||
**The revert still stands on its own merits**: ESH egress bought no measurable
|
|
||||||
improvement while making all fleet search depend on ESH WAN and mesh availability.
|
|
||||||
The simpler configuration is the better one. It is simply not the fix for the engines.
|
|
||||||
|
|
||||||
## What was built, because it was built correctly
|
|
||||||
|
|
||||||
The proxy side was not at fault and is worth keeping as a pattern. `microsocks` ran
|
|
||||||
as `nobody` under `searxng-egress.service`, bound `10.0.50.65:1080` only, and allowed
|
|
||||||
source `10.100.50.40` alone — everything else had to supply a password regenerated at
|
|
||||||
each start and never distributed. Allowed-host egress and denied-host rejection were
|
|
||||||
both tested. The unit is kept at `configs/esh-scale/searxng-egress.service`; the
|
|
||||||
service is **stopped and disabled** on esh-scale, and `tailscaled` there was not
|
|
||||||
touched (CT 108 is ESH's whole-site mesh SPOF).
|
|
||||||
|
|
||||||
Measured egress was `128.177.138.182` in use and `154.50.58.126` at cutover — ESH's
|
|
||||||
WAN address moves and nothing pins it, so the number in the stack README is historical.
|
|
||||||
|
|
||||||
## The transferable lesson
|
|
||||||
|
|
||||||
⚠ **I asserted causation from correlation with no baseline.** The only evidence that
|
|
||||||
residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was
|
|
||||||
no longer true of that address. Five samples of the post-change state and zero of the
|
|
||||||
working state is not a comparison. The rollback WAS the counterfactual, and it
|
|
||||||
falsified the claim I had already published in a commit message.
|
|
||||||
|
|
||||||
This was the first of three wrong causal attributions in a single afternoon. The
|
|
||||||
common shape: measure after a change, attribute the delta to *my* change, never check
|
|
||||||
what else moved.
|
|
||||||
|
|
||||||
Commits `156e126` (applied), `1a35181` (reverted).
|
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
# `[2026-09-18]` Worldtree's `env.sh` secrets are vaulted
|
|
||||||
|
|
||||||
**Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal, z-ai, zai). ⚠ `Z_AI_API_KEY` and `ZAI_API_KEY` are DIFFERENT keys despite the near-identical names (fp `26bde3fa` vs `a4884152`). `ANTHROPIC_API_KEY` was empty and skipped.
|
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# `headscale-ddns` exited 1 silently and the alarm carried no cause
|
||||||
|
|
||||||
|
`[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# A TCP connect as the "is Blender ready" probe
|
||||||
|
|
||||||
|
`[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# hermes-gateway restarted 0401 for highseat-dev
|
||||||
|
|
||||||
|
`[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Relying on the build cache surviving on esh-ml1
|
||||||
|
|
||||||
|
`[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel
|
||||||
|
|
||||||
|
`[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# `uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project
|
||||||
|
|
||||||
|
`[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Shorter Parakeet slices as Scriberr's memory fix
|
||||||
|
|
||||||
|
`[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# ⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)
|
||||||
|
|
||||||
|
`[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)
|
||||||
|
|
||||||
|
`[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed
|
||||||
|
|
||||||
|
`[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# Prime: delete the bench leftovers, no upstream for Scriberr, push
|
||||||
|
|
||||||
|
`[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Size + mtime scans to pick reclaimable "staging" on a model store
|
||||||
|
|
||||||
|
`[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`
|
||||||
|
|
||||||
|
`[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`
|
||||||
|
|
||||||
|
`[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)
|
||||||
|
|
||||||
|
`[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`
|
||||||
|
|
||||||
|
`[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629
|
||||||
|
|
||||||
|
`[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)
|
||||||
|
|
||||||
|
`[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03
|
||||||
|
|
||||||
|
`[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)
|
||||||
|
|
||||||
|
`[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Prime: dev-backup gets dailies + weeklies — DONE
|
||||||
|
|
||||||
|
`[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)
|
||||||
|
|
||||||
|
`[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1
|
||||||
|
|
||||||
|
`[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)
|
||||||
|
|
||||||
|
`[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone
|
||||||
|
|
||||||
|
`[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
# Worldtree memory-gate duty REFRAMED: regression check, not a gate
|
||||||
|
|
||||||
|
`[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
|
||||||
+68
-38
@@ -1,6 +1,6 @@
|
|||||||
# Persistent memory — eshpfi-management
|
# Persistent memory — eshpfi-management
|
||||||
|
|
||||||
_Last updated: 2026-10-01 ~0446 PT (gen-small OOM fixed by parakeet-nemo 0.1.1, with a hard cap and cache return; Prime: leave the gen-small KV and graphs-off as is; arbo webhook fixed (Gitea allow-list); leftovers deleted, repos pushed. Prior: Parakeet seat → unified-en LIVE, U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
|
_Last updated: 2026-10-03 ~1335 PT (nh3-pve PVE 8→9 upgrade by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_
|
||||||
|
|
||||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||||
@@ -115,7 +115,43 @@ no longer deployed sidecars here. See Recent decisions.)
|
|||||||
|
|
||||||
## Current state / in-flight
|
## Current state / in-flight
|
||||||
|
|
||||||
_As of 2026-10-01 ~0446 PT._
|
_As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections carry their own dates._
|
||||||
|
|
||||||
|
### 2026-10-03: live now (as of ~1335 PT)
|
||||||
|
|
||||||
|
- **nh3-pve PVE 8.4.1 → 9 upgrade: Prime is running it at the console now.** Steps given to him:
|
||||||
|
- latest 8.4, then `pve8to9 --full`;
|
||||||
|
- `apt remove systemd-boot` (safe: it boots GRUB via proxmox-boot-tool);
|
||||||
|
- `zfs snapshot -r rpool/ROOT@pre-pve9`;
|
||||||
|
- bookworm → trixie in `/etc/apt/sources.list`, then `apt dist-upgrade`;
|
||||||
|
- **before the reboot, `dkms status` must show nvidia 580.178.04 built for the NEW kernel** (nh3-ml1's GPU);
|
||||||
|
- reboot.
|
||||||
|
NIC names are already pinned (`/etc/systemd/network/10-pin-*.link`). **nh3-dev (VM 102, this session's host) goes down with it.**
|
||||||
|
- **Rollback for nh3-pve:** the ZFS snapshot, from a rescue boot. Never run `zpool upgrade rpool`.
|
||||||
|
- **POST-CHECK owed once nh3-pve and nh3-dev are back:**
|
||||||
|
- every guest is running: VMs 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3 (104 nh3-laser and 108 opnsense-lab were already stopped); CTs 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale, 109 nh3-ml1;
|
||||||
|
- NH3 DNS (AdGuard on nh3-docker 10.100.50.40);
|
||||||
|
- the althing post office (:8390) and albok-service (:8392);
|
||||||
|
- the mesh (headscale, nh3-scale);
|
||||||
|
- Miranda's channel (svos :8770 + hermes-gateway on nh3-dev);
|
||||||
|
- nh3-ml1's GPU (`nvidia-smi` in CT 109);
|
||||||
|
- `nh3-pve-amt` connected in MeshCentral;
|
||||||
|
- Beszel;
|
||||||
|
- `pveversion` shows 9.x.
|
||||||
|
- **nh3-pve-2** (MS-03, NH3) is live at 10.100.250.62 on nh3-mgmt:
|
||||||
|
- X710 `nic3` on USW port 25;
|
||||||
|
- AMT `nic0` with `auto` + `/etc/sysctl.d/90-amt-port.conf`;
|
||||||
|
- vmstore 913 GiB;
|
||||||
|
- subscription popup patched.
|
||||||
|
Open and NOT approved yet (Prime's call): disable the ceph enterprise apt source, apt upgrade, and one reboot to prove `.62` and nic0 survive. → `servers/nh3-pve-2/README.md`
|
||||||
|
- **esh-pve-2** (MS-03, ESH) is unplugged by Prime on purpose. vmstore + AMT phone-home are done. The permanent address is deferred by Prime ("address can wait"; `servers/esh-pve-2/README.md`). Run `playbooks/pve-nag-patch.yaml` on it when it is back.
|
||||||
|
- **cira-tunnel-watchdog** on pfi-tacticalrmm has been live since 1232 and drops :4433 tunnels silent >180 s. Watch `journalctl -t cira-tunnel-watchdog`.
|
||||||
|
- **Done today, nothing open:**
|
||||||
|
- tank verified healthy 0629, disk kept, Miranda told;
|
||||||
|
- albok-service 0.1.2 + the Nemi hourly timer;
|
||||||
|
- the post-deletion Worldtree gate (infra-hermes runs it as a regression check);
|
||||||
|
- the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve.
|
||||||
|
- **ESH PVE 8→9 plan:** written, NOT executed (needs Prime's green light). `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. The tank blocker is cleared.
|
||||||
|
|
||||||
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
|
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
|
||||||
|
|
||||||
@@ -287,25 +323,26 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
|
|
||||||
## Recent decisions
|
## Recent decisions
|
||||||
|
|
||||||
- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
|
- `[2026-10-03]` **PVE subscription popup patched out on nh3-pve, nh3-pve-2, pfi-pve, esh-pve (Prime: "every host")**: an anchored one-line edit plus an apt hook. esh-nas-pve was already de-nagged; esh-pve-2, sfsrv-ana and the PBS hosts are not done. → `services/pve-nag-patch/README.md`
|
||||||
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
|
- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone** → `persistent-memory.d/2026-10-03-nh3-pve-2-finished-prime-on-site-host-10-100-250.md`
|
||||||
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
|
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** → `persistent-memory.d/2026-10-03-cira-tunnel-watchdog-live-on-pfi-tacticalrmm-prime-go-1232.md`
|
||||||
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
|
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** → `persistent-memory.d/2026-10-03-worldtree-memory-gate-duty-reframed-regression-check-not-a-gate.md`
|
||||||
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
|
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1** → `persistent-memory.d/2026-10-03-albok-service-0-1-2-live-on-nh3-docker-albok-dev-ask.md`
|
||||||
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
|
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`** → `persistent-memory.d/2026-10-02-esh-pve-2-ms-03-at-esh-future-esh-dev-host-onboarded-infra-ops.md`
|
||||||
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
|
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684)** → `persistent-memory.d/2026-10-02-demo-outage-13-min-rolled-back-worldtree-dev-urgent-6684.md`
|
||||||
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
|
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`)** → `persistent-memory.d/2026-10-02-nh3-pve-s-amt-is-in-tacticalrmm-s-meshcentral-group.md`
|
||||||
- `[2026-10-02]` ✅ **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. **Prime (via Miranda) 1925: verify, then clear + full scrub.** Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. `zpool clear` 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. **UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107.** History explains it: after the Aug 20 replace, the resilver ended `errors=6676436`, someone ran only an error-scrub (`zpool scrub -e`, 4 s) and then `zpool clear` 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse `6.19M` as an integer); it still fires on completion. **0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m.** **0118: SECOND full scrub started WITHOUT `zpool clear`**, so any new error shows as count > 6494530 (`zpool status -p`); watcher tested with +1/null/finished controls; ETA ~0700. **RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. `zpool clear` 0629 → `pool 'tank' is healthy`.** Reported to Miranda once. **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
|
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE** → `persistent-memory.d/2026-10-02-prime-dev-backup-gets-dailies-weeklies-done.md`
|
||||||
- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
|
- `[2026-10-02]` **ESH `tank` (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629** → `persistent-memory.d/2026-10-02-esh-tank-esh-nas-pve-175-t-raidz2-2-was-degraded.md`
|
||||||
- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
|
- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** → `persistent-memory.d/2026-10-02-albok-service-0-1-0-live-on-nh3-docker-albok-dev-ask.md`
|
||||||
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
|
- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03** → `persistent-memory.d/2026-10-02-nh3-pve-2-status-at-end-of-day-amt-done-proxmox.md`
|
||||||
- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
|
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910)** → `persistent-memory.d/2026-10-02-prime-ms-03s-get-proxmox-dummy-plugs-in-hand-one.md`
|
||||||
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
|
- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot)** → `persistent-memory.d/2026-10-02-nh3-dev-root-grown-250-378-gb-prime-resized-scsi0.md`
|
||||||
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
|
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev)** → `persistent-memory.d/2026-10-01-mia-make-it-animatable-v2-auto-rigger-live-on-fv-ml1-gpu-3.md`
|
||||||
|
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed** → `persistent-memory.d/2026-10-01-prime-clean-up-the-1700-backups-dev-backup-retention-fixed.md`
|
||||||
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
|
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
|
||||||
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
|
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10)** → `persistent-memory.d/2026-10-01-gitea-s-webhook-allowed-host-list-external-10-0.md`
|
||||||
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
|
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`** → `persistent-memory.d/2026-10-01-vh-arbo-push-webhook-repointed-from-the-retired-wg0.md`
|
||||||
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
|
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push** → `persistent-memory.d/2026-10-01-prime-delete-the-bench-leftovers-no-upstream-for-scriberr.md`
|
||||||
- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
|
- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
|
||||||
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
|
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
|
||||||
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
|
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
|
||||||
@@ -316,7 +353,7 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
|
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
|
||||||
- `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3.
|
- `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3.
|
||||||
- `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions
|
- `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions
|
||||||
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
|
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** → `persistent-memory.d/2026-09-27-hermes-gateway-restarted-0401-for-highseat-dev.md`
|
||||||
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
|
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
|
||||||
- `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md`
|
- `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md`
|
||||||
- `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md`
|
- `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md`
|
||||||
@@ -370,7 +407,7 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
- `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md`
|
- `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md`
|
||||||
- `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md`
|
- `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md`
|
||||||
- `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md`
|
- `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md`
|
||||||
- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`.
|
- `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** → `persistent-memory.d/2026-09-22-headscale-ddns-exited-1-silently-and-the-alarm-carried-no.md`
|
||||||
|
|
||||||
- `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md`
|
- `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md`
|
||||||
- `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md`
|
- `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md`
|
||||||
@@ -418,13 +455,8 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
|
|
||||||
- `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md`
|
- `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md`
|
||||||
|
|
||||||
- `[2026-09-18]` ⭐⭐⭐ **NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure.** Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; `tailscale ping` 373–522 ms → **6 ms direct**, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs `central-nat`, so a policy `dstaddr` is the REAL internal address, not the VIP. No OOB access — back up with `show` to a local file and make additive changes ONLY. irv-ml1 still relayed. → `persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md`
|
|
||||||
- `[2026-09-18]` ⭐⭐⭐ **`.internal` DNS was failing ~10% of lookups fleet-wide, from two independent causes.** A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped `ratelimit: 20` shared across an entire **/24**, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ `resolv.conf` is DHCP-managed — change it at the UDM/FortiGate, not the file. → `persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md`
|
|
||||||
- `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md`
|
- `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md`
|
||||||
- `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md`
|
- `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md`
|
||||||
- `[2026-09-18]` ⭐ **FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.** `docs/fleettools/` + `~/FLEETTOOLS.md`; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → `persistent-memory.d/2026-09-18-fleettools-agent-index.md`
|
|
||||||
- `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md`
|
|
||||||
- `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md`
|
|
||||||
|
|
||||||
- `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md`
|
- `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md`
|
||||||
⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).
|
⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).
|
||||||
@@ -445,26 +477,26 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
|
|
||||||
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md`
|
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md`
|
||||||
|
|
||||||
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
|
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction** → `persistent-memory.d/2026-08-24-nconnect-8-on-mnt-smithy-approved-but-deferred-at.md`
|
||||||
|
|
||||||
_167 older entries archived to archival-memory.md._
|
_172 older entries archived to archival-memory.md._
|
||||||
|
|
||||||
## Tried and abandoned
|
## Tried and abandoned
|
||||||
|
|
||||||
- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it.
|
- `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** → `persistent-memory.d/2026-10-01-size-mtime-scans-to-pick-reclaimable-staging-on-a.md`
|
||||||
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
|
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix** → `persistent-memory.d/2026-09-30-shorter-parakeet-slices-as-scriberr-s-memory-fix.md`
|
||||||
- `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.
|
- `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.
|
||||||
- `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.
|
- `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.
|
||||||
- `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.
|
- `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.
|
||||||
- `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`.
|
- `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`.
|
||||||
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
|
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel** → `persistent-memory.d/2026-09-27-the-blender-mcp-server-on-nh3-dev-reaching-the-add-on.md`
|
||||||
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
|
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe** → `persistent-memory.d/2026-09-27-a-tcp-connect-as-the-is-blender-ready-probe.md`
|
||||||
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
|
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
|
||||||
- `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.
|
- `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.
|
||||||
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
|
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** → `persistent-memory.d/2026-09-27-uv-export-requirements-txt-uv-pip-install-for-a.md`
|
||||||
- `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`.
|
- `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`.
|
||||||
- `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`.
|
- `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`.
|
||||||
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
|
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** → `persistent-memory.d/2026-09-27-relying-on-the-build-cache-surviving-on-esh-ml1.md`
|
||||||
- `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md`
|
- `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md`
|
||||||
- `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`.
|
- `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`.
|
||||||
- `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`.
|
- `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`.
|
||||||
@@ -475,8 +507,6 @@ _167 older entries archived to archival-memory.md._
|
|||||||
|
|
||||||
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md`
|
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md`
|
||||||
|
|
||||||
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`
|
|
||||||
|
|
||||||
- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md`
|
|
||||||
|
|
||||||
_118 older entries archived to archival-memory.md._
|
_120 older entries archived to archival-memory.md._
|
||||||
|
|||||||
Reference in New Issue
Block a user