Session captured: the FV outbound-NAT root cause and its diagnostic signature, the fv-ml1 dead man's switch, fleet identity/group/path conventions and the root:docker normalization, nh3-dev's ts-input reachability fix, ESPHome modernisation and the kb KB-search tool, and the Hermes bearer rotation release. Six new detail files. Tried-and-abandoned gains three: probing OPNsense endpoints by POSTing at them (which rebooted the FV firewall), advertising a /32 from nh3-dev, and the nh3-scale remote-site masquerade rules that fired but were not the fix. Housekeeping: 8 Recent-decisions entries archived to archival-memory.md, and 21 oversized inline entries split into detail files per the two-tier rule -- they had been sitting fully inline in the index, which is what the split exists to prevent. Two pointers to a detail file archived this run were repointed at archival-memory.md. The index is 389 lines, still over the ~300 soft cap. The archival guards stop it there: only 4 further entries are old enough to move and every one carries an open deferred-work pointer. An over-cap file that keeps live decisions beats a scannable one that lost a deferred call.
1.8 KiB
[2026-09-13] FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load
⚠⚠⚠ FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load test; all other sites healthy. Operator's leading hypothesis: the 1500 VA Eaton UPS overloaded and DIED. It fits better than a breaker trip because a UPS's output rating sits far below the circuit's, making it the first protective device to give — which explains why the site let go at two cards loaded rather than four, and why the ~25 W OPNsense box died with it. ⚠ Will not self-recover (tripped needs a human, dead needs replacing) — do NOT poll FV. ⚠ Do NOT use surge-only outlets to exceed a UPS rating: both banks share one NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current limit. ⭐ Recover /tank/aimodels/flash-next-mtp-bench/power.log FIRST — all four cards every 10 s to the cut, on /tank not in a container, and the ONLY load measurement that exists. ⚠ 19 of 30 gateway aliases dark and NO local fallback — every free local model was on fv-ml1, irv-ml1 runs no chat seat at all; the only non-fv chat backends are paid, and any coverage must be a NEW opt-in alias, never a silent repoint. ⭐ OOB design gap: OPNsense-as-subnet-router covers box-down/gateway-up and nothing for a site-wide loss, since the BMC's only route out is that gateway. ⚠ Recovery hazard: ten restart: unless-stopped vLLM containers will all load at once on power-up — mask Docker first, then compose up -d seat by seat (which also finishes the stale-homepage-label fix, since labels attach only at creation). → docs/runbooks/fv-site-dark-20260913.md, persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md