memory: snapshot — Zigbee2MQTT + HA-MQTT route fix, Blender on GPU 3 (MCP per task + blender-run), semif 0.1.4, Worldtree reward config live and pushed; 3 entries archived; next: 2026-09-28 restic freshness check

This commit is contained in:
vh
2026-09-28 09:57:41 -07:00
parent 6f0c480693
commit 1756b891c6
5 changed files with 38 additions and 31 deletions
+27
View File
@@ -10629,3 +10629,30 @@ Canonical configs/restic/esh-vm-db and playbooks/esh-vm-db-restic-repair.yaml.
No DB restarts or auth-policy changes; changes saved locally, no commit.
_Archived 2026-09-26._
# `[2026-09-13]` FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load
⚠⚠⚠ **FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load test; all other sites healthy. Operator's leading hypothesis: the 1500 VA Eaton UPS overloaded and DIED.** It fits better than a breaker trip because a UPS's output rating sits far below the circuit's, making it the first protective device to give — which explains why the site let go at **two** cards loaded rather than four, and why the ~25 W OPNsense box died with it. ⚠ **Will not self-recover** (tripped needs a human, dead needs replacing) — do NOT poll FV. ⚠ **Do NOT use surge-only outlets to exceed a UPS rating**: both banks share one NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current limit. ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST** — all four cards every 10 s to the cut, on `/tank` not in a container, and the ONLY load measurement that exists. ⚠ **19 of 30 gateway aliases dark and NO local fallback** — every free local model was on fv-ml1, irv-ml1 runs no chat seat at all; the only non-fv chat backends are paid, and any coverage must be a NEW opt-in alias, never a silent repoint. ⭐ **OOB design gap**: OPNsense-as-subnet-router covers box-down/gateway-up and nothing for a site-wide loss, since the BMC's only route out is that gateway. ⚠ Recovery hazard: ten `restart: unless-stopped` vLLM containers will all load at once on power-up — mask Docker first, then `compose up -d` seat by seat (which also finishes the stale-homepage-label fix, since labels attach only at creation). → `docs/runbooks/fv-site-dark-20260913.md`, `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
_Archived 2026-09-28._
# `[2026-09-13]` Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.
**Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.** The sweep allowlist was built from files that mention the HOST and a `homepage.href` mentions only an IP, so every label-only stack fell outside it by construction; 24 files fixed plus host copies, allowlist extended with how to derive it next time. Two bugs fell out: `deploy-stack.sh` **rejected any stack name containing a dot** (so `qwen3.5-122b`/`qwopus3.5-122b`/`mistral-medium-3.5` could not be deployed at all), and scriberr's CORS allowlist held only the dead IP and the dead `scriberr.ana.internal`. ⚠ Incomplete: the 10 running containers were never recreated and a plain power-on will NOT apply labels — the staged `compose up -d` recovery does. Eight stacks deliberately not pushed (real host drift); three of those are untracked host-only stacks. Commits `3132a16`, `969a1b6`, `d79f104`.
_Archived 2026-09-28._
# FV→ANA NAT repaired safely
Operator approved narrow fix, caution not to strand subnet. OPNsense MESH/opt6
(tailscale0 100.64.0.8) lacked outbound NAT for forwarded LAN traffic. Temporary
fv-ml1→hub /32 NAT proved diagnosis; persisted hybrid NAT rule source
10.251.50.54/32 destination 10.250.0.0/16 translate interface address. Existing
WAN NAT retained, filter rules byte-identical, no routes/mesh/host changes.
Both stages guarded by independent rollback timers, disarmed after verification.
Verified hub HTTP200, ANA PostgreSQL TCP, Internet HTTPS200, reverse SSH, gateway
management; Beszel18/18 up. Agent intentionally stops SSH45876 when WebSocket
connects. BMC ping failed with no pre-change baseline; no BMC-health claim.
Other FV sources and other remote subnets not covered by this narrow fix.
Backup + rollback helper on gateway /root/fv-nat-repair-20260913. Full details:
docs/runbooks/fv-to-ana-nat.md. No commit made.
_Archived 2026-09-28._