Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md
T

1.9 KiB
Raw Blame History

[2026-09-24] NH3 power outage — recovered; three latent gaps closed, one needs a hand

Power at NH3 dropped and returned; every NH3 box booted at 1233–1234 PT. The core recovered unaided (post office, headscale, egress proxy, DNS, WAN IP unchanged at 70.230.226.88). Three things did not, each a latent config gap that only a site-wide power event exposes:

  1. pbs-nh3 (VM 105) had no onboot flag and stayed down. Set onboot=1 and started it; datastore NFS mounted (its line uses bg). qm guest exec 105 works (guest agent on) — the root path, since infra-ops is not provisioned there.
  2. NFS boot race. nh3-nas is the slowest box to serve NFS, so every plain fstab NFS line failed at boot (nh3-docker /mnt/compose /mnt/backup, nh3-dev /mnt/backup). nh3-dev's /mnt/smithy already had x-systemd.automount and self-healed on its next access (1234 failed, 1309 mounted, no hand) — the existence proof. playbooks/nh3-nfs-automount.yaml brought the rest to that shape (_netdev,nofail,x-systemd.automount, x-systemd.mount-timeout=30), commit 1cbde50. ⚠ nh3-dev's installer cdrom fstab line is a pre-existing findmnt --verify error; the playbook judges only errors a rewrite ADDS, after its first run aborted on it.
  3. pfi-gx10 (bare metal, no BMC) did not power on — ASUS ships "Restore AC Power Loss" = stay off. See [[2026-09-24-gx10-ac-restore-patch]].

Also: every Claude session on nh3-dev died with the reboot except infra-ops, infra-hermes, jekyll. Relaunch is Prime's call (dev-launch). nh3-laser (VM 104) is on-demand and stays off by Prime's ruling; servers/nh3-pve/README.md now carries an expected-after-power-loss column per guest so the next triage does not flag it.

The headscale-ddns failures 1141–1202 were the WAN being down, correctly refused (no public v4); every run since 1247 reads unchanged.