# `[2026-08-27]` Anaheim tripped a power breaker — and four guests including the NAS never came back Site-wide outage. **ana-ml2 was down 1h58m** — power lost between **18:14:45 and 18:17:00 PDT**, boot at **20:13:14** (the "~90 minutes" first recorded here was an estimate; the window is now pinned by run 3c's training log and the hourly cron that never ran — see [[2026-08-27-run3c-launch-count-reconstruction]]). **Operator-confirmed cause: a tripped power breaker**, not a fault and not the tunnel. The discriminator that established scope: `ana-srv1` (38.120.12.44:443, Anaheim's PUBLIC address) was dark **from the internet**, so it was not the NH3↔ANA IPsec tunnel stranding NH3 — the site was not answering on any path. `ana-ml2` returned with `up 1 min`, confirming a hard power event. ## ⚠ THE DURABLE FINDING — `onboot` was unset on four guests pfi-pve came back and auto-started everything **except**: CT109 ana-nas the storage SPOF CT113 ana-wg the WireGuard remote-access path CT112 ana-filebot VM106 corviduo-dev All four had `onboot` unset. **Recovery was manual and would have been manual every time** — including for the NAS that postgres/PBS/cross-site-restic depend on, and the WireGuard host that is the way in when the site misbehaves. **FIXED, with dependency ordering** (operator-authorised): CT109 ana-nas onboot=1 order=1,up=45 <- first; 45s for NFS to SERVE CT113 ana-wg onboot=1 order=2 <- remote access before anything can fail VM104/105 Mongo/Postgres order=3,up=60 (pre-existing) VM102 ANA-Docker order=4 (pre-existing) VM101 ANA-DC order=5,up=120 (pre-existing) CT112 ana-filebot onboot=1 order=10 VM106 corviduo-dev onboot=1 order=10 Every guest on pfi-pve now auto-starts. ana-nas precedes the databases deliberately; the `up=45` is for NFS to be *serving*, not merely for the container to be *running* — the exact distinction that killed `rest-server` on ana-docker, which came up before the NAS existed, found nothing to serve, and exited 255. ## ⚠ `/tank` came back DEGRADED — a disk is genuinely gone tank DEGRADED, raidz2-0, 7 devices ONLINE 9477159196657038377 FAULTED was /dev/nvme4n1p1 errors: No known data errors **Only 7 physical NVMe present where the pool expects 8** — checked, so not renumbering. One drive did not re-enumerate. raidz2 carries two disks of parity; one is spent. Operator taking it; chassis is a Supermicro AS-4125GS-TNRT2 with PCIe hot-plug slots, so a swap should not need a power-down. ## ⚠ `/mnt/smithy` is manual by design — it will be missing after EVERY reboot Not in fstab, and **deliberately so**: a cross-site NFS entry can hang boot on a GPU host, and it is `soft` rather than `hard` because ana-ml2 is cross-site from that NAS and a hard mount turns a link blip into unkillable D-state. Remount with the recorded spec, do NOT "fix" it into fstab: sudo mount -t nfs4 -o ro,soft,timeo=30,retrans=3,proto=tcp,vers=4.1 \ 10.100.50.50:/volume1/smithy /mnt/smithy Full rationale: [[2026-08-23-smithy-mount-ana-ml2]]. ## Power capacity is now the open item Operator: *"we'll triage this weekend, probably shut down some seats."* ana-ml2 alone was pulling ~600 W across both GPUs at their 300 W caps during training. `gen` stays up by instruction; everything else on that box is idle. ## Blast radius beyond us heid lost **both gateway-routed arms of a four-arm panel** mid-dispatch and discovered the outage by losing half a panel. That report produced the single most valuable artifact of the incident — see [[2026-08-27-empty-response-as-a-datum]].