The record read "power capacity is the open item" next to ana-ml2's ~600 W, which reads as a cause. It is not one. The trip and its timing are measured; the attribution to the training load is the operator's working read and the reason for the weekend triage. The observation that makes the single-load story incomplete on its own terms: a site-wide blackout is a larger blast radius than one GPU box accounts for. If ana-ml2's draw were the whole story, ana-nas, ana-wg and the public address would not have gone dark with it. Shedding seats may still be the right first move and it is cheap. That is not the same as having identified what loaded the circuit, and the distinction matters going into a triage that will act on it.
4.6 KiB
[2026-08-27] Anaheim tripped a power breaker — and four guests including the NAS never came back
Site-wide outage. ana-ml2 was down 1h58m — power lost between 18:14:45 and 18:17:00 PDT,
boot at 20:13:14 (the "~90 minutes" first recorded here was an estimate; the window is now
pinned by run 3c's training log and the hourly cron that never ran — see
2026-08-27-run3c-launch-count-reconstruction). Operator-confirmed cause: a tripped power breaker, not a
fault and not the tunnel. The discriminator that established scope: ana-srv1
(38.120.12.44:443, Anaheim's PUBLIC address) was dark from the internet, so it was not the
NH3↔ANA IPsec tunnel stranding NH3 — the site was not answering on any path. ana-ml2 returned
with up 1 min, confirming a hard power event.
⚠ THE DURABLE FINDING — onboot was unset on four guests
pfi-pve came back and auto-started everything except:
CT109 ana-nas the storage SPOF
CT113 ana-wg the WireGuard remote-access path
CT112 ana-filebot
VM106 corviduo-dev
All four had onboot unset. Recovery was manual and would have been manual every time —
including for the NAS that postgres/PBS/cross-site-restic depend on, and the WireGuard host
that is the way in when the site misbehaves.
FIXED, with dependency ordering (operator-authorised):
CT109 ana-nas onboot=1 order=1,up=45 <- first; 45s for NFS to SERVE
CT113 ana-wg onboot=1 order=2 <- remote access before anything can fail
VM104/105 Mongo/Postgres order=3,up=60 (pre-existing)
VM102 ANA-Docker order=4 (pre-existing)
VM101 ANA-DC order=5,up=120 (pre-existing)
CT112 ana-filebot onboot=1 order=10
VM106 corviduo-dev onboot=1 order=10
Every guest on pfi-pve now auto-starts. ana-nas precedes the databases deliberately; the
up=45 is for NFS to be serving, not merely for the container to be running — the exact
distinction that killed rest-server on ana-docker, which came up before the NAS existed,
found nothing to serve, and exited 255.
⚠ /tank came back DEGRADED — a disk is genuinely gone
tank DEGRADED, raidz2-0, 7 devices ONLINE
9477159196657038377 FAULTED was /dev/nvme4n1p1
errors: No known data errors
Only 7 physical NVMe present where the pool expects 8 — checked, so not renumbering. One drive did not re-enumerate. raidz2 carries two disks of parity; one is spent. Operator taking it; chassis is a Supermicro AS-4125GS-TNRT2 with PCIe hot-plug slots, so a swap should not need a power-down.
⚠ /mnt/smithy is manual by design — it will be missing after EVERY reboot
Not in fstab, and deliberately so: a cross-site NFS entry can hang boot on a GPU host, and
it is soft rather than hard because ana-ml2 is cross-site from that NAS and a hard mount
turns a link blip into unkillable D-state. Remount with the recorded spec, do NOT "fix" it into
fstab:
sudo mount -t nfs4 -o ro,soft,timeo=30,retrans=3,proto=tcp,vers=4.1 \
10.100.50.50:/volume1/smithy /mnt/smithy
Full rationale: 2026-08-23-smithy-mount-ana-ml2.
Power capacity is now the open item
Operator: "we'll triage this weekend, probably shut down some seats." ana-ml2 alone was
pulling ~600 W across both GPUs at their 300 W caps during training. gen stays up by
instruction; everything else on that box is idle.
⚠ Keep the measured and the hypothesised apart here — the load attribution is NOT established:
MEASURED a breaker tripped (operator-confirmed)
MEASURED power lost between 18:14:45 and 18:17:00 PDT
MEASURED the WHOLE Anaheim site went dark from the internet
HYPOTHESIS the training load tripped it — the operator's working read and the reason
for the weekend triage. There is no per-circuit meter on that panel, and
a site-wide blackout is a larger blast radius than one GPU box's ~600 W
accounts for on its own: if ana-ml2's draw were the whole story, why did
ana-nas, ana-wg and the public address go dark with it?
Held as the right thing to triage first, not as the cause. Shedding GPU seats may well be correct and is cheap; it is not the same as having identified what loaded the circuit.
Blast radius beyond us
heid lost both gateway-routed arms of a four-arm panel mid-dispatch and discovered the outage by losing half a panel. That report produced the single most valuable artifact of the incident — see 2026-08-27-empty-response-as-a-datum.