ea818380ff
Operator supplied the topology: "the entire rack is on the same circuit,
public ip is served by firewall on the same circuit. load tripped
breaker, entire rack goes dark."
That inverts the argument I committed one commit ago in 3cc55b4. I held
the load attribution partly because a site-wide blackout looked like too
large a blast radius for one GPU box. It is not — it is exactly the
blast radius a shared circuit predicts. The observation I treated as
unexplained was the signature of the mechanism.
Kept the wrong version visible with a superseded marker rather than
editing it away, per the superseded-claims convention: the reasoning
circulated in two althing messages and a peer adopted it.
Holding the attribution as unmeasured was still correct; the argument
used to hold it was reasoning about wiring neither side had seen and
neither side asked for.
Records the finding that outlives the incident: the same circuit powers
ana-ml2, ana-gw, ana-wg, ana-nas and the BMCs, so the load takes every
recovery path with it. Composed with ana-gw's WAN admin being closed
(tunnel-only) and the tunnel terminating on ana-wg, there is no remote
management path to Anaheim during a rack power event, in-band or out.
Open question for the weekend triage: is any of that rack on UPS or a
second circuit?
124 lines
6.7 KiB
Markdown
124 lines
6.7 KiB
Markdown
# `[2026-08-27]` Anaheim tripped a power breaker — and four guests including the NAS never came back
|
|
|
|
Site-wide outage. **ana-ml2 was down 1h58m** — power lost between **18:14:45 and 18:17:00 PDT**,
|
|
boot at **20:13:14** (the "~90 minutes" first recorded here was an estimate; the window is now
|
|
pinned by run 3c's training log and the hourly cron that never ran — see
|
|
[[2026-08-27-run3c-launch-count-reconstruction]]). **Operator-confirmed cause: a tripped power breaker**, not a
|
|
fault and not the tunnel. The discriminator that established scope: `ana-srv1`
|
|
(38.120.12.44:443, Anaheim's PUBLIC address) was dark **from the internet**, so it was not the
|
|
NH3↔ANA IPsec tunnel stranding NH3 — the site was not answering on any path. `ana-ml2` returned
|
|
with `up 1 min`, confirming a hard power event.
|
|
|
|
## ⚠ THE DURABLE FINDING — `onboot` was unset on four guests
|
|
|
|
pfi-pve came back and auto-started everything **except**:
|
|
|
|
CT109 ana-nas the storage SPOF
|
|
CT113 ana-wg the WireGuard remote-access path
|
|
CT112 ana-filebot
|
|
VM106 corviduo-dev
|
|
|
|
All four had `onboot` unset. **Recovery was manual and would have been manual every time** —
|
|
including for the NAS that postgres/PBS/cross-site-restic depend on, and the WireGuard host
|
|
that is the way in when the site misbehaves.
|
|
|
|
**FIXED, with dependency ordering** (operator-authorised):
|
|
|
|
CT109 ana-nas onboot=1 order=1,up=45 <- first; 45s for NFS to SERVE
|
|
CT113 ana-wg onboot=1 order=2 <- remote access before anything can fail
|
|
VM104/105 Mongo/Postgres order=3,up=60 (pre-existing)
|
|
VM102 ANA-Docker order=4 (pre-existing)
|
|
VM101 ANA-DC order=5,up=120 (pre-existing)
|
|
CT112 ana-filebot onboot=1 order=10
|
|
VM106 corviduo-dev onboot=1 order=10
|
|
|
|
Every guest on pfi-pve now auto-starts. ana-nas precedes the databases deliberately; the
|
|
`up=45` is for NFS to be *serving*, not merely for the container to be *running* — the exact
|
|
distinction that killed `rest-server` on ana-docker, which came up before the NAS existed,
|
|
found nothing to serve, and exited 255.
|
|
|
|
## ⚠ `/tank` came back DEGRADED — a disk is genuinely gone
|
|
|
|
tank DEGRADED, raidz2-0, 7 devices ONLINE
|
|
9477159196657038377 FAULTED was /dev/nvme4n1p1
|
|
errors: No known data errors
|
|
|
|
**Only 7 physical NVMe present where the pool expects 8** — checked, so not renumbering. One
|
|
drive did not re-enumerate. raidz2 carries two disks of parity; one is spent. Operator taking
|
|
it; chassis is a Supermicro AS-4125GS-TNRT2 with PCIe hot-plug slots, so a swap should not
|
|
need a power-down.
|
|
|
|
## ⚠ `/mnt/smithy` is manual by design — it will be missing after EVERY reboot
|
|
|
|
Not in fstab, and **deliberately so**: a cross-site NFS entry can hang boot on a GPU host, and
|
|
it is `soft` rather than `hard` because ana-ml2 is cross-site from that NAS and a hard mount
|
|
turns a link blip into unkillable D-state. Remount with the recorded spec, do NOT "fix" it into
|
|
fstab:
|
|
|
|
sudo mount -t nfs4 -o ro,soft,timeo=30,retrans=3,proto=tcp,vers=4.1 \
|
|
10.100.50.50:/volume1/smithy /mnt/smithy
|
|
|
|
Full rationale: [[2026-08-23-smithy-mount-ana-ml2]].
|
|
|
|
## Power capacity is now the open item
|
|
|
|
Operator: *"we'll triage this weekend, probably shut down some seats."* ana-ml2 alone was
|
|
pulling ~600 W across both GPUs at their 300 W caps during training. `gen` stays up by
|
|
instruction; everything else on that box is idle.
|
|
|
|
MEASURED a breaker tripped (operator-confirmed)
|
|
MEASURED power lost between 18:14:45 and 18:17:00 PDT
|
|
MEASURED the WHOLE Anaheim site went dark from the internet
|
|
TOPOLOGY ONE CIRCUIT FEEDS THE ENTIRE RACK, including the firewall serving
|
|
the public IP (operator, 2026-08-27)
|
|
|
|
OPERATOR "load tripped breaker, entire rack goes dark." Consistent with every
|
|
DETERMINATION observation. Not a per-circuit measurement, and no longer needs to be —
|
|
the operator has the site knowledge and has ruled.
|
|
|
|
⚠ **SUPERSEDED — the blast-radius objection I raised here was WRONG, and it was committed.**
|
|
An earlier revision of this section (commit `3cc55b4`) argued that *"a site-wide blackout is a
|
|
larger blast radius than one GPU box's ~600 W accounts for on its own: if ana-ml2's draw were
|
|
the whole story, why did ana-nas, ana-wg and the public address go dark with it?"*
|
|
|
|
**One shared circuit inverts that completely.** Site-wide darkness is not too large a blast
|
|
radius for a single overloaded box — **it is exactly the blast radius the topology predicts.**
|
|
The fact I treated as unexplained was the signature of the mechanism.
|
|
|
|
Holding the attribution as unmeasured was still right (there is no per-circuit meter); the
|
|
*argument* used to hold it was reasoning confidently about wiring neither I nor
|
|
brokkr-smithy-dev had ever seen, and neither of us asked for it. Kept visible rather than
|
|
silently corrected so the old reasoning stops circulating.
|
|
|
|
## ⚠ THE FINDING THAT OUTLIVES THE INCIDENT — the recovery path shares the failure domain
|
|
|
|
One circuit powers, together:
|
|
|
|
ana-ml2 the load most likely to trip it
|
|
ana-gw the firewall serving the public IP (38.120.12.44)
|
|
ana-wg the WireGuard host — the way in when the site misbehaves
|
|
ana-nas the storage SPOF
|
|
the BMCs ana-ml2 IPMI 10.250.250.50, pfi-pve iDRAC 10.250.250.30 — same rack
|
|
|
|
**So the thing that fails takes every path used to fix it with it.** Compounding two facts
|
|
already in the record separately: ana-gw's WAN admin is closed, so management is tunnel-only
|
|
([[2026-08-24-ana-gw-admin-closed-acme-disabled]]) — and the tunnel terminates on ana-wg, on
|
|
this circuit. **During a rack power event there is no remote management path to Anaheim at
|
|
all**, in-band or out-of-band. Recovery is: wait for power. Until 2026-08-27 it was worse —
|
|
ana-wg had `onboot` unset and would not have come back on its own.
|
|
|
|
Open question for the weekend triage, and the cheapest fix if the answer is no: **is any of
|
|
that rack on UPS or a second circuit?** Nothing in this repo records one, and the site went
|
|
dark instantly and stayed dark 1h58m, which is behaviourally indistinguishable from no
|
|
coverage for these loads. Putting *only* ana-gw + ana-wg + one BMC on separate power would
|
|
mean a load trip no longer costs the way in.
|
|
|
|
Load-shed on ana-ml2 is still the right first move and it is cheap. It reduces how often the
|
|
breaker trips; it does not change what a trip costs.
|
|
|
|
## Blast radius beyond us
|
|
|
|
heid lost **both gateway-routed arms of a four-arm panel** mid-dispatch and discovered the
|
|
outage by losing half a panel. That report produced the single most valuable artifact of the
|
|
incident — see [[2026-08-27-empty-response-as-a-datum]].
|