memory: separate the measured breaker trip from the load hypothesis

The record read "power capacity is the open item" next to ana-ml2's
~600 W, which reads as a cause. It is not one. The trip and its timing
are measured; the attribution to the training load is the operator's
working read and the reason for the weekend triage.

The observation that makes the single-load story incomplete on its own
terms: a site-wide blackout is a larger blast radius than one GPU box
accounts for. If ana-ml2's draw were the whole story, ana-nas, ana-wg
and the public address would not have gone dark with it.

Shedding seats may still be the right first move and it is cheap. That
is not the same as having identified what loaded the circuit, and the
distinction matters going into a triage that will act on it.
This commit is contained in:
2026-08-27 11:01:44 -07:00
parent 88d79375f7
commit 3cc55b4b40
@@ -66,6 +66,21 @@ Operator: *"we'll triage this weekend, probably shut down some seats."* ana-ml2
pulling ~600 W across both GPUs at their 300 W caps during training. `gen` stays up by
instruction; everything else on that box is idle.
**Keep the measured and the hypothesised apart here — the load attribution is NOT established:**
MEASURED a breaker tripped (operator-confirmed)
MEASURED power lost between 18:14:45 and 18:17:00 PDT
MEASURED the WHOLE Anaheim site went dark from the internet
HYPOTHESIS the training load tripped it — the operator's working read and the reason
for the weekend triage. There is no per-circuit meter on that panel, and
a site-wide blackout is a larger blast radius than one GPU box's ~600 W
accounts for on its own: if ana-ml2's draw were the whole story, why did
ana-nas, ana-wg and the public address go dark with it?
**Held as the right thing to triage first, not as the cause.** Shedding GPU seats may well be
correct and is cheap; it is not the same as having identified what loaded the circuit.
## Blast radius beyond us
heid lost **both gateway-routed arms of a four-arm panel** mid-dispatch and discovered the