correct the hardware: fv-ml1 is 4x Blackwell Max-Q 300W, ana-ml3 is 2x Ada RTX 6000 — and four cards is a breaker problem
Operator clarification, and it separates two boxes I had been conflating. fv-ml1 is 4x Blackwell RTX PRO 6000 Max-Q at 300 W each (Max-Q being the reduced-TGP SKU; the Workstation Edition is the 600 W part), 391 GB VRAM, deployed and currently dark. ana-ml3 is 2x Ada Generation RTX 6000 at 300 W, 96 GB VRAM, not yet deployed. The 200 W cap directive is ana-ml3's. With the TGP known, the outage stops being a vague 'undersized' and acquires a mechanism: two Max-Q cards at 300 W is ~600 W of card, plus a host carrying 566 GB of RAM, drives, fans and PSU conversion loss at perhaps 200-350 W, against an Eaton 1500 VA's real ~900-1200 W. That lands at or just over the rating, which is precisely what explains a full day of service on one card and failure minutes into the second. The host term is the only one being guessed; idle-at-the-plug measures it directly. It also surfaces something that is not a UPS question at all. Four cards at 300 W plus ~300 W of host is ~1500 W against a 15 A circuit's 1440 W continuous derating, so four cards uncapped is marginal on the breaker with no UPS in the path. Capping therefore belongs at fv-ml1 as well as ana-ml3, or fv-ml1 needs a 20 A feed -- and worth noting today's incident only ever had two of the four cards working. ana-ml3's placement constraints sharpen too: sm_89 has native FP8 but no NVFP4, so the in-house NVFP4 quants stay at FV, and at 96 GB total it cannot host the Flash-Next seat at all -- that needs 74 GiB resident on a single card, and the offload moves the n-gram table rather than the experts.
This commit is contained in:
@@ -351,7 +351,15 @@ probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
||||
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
|
||||
> power limited to 200w"
|
||||
>
|
||||
> **CONFIRMED 2026-09-13: they are RTX 6000 Ada — 300 W, not 600 W.**
|
||||
> **CLARIFIED BY OPERATOR 2026-09-13 — the two boxes are different hardware:**
|
||||
>
|
||||
> | box | cards | TGP each | VRAM total | status |
|
||||
> |---|---|---|---|---|
|
||||
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
|
||||
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
|
||||
>
|
||||
> The 200 W cap directive applies to **ana-ml3**. It very likely wants applying to
|
||||
> fv-ml1 as well — see the four-card circuit arithmetic below.
|
||||
|
||||
The generalised lesson from this outage: **decide the power envelope first and size the
|
||||
cards into it**, rather than installing cards and discovering the constraint by tripping
|
||||
@@ -369,15 +377,41 @@ Three things to settle before that is a plan:
|
||||
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
|
||||
feed, not from the cap.
|
||||
|
||||
2. ✅ **RESOLVED — RTX 6000 Ada, 300 W.** So 200 W is a cap to **67% of TGP**, which is
|
||||
the favourable part of the curve, not the severe 33% cap a 600 W part would have
|
||||
implied. The floor concern largely goes away too: 200 W was borderline against a
|
||||
600 W card's minimum and is very unlikely to sit below a 300 W card's. Still worth the
|
||||
one command, but expect it to take.
|
||||
2. ✅ **RESOLVED — both card types are 300 W.** So 200 W is a cap to **67% of TGP**, the
|
||||
favourable part of the concave curve, not the severe 33% cap a 600 W part would have
|
||||
implied. The enforceable-floor concern largely goes away too: 200 W was borderline
|
||||
against a 600 W card's minimum and is very unlikely to sit below a 300 W card's. Worth
|
||||
the one command; expect it to take.
|
||||
|
||||
⭐ **The protective value is real**: 4 x 300 W uncapped is ~1200 W of card, which is
|
||||
roughly the neighbourhood that just overwhelmed a 1500 VA unit at FV with only TWO
|
||||
Blackwell cards drawing. Capping to 800 W makes a repeat a non-event.
|
||||
**ana-ml3: 2 x 200 W = 400 W of card.** Modest, and pointed — ana-ml3 lands in the
|
||||
**Anaheim** rack whose breaker tripped on 2026-08-26 and 2026-09-11, one of those
|
||||
caused by this very chassis before it relocated. The cap there is remediation of a
|
||||
circuit with a track record, not precaution.
|
||||
|
||||
### ⭐⭐ The outage arithmetic, now that the TGP is known
|
||||
|
||||
2 x Blackwell Max-Q @ 300 W ~ 600 W of card under load
|
||||
host (board, 566 GB RAM, drives,
|
||||
fans, PSU conversion loss) ~ 200-350 W <-- UNMEASURED, the gap
|
||||
-----------
|
||||
~ 800-950 W
|
||||
|
||||
Eaton 1500 VA real watt rating ~ 900-1200 W depending on model
|
||||
|
||||
**At or just over the line** — and this is what a vague "undersized" could not explain:
|
||||
why it ran a full day on one card (~500-650 W, comfortably inside) and died minutes into
|
||||
the second (~800-950 W, at or past the rating). The host term is the only one being
|
||||
guessed at, and idle-at-the-plug with all seats down measures it directly.
|
||||
|
||||
### ⚠⚠ FOUR cards is a BREAKER problem, not a UPS problem
|
||||
|
||||
4 x 300 W card + ~300 W host ~ 1500 W
|
||||
15 A circuit, 80% continuous = 1440 W
|
||||
|
||||
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all.** So
|
||||
capping belongs at **fv-ml1 too**, not only ana-ml3. If the four-card ammeter reading
|
||||
confirms it, fv-ml1 needs a per-card cap or a 20 A feed before anyone loads all four
|
||||
again — and note that today's incident only ever had TWO cards working.
|
||||
|
||||
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
||||
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
||||
|
||||
Reference in New Issue
Block a user