policy: cap GPU power limits at build time — 200 W for the other colo's cards
Operator directive, and the right generalisation of the FV outage: decide the power envelope first and size the cards into it, rather than installing cards and discovering the constraint by tripping it. Four cards at 200 W is 800 W, which fits a real circuit with a real UPS and headroom. Records three things to settle before it is a plan. First, 200 W may sit below the card's enforceable floor: nvidia-smi -pl is bounded by Min Power Limit, often around half of TGP on a high-TGP part, and a sub-floor request is refused -- quietly, depending on how it is scripted. Run nvidia-smi -q -d POWER before any build planning depends on the number. Second, the 600 W figure wants confirming against the actual SKU. The Ada parts do not land there -- RTX 6000 Ada is 300 W, L40/L40S 300/350 W, 4090 450 W -- while 600 W is Blackwell RTX PRO 6000 Workstation territory, so these may be Blackwell or the figure may be a two-card total. Read it off the device rather than a spec sheet. Third, the workload asymmetry is in this fleet's favour: decode is memory-bandwidth-bound and tolerates a cap far better than training does, with a concave perf/watt curve where 60-70% of TGP costs roughly 10-15% of throughput. A cap to a third of TGP is deeper into the steep region; measure it on the first card rather than predicting, and expect prefill-heavy and training work to pay more than a serving seat. And persist the cap. A hand-set limit holds until the next reboot and then silently stops holding, which is the worst shape available given that the thing rebooting the box is likely to be the power event the cap existed to prevent.
This commit is contained in:
@@ -345,3 +345,42 @@ probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
||||
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
|
||||
up to the cut. Recover it after boot — it is the only measurement of what the load
|
||||
actually drew, and it survives on `/tank`, not in the container.
|
||||
|
||||
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
|
||||
|
||||
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
|
||||
> power limited to 200w"
|
||||
|
||||
The generalised lesson from this outage: **decide the power envelope first and size the
|
||||
cards into it**, rather than installing cards and discovering the constraint by tripping
|
||||
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
|
||||
|
||||
Three things to settle before that is a plan:
|
||||
|
||||
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
|
||||
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
|
||||
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
|
||||
is scripted, refused quietly. **First command on the new hardware:**
|
||||
|
||||
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
|
||||
|
||||
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
|
||||
feed, not from the cap.
|
||||
|
||||
2. ⚠ **The 600 W figure wants confirming against the actual SKU.** The Ada parts do not
|
||||
land there — RTX 6000 Ada is 300 W, L40/L40S 300/350 W, RTX 4090 450 W. 600 W is
|
||||
*Blackwell* RTX PRO 6000 Workstation Edition territory. So either these are Blackwell
|
||||
rather than Ada, or 600 W is a two-card/total figure. Read it off the device
|
||||
(`nvidia-smi -q -d POWER`), not off a spec sheet or a recollection.
|
||||
|
||||
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
||||
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
||||
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
|
||||
33% of TGP is deeper into the steep region — the cost is real and should be measured
|
||||
on the first card rather than predicted, and it will hurt a prefill-heavy or training
|
||||
workload considerably more than a serving seat.
|
||||
|
||||
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
|
||||
stops holding — the worst possible failure shape, because the thing that reboots the box
|
||||
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
|
||||
mode, ordered before Docker starts.
|
||||
|
||||
Reference in New Issue
Block a user