policy: cap GPU power limits at build time — 200 W for the other colo's cards

Operator directive, and the right generalisation of the FV outage: decide the power
envelope first and size the cards into it, rather than installing cards and discovering
the constraint by tripping it. Four cards at 200 W is 800 W, which fits a real circuit
with a real UPS and headroom.

Records three things to settle before it is a plan. First, 200 W may sit below the
card's enforceable floor: nvidia-smi -pl is bounded by Min Power Limit, often around
half of TGP on a high-TGP part, and a sub-floor request is refused -- quietly, depending
on how it is scripted. Run nvidia-smi -q -d POWER before any build planning depends on
the number.

Second, the 600 W figure wants confirming against the actual SKU. The Ada parts do not
land there -- RTX 6000 Ada is 300 W, L40/L40S 300/350 W, 4090 450 W -- while 600 W is
Blackwell RTX PRO 6000 Workstation territory, so these may be Blackwell or the figure
may be a two-card total. Read it off the device rather than a spec sheet.

Third, the workload asymmetry is in this fleet's favour: decode is
memory-bandwidth-bound and tolerates a cap far better than training does, with a concave
perf/watt curve where 60-70% of TGP costs roughly 10-15% of throughput. A cap to a third
of TGP is deeper into the steep region; measure it on the first card rather than
predicting, and expect prefill-heavy and training work to pay more than a serving seat.

And persist the cap. A hand-set limit holds until the next reboot and then silently
stops holding, which is the worst shape available given that the thing rebooting the box
is likely to be the power event the cap existed to prevent.
This commit is contained in:
vh
2026-09-13 00:23:20 -07:00
parent b6335bf6ad
commit 3e61d7d4e0
2 changed files with 41 additions and 0 deletions
+39
View File
@@ -345,3 +345,42 @@ probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
up to the cut. Recover it after boot — it is the only measurement of what the load
actually drew, and it survives on `/tank`, not in the container.
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
> power limited to 200w"
The generalised lesson from this outage: **decide the power envelope first and size the
cards into it**, rather than installing cards and discovering the constraint by tripping
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
Three things to settle before that is a plan:
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
is scripted, refused quietly. **First command on the new hardware:**
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
feed, not from the cap.
2. ⚠ **The 600 W figure wants confirming against the actual SKU.** The Ada parts do not
land there — RTX 6000 Ada is 300 W, L40/L40S 300/350 W, RTX 4090 450 W. 600 W is
*Blackwell* RTX PRO 6000 Workstation Edition territory. So either these are Blackwell
rather than Ada, or 600 W is a two-card/total figure. Read it off the device
(`nvidia-smi -q -d POWER`), not off a spec sheet or a recollection.
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
33% of TGP is deeper into the steep region — the cost is real and should be measured
on the first card rather than predicted, and it will hurt a prefill-heavy or training
workload considerably more than a serving seat.
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
stops holding — the worst possible failure shape, because the thing that reboots the box
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
mode, ordered before Docker starts.