DCGM_CONFIG_POWER_BUDGET_GROUP ('the power budget for the entire group') exists
alongside DCGM_CONFIG_POWER_CAP_INDIVIDUAL, so the concept is first-class. The docs do
not state how a group budget is distributed, and the deduction is that it cannot be
anything exotic: the only enforcement primitive underneath is NVML's per-GPU
nvmlDeviceSetPowerManagementLimit and there is no bank-level register, so any group
budget resolves to N per-GPU writes. Static even division is one write each; 'each card
free until they are all loaded' requires continuous re-writing, which is a control loop
rather than a hardware feature. Worth a ten-minute test when the box returns, in case
NVIDIA already runs that loop.
Records the design constraint that matters more than the logic: power readings lag and
-pl application takes tens of milliseconds, so a reactive daemon overshoots during a load
ramp -- and the ramp is the dangerous moment, being the same all-cards-at-once shape as
this box's ten restart:unless-stopped containers starting together. So any such loop must
be safe-by-default and opportunistic upward: boot at budget/N, only ever raise after
observing idle neighbours. Inverted, it works for weeks and then fails on precisely the
event it existed to prevent.
And the reason to defer it: 4x250 W is already 1000 W, so the static cap is the
conservative floor of the dynamic scheme rather than an alternative. The daemon's entire
contribution is the one-card-busy case, worth perhaps 5% throughput, which is rare for a
serving fleet that puts one seat per card and common only for a training window.
534 lines
29 KiB
Markdown
534 lines
29 KiB
Markdown
# FV site dark — 2026-09-13 ~06:56Z
|
|
|
|
**Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley.**
|
|
Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one
|
|
needs replacing, so there is no recovery to wait for and no point polling the site.
|
|
Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is
|
|
down so recovery does not have to be reconstructed from memory.
|
|
|
|
## What is down
|
|
|
|
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
|
|
|
|
| target | result |
|
|
|---|---|
|
|
| `fv-ml1` 10.251.50.54 | 100% loss |
|
|
| `fv-ml1` mesh 100.64.0.7 | 100% loss |
|
|
| FV gateway 10.251.50.1 | 100% loss |
|
|
| FV gateway mesh 100.64.0.8 | 100% loss |
|
|
| **fv-ml1-bmc 10.251.250.50** | **100% loss** |
|
|
| `fv.phasefinal.com` (172.83.89.66) | no answer |
|
|
|
|
**My own path is healthy** — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all
|
|
answer with 0% loss, so this is FV-local, not an isolation of the observer.
|
|
|
|
## Timeline
|
|
|
|
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
|
|
the two-card power risk. Production seat already live on GPU 2.
|
|
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
|
|
06:54:39 off_A arm healthy after 211 s
|
|
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
|
|
06:56:04 off_A rep 2 started <-- LAST LOG LINE
|
|
06:58:40 every FV address unreachable, including the BMC
|
|
|
|
So the site went dark inside a ~2.5 minute window, roughly five minutes into
|
|
sustained bench load on GPU 3 with GPU 2 also resident/serving.
|
|
|
|
## ⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
|
|
|
|
Operator's read, and it fits the evidence **better** than the breaker-trip theory
|
|
below, for a reason worth writing down:
|
|
|
|
**A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective
|
|
device to give.** A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit
|
|
carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM,
|
|
plus the firewall, plausibly clears the UPS rating while staying comfortably under the
|
|
breaker. That explains the thing a breaker trip explains poorly: **why it let go at
|
|
TWO cards loaded and not four.** It also explains the OPNsense box dying -- same UPS.
|
|
|
|
⚠ **"Died" is the operationally important word.** A tripped UPS resets; an overloaded
|
|
one can kill its output stage or battery pack permanently. If it is dead rather than
|
|
tripped, **nothing on site can be reset back to life** and the trip is wasted without
|
|
the means to bypass it.
|
|
|
|
### Bring / check list for the site visit
|
|
|
|
- **The means to BYPASS the UPS entirely** -- chassis and firewall straight to the
|
|
PDU or wall. Do this first, get the site back, diagnose the UPS after.
|
|
- **Read off the UPS before moving it:** make/model, VA and W rating, fault LEDs, and
|
|
any LCD or event log entry (many units record "overload" explicitly).
|
|
- **Note which outlets were battery-backed vs surge-only.** Mixed-bank units are
|
|
common, and a GPU chassis on the battery-backed bank overloads soonest.
|
|
- ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST, before any
|
|
bring-up.** It sampled all four cards every 10 s right up to the cut and is the
|
|
ONLY measurement of what the load actually drew. It lives on `/tank`, not in a
|
|
container, so it survived. Without it, the replacement UPS gets sized by guesswork.
|
|
- ⚠ **Size the replacement for FOUR cards under load, not two.** Anything sized to
|
|
today's failure just relocates the trip point to the next person who loads all four.
|
|
|
|
⚠ **No measured load figure exists yet.** Idle draw was measured (GPU0 3.80 W, GPU1
|
|
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
|
|
and is recoverable only from `power.log`. Do not let anyone put a wattage in a
|
|
purchasing decision until that file has been read.
|
|
|
|
## ⭐ OPERATOR RULING 2026-09-13: undersized UPS. NAT hypothesis DEMOTED.
|
|
|
|
> "highly doubt the nat hypothesis, it went effective (connectivity confirmed for
|
|
> previously dark path) -- then 20 minutes later, during load, site went dark. that's
|
|
> pretty unlikely to be the cause. for sure I think the ups was undersized."
|
|
|
|
Accepted, and the reasoning is better than the hypothesis-space argument below: **a
|
|
config change that went effective, was verified bidirectional, and then ran correctly
|
|
for twenty minutes does not spontaneously fail when an unrelated physical variable --
|
|
someone else's GPU load -- is introduced.** The load correlation is tight; the NAT
|
|
correlation is merely adjacent in time. Undersized UPS is the only candidate that
|
|
explains the *trigger*.
|
|
|
|
The NAT material below is retained as record, not as a live competing theory, and the
|
|
`power.log` / `uptime` check is retained as **free confirmation** rather than as a
|
|
decision point.
|
|
|
|
## Measurement plan for the site visit (operator bringing a PDU + ammeter)
|
|
|
|
⚠ **`power.log` is GPU-ONLY.** It samples `nvidia-smi` per-card draw and does **not**
|
|
include the host: CPU, 566 GB of RAM, drives, fans, or PSU conversion losses. The
|
|
number that matters against a UPS rating is the whole chassis **at the plug**. The
|
|
ammeter is therefore the primary instrument and `power.log` is a cross-check on the
|
|
GPU share.
|
|
|
|
Capture four states -- this is the first real sizing data that has ever existed for
|
|
this box:
|
|
|
|
| state | why it matters |
|
|
|---|---|
|
|
| all seats down, idle | the floor (GPU idle measured 3.80 / 3.88 / 14.27 / 6.75 W) |
|
|
| one card loaded | the condition that ran fine for a day |
|
|
| **two cards loaded** | the condition that took the site down |
|
|
| **four cards loaded** | the only number that can size a replacement honestly |
|
|
|
|
⚠⚠ **CAPTURE PEAK, NOT AVERAGE.** GPU power has fast transients and UPS overload
|
|
protection responds to short-term overload, so a 1-second-average reading can
|
|
under-read peaks badly. Use max-hold/peak capture if the meter has it. A figure like
|
|
"1100 W average" that hides 1600 W spikes will mis-size the replacement the same way
|
|
the current unit got mis-sized. **If the meter is average-only, record the number as a
|
|
FLOOR, not as the draw.**
|
|
|
|
⭐ **The four-card figure goes into `servers/fv-ml1/README.md` permanently.** The
|
|
cutover notes flagged that the FV circuit was "likely specced against half the real
|
|
draw" while every record still said two GPUs; this closes that with a measurement
|
|
instead of an assumption.
|
|
|
|
## The NAT change, retained as record (DEMOTED — see the ruling above)
|
|
|
|
**Another session applied a Tailscale SNAT rule to the FV gateway at ~06:22Z — 34
|
|
minutes before the site went dark.** See `docs/runbooks/fv-to-ana-nat.md` (uncommitted
|
|
as of this writing; not my work, left alone). So "we overloaded the power" is no longer
|
|
the only live hypothesis, and the UPS should not be replaced on the strength of a theory
|
|
until the one below is run.
|
|
|
|
**On the evidence, that change is the WRONG SHAPE to have caused this**, and I want that
|
|
on the record so nobody wastes the visit chasing it:
|
|
|
|
- It is one **outbound** SNAT rule, source-scoped to `10.251.50.54/32`, destination-
|
|
scoped to `10.250.0.0/16`. An outbound NAT rule cannot stop the gateway, the BMC or
|
|
the public WAN address from answering **inbound**.
|
|
- The runbook states no routes, filter rules, WAN settings or subnet advertisements were
|
|
touched, and that `pfctl -sr` came back byte-identical.
|
|
- It was verified working in **both** directions afterwards: FV→hub HTTP 200, FV→ANA
|
|
TCP 5432, FV internet HTTPS 200, **ANA→FV SSH reachable**, gateway management intact,
|
|
Beszel **18/18 up**.
|
|
|
|
⚠ Note their BMC observation used **`10.251.50.50`**, which is not the BMC — the BMC is
|
|
**`10.251.250.50`**, a different subnet. They correctly declined to claim BMC health, but
|
|
the datapoint is *void*, not negative. Do not reason from it either way.
|
|
|
|
### ⭐⭐ THE DISCRIMINATOR — run this before forming any conclusion
|
|
|
|
`power.log` is written **locally on `/tank`, every 10 seconds, by a shell loop on the
|
|
box.** It does not depend on the network. So:
|
|
|
|
| `power.log` last entry | what it means |
|
|
|---|---|
|
|
| **past 06:56Z** | the box **never lost power**. This is a routing/gateway fault, and the UPS is innocent. |
|
|
| **stops at ~06:56Z** | the box lost power. UPS/circuit confirmed. |
|
|
|
|
Cross-check with `uptime` and `journalctl --list-boots` the moment there is a console:
|
|
**continuous uptime across 06:56Z kills the UPS theory outright.**
|
|
|
|
⚠ **So the FIRST action on site is to read, not to fix.** `uptime`,
|
|
`journalctl --list-boots`, then `tail power.log`. Establish whether the machine ever
|
|
went down before anyone buys hardware or flips anything — the two hypotheses lead to
|
|
completely different remediations and only one of them needs a new UPS.
|
|
|
|
## Three further candidate causes, and what distinguishes them
|
|
|
|
Cannot be distinguished remotely, because every FV path — including the BMC —
|
|
traverses the OPNsense gateway, and the gateway is also dark.
|
|
|
|
1. **Circuit tripped under two-card load.** Fits the timing and is the predicted
|
|
failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this
|
|
same chassis, and the FV circuit was specced while every record still said the box
|
|
had **two** GPUs rather than four. ⭐ Distinguishing evidence: **the breaker is
|
|
visibly tripped**, and the OPNsense box is dark too (it draws ~20-30 W and would
|
|
survive anything short of a circuit/utility loss).
|
|
2. **OPNsense gateway crashed or rebooted**, taking all FV routing with it while the
|
|
GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and
|
|
the OPNsense box is the only dead thing.
|
|
3. **Upstream utility or colo-side power/network loss**, unrelated to us.
|
|
Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has
|
|
power; neighbouring equipment is also dark.
|
|
|
|
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong
|
|
circumstantial evidence, not a measurement — and the instrument that would have
|
|
measured it (the power log on fv-ml1) died with the box.
|
|
|
|
## Blast radius
|
|
|
|
**19 of 30 LiteLLM aliases are dark** — fv-ml1 backs most of the fleet's inference:
|
|
|
|
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
|
|
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
|
|
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
|
|
summarizer-large
|
|
|
|
⚠⚠ **THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied
|
|
there was.** Probed 2026-09-13: every free local model on the gateway lives on
|
|
fv-ml1. irv-ml1 runs **no LLM chat seat at all** -- it carries TTS (tts-gateway,
|
|
breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and
|
|
waterland-studio, and its two Ampere cards are partly occupied by them. The only
|
|
non-fv chat backends on the gateway are **paid**: api.z.ai (9 aliases) and
|
|
Moonshot/Kimi (2).
|
|
|
|
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
|
|
vendor credits. ⚠ **If credits are spent, it must be under a NEW alias name that
|
|
callers opt into** -- never by silently repointing `summarizer`/`gen`/`classifier` at
|
|
GLM. Silent model substitution behind a familiar name is a standing prohibition here
|
|
and has already been violated twice; an outage is not an exemption.
|
|
|
|
The gateway itself on ana-docker is healthy — it is the backends that are gone.
|
|
|
|
## ⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
|
|
|
|
Every vLLM seat on fv-ml1 carries `restart: unless-stopped`. On boot **all of them
|
|
start loading simultaneously** — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
|
|
embed, rerank, reward, coder, scriberr — which is the single largest power transient
|
|
the box can produce, fed straight into a circuit that may have just tripped. That is a
|
|
re-trip, and a re-trip during model load can leave a half-written page cache and a
|
|
much longer recovery.
|
|
|
|
**Preferred sequence:**
|
|
|
|
1. Power the chassis on with **Docker masked**, so nothing auto-starts:
|
|
at the BMC/console, boot to the OS and before the network comes up run
|
|
`systemctl mask docker containerd` — or if the box is already up and loading,
|
|
`systemctl stop docker` immediately.
|
|
2. Confirm `nvidia-smi` sees all four cards and `zpool status tank` is ONLINE.
|
|
3. Unmask, then bring seats up **one at a time**, waiting for each to report healthy:
|
|
`gen` first (19 aliases depend on it), then embed/rerank/reward/coder, then
|
|
mog-sec, then the rest. `flash-next` LAST — it is the newest and least depended-on.
|
|
4. **Do not restart the MTP campaign.** It is the prime suspect.
|
|
5. Watch power while seats come up: `nvidia-smi --query-gpu=index,power.draw --format=csv`.
|
|
|
|
### ⭐ The staged bring-up also finishes the homepage-label fix, for free
|
|
|
|
The 2026-09-13 renumber (commit `3132a16`) repaired 25 compose files on this box but
|
|
the 10 RUNNING containers were never recreated, so their labels still carried the dead
|
|
10.250.50.54. Those containers are gone with the power loss.
|
|
|
|
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. `restart: unless-stopped` restarts the existing
|
|
container with its existing labels; labels only attach at container CREATION. But the
|
|
staged `docker compose up -d <svc>` sequence above **is** a recreate, and the compose
|
|
files on disk are already corrected — so bringing seats up that way applies the new
|
|
labels as a side effect and the dashboard comes back correct. Bring them up with
|
|
`compose up -d`, not by letting Docker restore the old containers.
|
|
|
|
Afterwards, confirm with:
|
|
|
|
curl -s http://10.0.50.45:5100/api/services | \
|
|
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
|
|
|
|
Expect `[]`. Before the outage that query returned 16 entries.
|
|
|
|
**Do not reboot the OPNsense firewall** (standing operator directive; its reboot API
|
|
403s anyway).
|
|
|
|
## ⚠ THE CIRCUIT CASE — what split power does and does not buy (operator, 2026-09-13)
|
|
|
|
> "unless of course the thing trips the circuit anyway."
|
|
|
|
**It still helps, but only halfway, and the halfway matters.**
|
|
|
|
- ✅ **A breaker trip is exactly what the split survives.** Firewall + BMC on the UPS is
|
|
~25-40 W of load on a 1500 VA unit — hours of battery, not minutes. On a trip the UPS
|
|
stops being a load-bearing supply and goes back to being what it is for.
|
|
- ❌ **A live firewall is useless if the path OUT of the site is dead.** Our UPS covers
|
|
our gear; it does not cover the **colo's handoff** — their switch, ONT or demarc. If
|
|
that sits on the circuit we just tripped, the result is a firewall running happily on
|
|
battery with nothing upstream to talk to, and the drive happens anyway.
|
|
⭐ **ASK THE FACILITY: is the network handoff on our circuit or theirs, and is theirs
|
|
on facility UPS?** This is the question that decides whether split power actually
|
|
delivers remote diagnosis or merely feels like it does.
|
|
|
|
### ⚠⚠ And the case where none of the above matters
|
|
|
|
**If four cards plus host exceeds the circuit, no UPS arrangement helps** — the box does
|
|
not fit its feed. Removing an undersized UPS does not remove the constraint, it promotes
|
|
the next one:
|
|
|
|
UPS ~900-1200 W (the one that just gave way)
|
|
circuit ~1800 W @ 15 A / ~2400 W @ 20 A
|
|
|
|
Which side of those the four-card figure lands on decides everything, which is why that
|
|
single ammeter reading is the load-bearing measurement of the visit.
|
|
|
|
### ⭐ The lever that may avoid an electrician: per-card power limits
|
|
|
|
`nvidia-smi -pl <watts>` caps TGP per card. The box can be made to fit whatever the feed
|
|
turns out to be, at a **throughput** cost rather than a **rewiring** cost — four capped
|
|
cards on a 15 A circuit is a dial we control today, where a 20 A drop is a ticket and a
|
|
site visit.
|
|
|
|
- Read `nvidia-smi -q -d POWER` first for the enforced min/max range per card; do not
|
|
assume how much room the dial has.
|
|
- ⚠ **If capping is the answer it MUST be persisted** (systemd unit, or an `if-up`
|
|
equivalent). A limit that evaporates on reboot is worse than no limit, because it will
|
|
hold right up until the next power event and then silently stop holding.
|
|
|
|
### Three questions for the site visit
|
|
|
|
1. What is the **breaker rating** on that circuit?
|
|
2. Is the circuit **dedicated** to us, or shared with other racks/tenants?
|
|
3. Is the **network handoff** on our circuit or the facility's, and is the facility's on
|
|
their UPS?
|
|
|
|
## ⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
|
|
|
|
**This is a recommendation awaiting the operator's call, not settled intent.** Written
|
|
down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring
|
|
the arrangement that just failed by default.
|
|
|
|
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
|
|
PDU / wall <- GPU chassis (no UPS in series)
|
|
|
|
Two reasons:
|
|
|
|
1. **It fixes the OOB gap this outage exposed.** The BMC's only route to the fleet is
|
|
through the firewall, so a power event at the GPU box takes out the management plane
|
|
with it -- which is precisely why this incident needs a drive rather than a console
|
|
session. Separate the two and a repeat leaves a live firewall, a live BMC, and
|
|
remote eyes on a dark chassis.
|
|
2. **A 1500 VA unit was never going to hold this box.** It has four cards, not the two
|
|
every record claimed until 2026-09-12.
|
|
|
|
⚠ **Do NOT use a UPS's surge-only outlets to get around its rating.** Both outlet banks
|
|
sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA
|
|
unit that is a single NEMA 5-15P rated **12 A at maximum load**, total across all
|
|
outlets. The surge bank bypasses the inverter, not the current rating. Overloading the
|
|
inverter trips or kills the unit; overloading the cord is a thermal problem in an
|
|
unattended rack. Bypass the UPS entirely instead.
|
|
|
|
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
|
|
probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
|
|
|
## Afterwards
|
|
|
|
- ⭐ **The OOB design gap this exposes.** The cutover chose OPNsense-as-subnet-router
|
|
specifically so the BMC stays reachable when the GPU box is down. That works for
|
|
box-down/gateway-up. It does **nothing** for a site-wide power or gateway loss —
|
|
exactly what happened — because the BMC's only path to the fleet is through that
|
|
gateway. A genuine OOB path at FV needs something the FV circuit cannot take down:
|
|
an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
|
|
- Get the actual circuit rating and the box's real peak draw, now that it is known to
|
|
have four cards and not two. Until then, treat concurrent multi-card load at FV as
|
|
unproven rather than safe.
|
|
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
|
|
up to the cut. Recover it after boot — it is the only measurement of what the load
|
|
actually drew, and it survives on `/tank`, not in the container.
|
|
|
|
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
|
|
|
|
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
|
|
> power limited to 200w"
|
|
>
|
|
> **CLARIFIED BY OPERATOR 2026-09-13 — the two boxes are different hardware:**
|
|
>
|
|
> | box | cards | TGP each | VRAM total | status |
|
|
> |---|---|---|---|---|
|
|
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
|
|
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
|
|
>
|
|
> **CAPS (operator, 2026-09-13):** **fv-ml1 → 250 W** per card (83% of TGP, ~5%
|
|
> throughput). **ana-ml3 → 200 W** per card (67% of TGP, ~10-15%). See the plug-side
|
|
> arithmetic below — 250 W is marginal on a 15 A circuit once PSU efficiency is counted,
|
|
> and must be verified with the ammeter rather than assumed.
|
|
|
|
The generalised lesson from this outage: **decide the power envelope first and size the
|
|
cards into it**, rather than installing cards and discovering the constraint by tripping
|
|
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
|
|
|
|
Three things to settle before that is a plan:
|
|
|
|
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
|
|
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
|
|
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
|
|
is scripted, refused quietly. **First command on the new hardware:**
|
|
|
|
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
|
|
|
|
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
|
|
feed, not from the cap.
|
|
|
|
2. ✅ **RESOLVED — both card types are 300 W.** So 200 W is a cap to **67% of TGP**, the
|
|
favourable part of the concave curve, not the severe 33% cap a 600 W part would have
|
|
implied. The enforceable-floor concern largely goes away too: 200 W was borderline
|
|
against a 600 W card's minimum and is very unlikely to sit below a 300 W card's. Worth
|
|
the one command; expect it to take.
|
|
|
|
**ana-ml3: 2 x 200 W = 400 W of card.** Modest, and pointed — ana-ml3 lands in the
|
|
**Anaheim** rack whose breaker tripped on 2026-08-26 and 2026-09-11, one of those
|
|
caused by this very chassis before it relocated. The cap there is remediation of a
|
|
circuit with a track record, not precaution.
|
|
|
|
### ⭐⭐ The outage arithmetic, now that the TGP is known
|
|
|
|
2 x Blackwell Max-Q @ 300 W ~ 600 W of card under load
|
|
host (board, 566 GB RAM, drives,
|
|
fans, PSU conversion loss) ~ 200-350 W <-- UNMEASURED, the gap
|
|
-----------
|
|
~ 800-950 W
|
|
|
|
Eaton 1500 VA real watt rating ~ 900-1200 W depending on model
|
|
|
|
**At or just over the line** — and this is what a vague "undersized" could not explain:
|
|
why it ran a full day on one card (~500-650 W, comfortably inside) and died minutes into
|
|
the second (~800-950 W, at or past the rating). The host term is the only one being
|
|
guessed at, and idle-at-the-plug with all seats down measures it directly.
|
|
|
|
### ⚠⚠ FOUR cards is a BREAKER problem, not a UPS problem
|
|
|
|
4 x 300 W card + ~300 W host ~ 1500 W
|
|
15 A circuit, 80% continuous = 1440 W
|
|
|
|
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all**, so
|
|
capping belongs at fv-ml1 too. Note today's incident only ever had TWO of the four cards
|
|
working; nobody has loaded all four.
|
|
|
|
### ⚠⚠⚠ AND `nvidia-smi -pl` CAPS BOARD POWER, NOT WALL POWER
|
|
|
|
This term is easy to drop and it is the one that decides whether 250 W clears a 15 A
|
|
feed:
|
|
|
|
4 x 250 W board = 1000 W
|
|
host components (board, 566 GB RAM,
|
|
drives, fans) ~ 180-300 W
|
|
-----------
|
|
component total ~ 1180-1300 W
|
|
/ PSU efficiency (~0.90) -> AT THE PLUG ~ 1310-1445 W
|
|
|
|
15 A circuit, NEC 80% continuous derating = 1440 W
|
|
|
|
**250 W lands ON the limit, not under it.** An inference box serving all day is a
|
|
continuous load, so 1440 W is the design figure, not 1800.
|
|
|
|
At **200 W** the same arithmetic gives **~1090-1220 W at the plug** — comfortable, 200+ W
|
|
of margin.
|
|
|
|
⚠ **So 250 W is PROBABLY fine and POSSIBLY not, and the deciding term is the host draw,
|
|
which is still an estimate.** Procedure: set 250 W, then **verify at the plug under
|
|
four-card load** before calling it done; fall back to 200 W if the reading comes in near
|
|
1440 W. A cap is a claim; the ammeter is the verification.
|
|
|
|
⚠ **Power limits cap SUSTAINED draw, not transients.** The enforcement window is short
|
|
but not instantaneous, so four cards at 250 W can momentarily exceed 1000 W of board.
|
|
A breaker tolerates that (thermal-magnetic curves are forgiving of brief overload); a
|
|
UPS's overload protection is not. Which means **250 W implicitly commits the chassis to
|
|
the PDU rather than behind the 1500 VA unit** — even capped, 4 x 250 W + host exceeds
|
|
that UPS's real rating.
|
|
|
|
⚠ **200 W on ana-ml3's TWO cards is deliberately conservative** (2 x 200 = 400 W is
|
|
trivial on any circuit). Relaxable later if Anaheim's measured headroom beats its trip
|
|
history; not a permanent figure.
|
|
|
|
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
|
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
|
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
|
|
33% of TGP is deeper into the steep region — the cost is real and should be measured
|
|
on the first card rather than predicted, and it will hurt a prefill-heavy or training
|
|
workload considerably more than a serving seat.
|
|
|
|
### ⚠ Two placement consequences of Ada, independent of power
|
|
|
|
- **sm_89 has native FP8 but NOT NVFP4** (Blackwell-only, sm_100/sm_120). Most of our
|
|
in-house quants are NVFP4, so **they will not run accelerated on that colo's cards.**
|
|
Its seats want FP8 W8A8 builds, or the NVFP4 checkpoints stay on fv-ml1. Same class of
|
|
constraint as the Ampere finding for irv-ml1, one generation up.
|
|
- ⭐ **It unparks the triton-backend item.** That is a hard no on Ampere — crashes every
|
|
render on the A6000, `fp8e4nv` unsupported on sm_86 — and was explicitly deferred TO
|
|
Ada. sm_89 has the FP8 support it needs, so it becomes testable on this hardware.
|
|
- **VRAM:** 4 x 48 GB = 192 GB, against fv-ml1's 4 x 96 = 391 GB. Big-model placement
|
|
stays at FV. The Flash-Next seat needs 74 GiB resident on ONE card and would not fit a
|
|
48 GB Ada card even with the n-gram table offloaded — the offload moves the *table*,
|
|
not the experts.
|
|
|
|
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
|
|
stops holding — the worst possible failure shape, because the thing that reboots the box
|
|
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
|
|
mode, ordered before Docker starts.
|
|
|
|
## Can the whole BANK be capped at 1000 W instead of per-card? (operator question)
|
|
|
|
> "is it possible to cap the ENTIRE bank to 1000w? meaning that each card can go to max
|
|
> until they're all loaded?"
|
|
|
|
**The concept is first-class in DCGM, the dynamic behaviour is not free, and the static
|
|
cap already equals the bank budget.**
|
|
|
|
### What exists
|
|
|
|
`dcgmConfigPowerLimitType_enum` (DCGM API) carries exactly this distinction:
|
|
|
|
DCGM_CONFIG_POWER_CAP_INDIVIDUAL "the power cap to be applied for each member of the group"
|
|
DCGM_CONFIG_POWER_BUDGET_GROUP "the power budget for the entire group"
|
|
|
|
⚠ **The documentation does not state how a group budget is distributed.** Deduction, not
|
|
a quote: the only enforcement primitive underneath is NVML's per-GPU
|
|
`nvmlDeviceSetPowerManagementLimit` — **there is no bank-level register** — so any group
|
|
budget ultimately resolves to N per-GPU writes. Static even division needs one write
|
|
each; "each card free until they are all loaded" needs **continuous re-writing**, i.e. a
|
|
control loop rather than a hardware feature.
|
|
|
|
⭐ **Worth a 10-minute test when the box is back:** set `DCGM_CONFIG_POWER_BUDGET_GROUP`
|
|
to 1000 W and watch whether per-GPU limits move as load shifts. If NVIDIA already runs
|
|
the loop, take it for free.
|
|
|
|
### ⚠⚠ If a loop is written, its failure direction matters more than its logic
|
|
|
|
Power readings lag and `-pl` application takes tens of ms, so a reactive daemon
|
|
overshoots during a load RAMP — and the ramp is exactly the dangerous moment, because it
|
|
is the all-four-cards-loading-at-once case (the same shape as this box's ten
|
|
`restart: unless-stopped` containers starting together).
|
|
|
|
**Therefore: safe-by-default, opportunistic upward.** Boot every card at budget/N and
|
|
only ever RAISE a card's cap after observing idle neighbours. Never boot high and react
|
|
down. Failure mode then becomes "slower than it could have been" instead of "tripped the
|
|
breaker." Inverted, it is a thing that works for weeks and then fails on precisely the
|
|
event it existed to prevent.
|
|
|
|
### Why not yet
|
|
|
|
**4 x 250 W = 1000 W — the static cap IS the bank budget**, and it is the conservative
|
|
floor of the dynamic scheme rather than an alternative to it. The daemon's entire
|
|
contribution is the one-card-busy case: ~300 W instead of ~250 W on a single card, ~17%
|
|
more board power, which on a concave perf/watt curve is perhaps ~5% throughput. That is
|
|
the whole prize, against a control loop whose failure mode points at a breaker.
|
|
|
|
And whether that case is even common depends on workload mix. A **serving** fleet spreads
|
|
across cards by construction — one seat per card, gateway traffic split across aliases —
|
|
so single-card-busy is rare. A **training window** is the opposite: one card hammered,
|
|
three idle, which is where the dynamic scheme actually pays.
|
|
|
|
**Recommendation: static 250 W now, measure at the plug, build the loop only if the
|
|
measurements show single-card-busy is the common case.** It is a pure optimization on
|
|
top; adding it later re-architects nothing and would be built against real numbers.
|