I claimed in 100670e that DCGM's config enforcement was plausibly gated to datacenter
SKUs and told the operator not to plan around it. That was a guess presented as a caveat
and it is wrong. Verified against NVIDIA's own documentation at the operator's request.
Supported platforms explicitly cover 'All NVIDIA Maxwell and newer non-datacenter (e.g.
NVIDIA GeForce or NVIDIA Quadro) GPUs', and the feature-overview table marks
Configuration Management as supported for Tesla, Titan, Quadro and GeForce alike --
where Configuration Management explicitly includes 'Power Limit: Set the maximum allowed
power consumption'. What is actually gated on non-datacenter cards is diagnostics: Level
1 only, against All Levels on Tesla. Configuration was never the restricted part.
One soft edge retained rather than papered over: the table says 'Quadro', the former name
for the professional line, and RTX 6000 Ada / RTX PRO 6000 are its successors, so placing
them in that column is inference rather than quotation. One command on the box settles it.
What does not change is the distribution question. DCGM_CONFIG_POWER_BUDGET_GROUP is
available to us, but the docs still never state how a group budget is divided, and the
NVML argument is untouched -- there is no bank-level register, so it resolves to per-GPU
writes either way and the likely finding is static even division, which is exactly
4 x 250 W. The experiment is therefore promoted from curiosity back to a real test.
573 lines
31 KiB
Markdown
573 lines
31 KiB
Markdown
# FV site dark — 2026-09-13 ~06:56Z
|
|
|
|
**Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley.**
|
|
Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one
|
|
needs replacing, so there is no recovery to wait for and no point polling the site.
|
|
Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is
|
|
down so recovery does not have to be reconstructed from memory.
|
|
|
|
## What is down
|
|
|
|
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
|
|
|
|
| target | result |
|
|
|---|---|
|
|
| `fv-ml1` 10.251.50.54 | 100% loss |
|
|
| `fv-ml1` mesh 100.64.0.7 | 100% loss |
|
|
| FV gateway 10.251.50.1 | 100% loss |
|
|
| FV gateway mesh 100.64.0.8 | 100% loss |
|
|
| **fv-ml1-bmc 10.251.250.50** | **100% loss** |
|
|
| `fv.phasefinal.com` (172.83.89.66) | no answer |
|
|
|
|
**My own path is healthy** — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all
|
|
answer with 0% loss, so this is FV-local, not an isolation of the observer.
|
|
|
|
## Timeline
|
|
|
|
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
|
|
the two-card power risk. Production seat already live on GPU 2.
|
|
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
|
|
06:54:39 off_A arm healthy after 211 s
|
|
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
|
|
06:56:04 off_A rep 2 started <-- LAST LOG LINE
|
|
06:58:40 every FV address unreachable, including the BMC
|
|
|
|
So the site went dark inside a ~2.5 minute window, roughly five minutes into
|
|
sustained bench load on GPU 3 with GPU 2 also resident/serving.
|
|
|
|
## ⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
|
|
|
|
Operator's read, and it fits the evidence **better** than the breaker-trip theory
|
|
below, for a reason worth writing down:
|
|
|
|
**A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective
|
|
device to give.** A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit
|
|
carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM,
|
|
plus the firewall, plausibly clears the UPS rating while staying comfortably under the
|
|
breaker. That explains the thing a breaker trip explains poorly: **why it let go at
|
|
TWO cards loaded and not four.** It also explains the OPNsense box dying -- same UPS.
|
|
|
|
⚠ **"Died" is the operationally important word.** A tripped UPS resets; an overloaded
|
|
one can kill its output stage or battery pack permanently. If it is dead rather than
|
|
tripped, **nothing on site can be reset back to life** and the trip is wasted without
|
|
the means to bypass it.
|
|
|
|
### Bring / check list for the site visit
|
|
|
|
- **The means to BYPASS the UPS entirely** -- chassis and firewall straight to the
|
|
PDU or wall. Do this first, get the site back, diagnose the UPS after.
|
|
- **Read off the UPS before moving it:** make/model, VA and W rating, fault LEDs, and
|
|
any LCD or event log entry (many units record "overload" explicitly).
|
|
- **Note which outlets were battery-backed vs surge-only.** Mixed-bank units are
|
|
common, and a GPU chassis on the battery-backed bank overloads soonest.
|
|
- ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST, before any
|
|
bring-up.** It sampled all four cards every 10 s right up to the cut and is the
|
|
ONLY measurement of what the load actually drew. It lives on `/tank`, not in a
|
|
container, so it survived. Without it, the replacement UPS gets sized by guesswork.
|
|
- ⚠ **Size the replacement for FOUR cards under load, not two.** Anything sized to
|
|
today's failure just relocates the trip point to the next person who loads all four.
|
|
|
|
⚠ **No measured load figure exists yet.** Idle draw was measured (GPU0 3.80 W, GPU1
|
|
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
|
|
and is recoverable only from `power.log`. Do not let anyone put a wattage in a
|
|
purchasing decision until that file has been read.
|
|
|
|
## ⭐ OPERATOR RULING 2026-09-13: undersized UPS. NAT hypothesis DEMOTED.
|
|
|
|
> "highly doubt the nat hypothesis, it went effective (connectivity confirmed for
|
|
> previously dark path) -- then 20 minutes later, during load, site went dark. that's
|
|
> pretty unlikely to be the cause. for sure I think the ups was undersized."
|
|
|
|
Accepted, and the reasoning is better than the hypothesis-space argument below: **a
|
|
config change that went effective, was verified bidirectional, and then ran correctly
|
|
for twenty minutes does not spontaneously fail when an unrelated physical variable --
|
|
someone else's GPU load -- is introduced.** The load correlation is tight; the NAT
|
|
correlation is merely adjacent in time. Undersized UPS is the only candidate that
|
|
explains the *trigger*.
|
|
|
|
The NAT material below is retained as record, not as a live competing theory, and the
|
|
`power.log` / `uptime` check is retained as **free confirmation** rather than as a
|
|
decision point.
|
|
|
|
## Measurement plan for the site visit (operator bringing a PDU + ammeter)
|
|
|
|
⚠ **`power.log` is GPU-ONLY.** It samples `nvidia-smi` per-card draw and does **not**
|
|
include the host: CPU, 566 GB of RAM, drives, fans, or PSU conversion losses. The
|
|
number that matters against a UPS rating is the whole chassis **at the plug**. The
|
|
ammeter is therefore the primary instrument and `power.log` is a cross-check on the
|
|
GPU share.
|
|
|
|
Capture four states -- this is the first real sizing data that has ever existed for
|
|
this box:
|
|
|
|
| state | why it matters |
|
|
|---|---|
|
|
| all seats down, idle | the floor (GPU idle measured 3.80 / 3.88 / 14.27 / 6.75 W) |
|
|
| one card loaded | the condition that ran fine for a day |
|
|
| **two cards loaded** | the condition that took the site down |
|
|
| **four cards loaded** | the only number that can size a replacement honestly |
|
|
|
|
⚠⚠ **CAPTURE PEAK, NOT AVERAGE.** GPU power has fast transients and UPS overload
|
|
protection responds to short-term overload, so a 1-second-average reading can
|
|
under-read peaks badly. Use max-hold/peak capture if the meter has it. A figure like
|
|
"1100 W average" that hides 1600 W spikes will mis-size the replacement the same way
|
|
the current unit got mis-sized. **If the meter is average-only, record the number as a
|
|
FLOOR, not as the draw.**
|
|
|
|
⭐ **The four-card figure goes into `servers/fv-ml1/README.md` permanently.** The
|
|
cutover notes flagged that the FV circuit was "likely specced against half the real
|
|
draw" while every record still said two GPUs; this closes that with a measurement
|
|
instead of an assumption.
|
|
|
|
## The NAT change, retained as record (DEMOTED — see the ruling above)
|
|
|
|
**Another session applied a Tailscale SNAT rule to the FV gateway at ~06:22Z — 34
|
|
minutes before the site went dark.** See `docs/runbooks/fv-to-ana-nat.md` (uncommitted
|
|
as of this writing; not my work, left alone). So "we overloaded the power" is no longer
|
|
the only live hypothesis, and the UPS should not be replaced on the strength of a theory
|
|
until the one below is run.
|
|
|
|
**On the evidence, that change is the WRONG SHAPE to have caused this**, and I want that
|
|
on the record so nobody wastes the visit chasing it:
|
|
|
|
- It is one **outbound** SNAT rule, source-scoped to `10.251.50.54/32`, destination-
|
|
scoped to `10.250.0.0/16`. An outbound NAT rule cannot stop the gateway, the BMC or
|
|
the public WAN address from answering **inbound**.
|
|
- The runbook states no routes, filter rules, WAN settings or subnet advertisements were
|
|
touched, and that `pfctl -sr` came back byte-identical.
|
|
- It was verified working in **both** directions afterwards: FV→hub HTTP 200, FV→ANA
|
|
TCP 5432, FV internet HTTPS 200, **ANA→FV SSH reachable**, gateway management intact,
|
|
Beszel **18/18 up**.
|
|
|
|
⚠ Note their BMC observation used **`10.251.50.50`**, which is not the BMC — the BMC is
|
|
**`10.251.250.50`**, a different subnet. They correctly declined to claim BMC health, but
|
|
the datapoint is *void*, not negative. Do not reason from it either way.
|
|
|
|
### ⭐⭐ THE DISCRIMINATOR — run this before forming any conclusion
|
|
|
|
`power.log` is written **locally on `/tank`, every 10 seconds, by a shell loop on the
|
|
box.** It does not depend on the network. So:
|
|
|
|
| `power.log` last entry | what it means |
|
|
|---|---|
|
|
| **past 06:56Z** | the box **never lost power**. This is a routing/gateway fault, and the UPS is innocent. |
|
|
| **stops at ~06:56Z** | the box lost power. UPS/circuit confirmed. |
|
|
|
|
Cross-check with `uptime` and `journalctl --list-boots` the moment there is a console:
|
|
**continuous uptime across 06:56Z kills the UPS theory outright.**
|
|
|
|
⚠ **So the FIRST action on site is to read, not to fix.** `uptime`,
|
|
`journalctl --list-boots`, then `tail power.log`. Establish whether the machine ever
|
|
went down before anyone buys hardware or flips anything — the two hypotheses lead to
|
|
completely different remediations and only one of them needs a new UPS.
|
|
|
|
## Three further candidate causes, and what distinguishes them
|
|
|
|
Cannot be distinguished remotely, because every FV path — including the BMC —
|
|
traverses the OPNsense gateway, and the gateway is also dark.
|
|
|
|
1. **Circuit tripped under two-card load.** Fits the timing and is the predicted
|
|
failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this
|
|
same chassis, and the FV circuit was specced while every record still said the box
|
|
had **two** GPUs rather than four. ⭐ Distinguishing evidence: **the breaker is
|
|
visibly tripped**, and the OPNsense box is dark too (it draws ~20-30 W and would
|
|
survive anything short of a circuit/utility loss).
|
|
2. **OPNsense gateway crashed or rebooted**, taking all FV routing with it while the
|
|
GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and
|
|
the OPNsense box is the only dead thing.
|
|
3. **Upstream utility or colo-side power/network loss**, unrelated to us.
|
|
Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has
|
|
power; neighbouring equipment is also dark.
|
|
|
|
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong
|
|
circumstantial evidence, not a measurement — and the instrument that would have
|
|
measured it (the power log on fv-ml1) died with the box.
|
|
|
|
## Blast radius
|
|
|
|
**19 of 30 LiteLLM aliases are dark** — fv-ml1 backs most of the fleet's inference:
|
|
|
|
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
|
|
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
|
|
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
|
|
summarizer-large
|
|
|
|
⚠⚠ **THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied
|
|
there was.** Probed 2026-09-13: every free local model on the gateway lives on
|
|
fv-ml1. irv-ml1 runs **no LLM chat seat at all** -- it carries TTS (tts-gateway,
|
|
breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and
|
|
waterland-studio, and its two Ampere cards are partly occupied by them. The only
|
|
non-fv chat backends on the gateway are **paid**: api.z.ai (9 aliases) and
|
|
Moonshot/Kimi (2).
|
|
|
|
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
|
|
vendor credits. ⚠ **If credits are spent, it must be under a NEW alias name that
|
|
callers opt into** -- never by silently repointing `summarizer`/`gen`/`classifier` at
|
|
GLM. Silent model substitution behind a familiar name is a standing prohibition here
|
|
and has already been violated twice; an outage is not an exemption.
|
|
|
|
The gateway itself on ana-docker is healthy — it is the backends that are gone.
|
|
|
|
## ⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
|
|
|
|
Every vLLM seat on fv-ml1 carries `restart: unless-stopped`. On boot **all of them
|
|
start loading simultaneously** — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
|
|
embed, rerank, reward, coder, scriberr — which is the single largest power transient
|
|
the box can produce, fed straight into a circuit that may have just tripped. That is a
|
|
re-trip, and a re-trip during model load can leave a half-written page cache and a
|
|
much longer recovery.
|
|
|
|
**Preferred sequence:**
|
|
|
|
1. Power the chassis on with **Docker masked**, so nothing auto-starts:
|
|
at the BMC/console, boot to the OS and before the network comes up run
|
|
`systemctl mask docker containerd` — or if the box is already up and loading,
|
|
`systemctl stop docker` immediately.
|
|
2. Confirm `nvidia-smi` sees all four cards and `zpool status tank` is ONLINE.
|
|
3. Unmask, then bring seats up **one at a time**, waiting for each to report healthy:
|
|
`gen` first (19 aliases depend on it), then embed/rerank/reward/coder, then
|
|
mog-sec, then the rest. `flash-next` LAST — it is the newest and least depended-on.
|
|
4. **Do not restart the MTP campaign.** It is the prime suspect.
|
|
5. Watch power while seats come up: `nvidia-smi --query-gpu=index,power.draw --format=csv`.
|
|
|
|
### ⭐ The staged bring-up also finishes the homepage-label fix, for free
|
|
|
|
The 2026-09-13 renumber (commit `3132a16`) repaired 25 compose files on this box but
|
|
the 10 RUNNING containers were never recreated, so their labels still carried the dead
|
|
10.250.50.54. Those containers are gone with the power loss.
|
|
|
|
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. `restart: unless-stopped` restarts the existing
|
|
container with its existing labels; labels only attach at container CREATION. But the
|
|
staged `docker compose up -d <svc>` sequence above **is** a recreate, and the compose
|
|
files on disk are already corrected — so bringing seats up that way applies the new
|
|
labels as a side effect and the dashboard comes back correct. Bring them up with
|
|
`compose up -d`, not by letting Docker restore the old containers.
|
|
|
|
Afterwards, confirm with:
|
|
|
|
curl -s http://10.0.50.45:5100/api/services | \
|
|
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
|
|
|
|
Expect `[]`. Before the outage that query returned 16 entries.
|
|
|
|
**Do not reboot the OPNsense firewall** (standing operator directive; its reboot API
|
|
403s anyway).
|
|
|
|
## ⚠ THE CIRCUIT CASE — what split power does and does not buy (operator, 2026-09-13)
|
|
|
|
> "unless of course the thing trips the circuit anyway."
|
|
|
|
**It still helps, but only halfway, and the halfway matters.**
|
|
|
|
- ✅ **A breaker trip is exactly what the split survives.** Firewall + BMC on the UPS is
|
|
~25-40 W of load on a 1500 VA unit — hours of battery, not minutes. On a trip the UPS
|
|
stops being a load-bearing supply and goes back to being what it is for.
|
|
- ❌ **A live firewall is useless if the path OUT of the site is dead.** Our UPS covers
|
|
our gear; it does not cover the **colo's handoff** — their switch, ONT or demarc. If
|
|
that sits on the circuit we just tripped, the result is a firewall running happily on
|
|
battery with nothing upstream to talk to, and the drive happens anyway.
|
|
⭐ **ASK THE FACILITY: is the network handoff on our circuit or theirs, and is theirs
|
|
on facility UPS?** This is the question that decides whether split power actually
|
|
delivers remote diagnosis or merely feels like it does.
|
|
|
|
### ⚠⚠ And the case where none of the above matters
|
|
|
|
**If four cards plus host exceeds the circuit, no UPS arrangement helps** — the box does
|
|
not fit its feed. Removing an undersized UPS does not remove the constraint, it promotes
|
|
the next one:
|
|
|
|
UPS ~900-1200 W (the one that just gave way)
|
|
circuit ~1800 W @ 15 A / ~2400 W @ 20 A
|
|
|
|
Which side of those the four-card figure lands on decides everything, which is why that
|
|
single ammeter reading is the load-bearing measurement of the visit.
|
|
|
|
### ⭐ The lever that may avoid an electrician: per-card power limits
|
|
|
|
`nvidia-smi -pl <watts>` caps TGP per card. The box can be made to fit whatever the feed
|
|
turns out to be, at a **throughput** cost rather than a **rewiring** cost — four capped
|
|
cards on a 15 A circuit is a dial we control today, where a 20 A drop is a ticket and a
|
|
site visit.
|
|
|
|
- Read `nvidia-smi -q -d POWER` first for the enforced min/max range per card; do not
|
|
assume how much room the dial has.
|
|
- ⚠ **If capping is the answer it MUST be persisted** (systemd unit, or an `if-up`
|
|
equivalent). A limit that evaporates on reboot is worse than no limit, because it will
|
|
hold right up until the next power event and then silently stop holding.
|
|
|
|
### Three questions for the site visit
|
|
|
|
1. What is the **breaker rating** on that circuit?
|
|
2. Is the circuit **dedicated** to us, or shared with other racks/tenants?
|
|
3. Is the **network handoff** on our circuit or the facility's, and is the facility's on
|
|
their UPS?
|
|
|
|
## ⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
|
|
|
|
**This is a recommendation awaiting the operator's call, not settled intent.** Written
|
|
down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring
|
|
the arrangement that just failed by default.
|
|
|
|
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
|
|
PDU / wall <- GPU chassis (no UPS in series)
|
|
|
|
Two reasons:
|
|
|
|
1. **It fixes the OOB gap this outage exposed.** The BMC's only route to the fleet is
|
|
through the firewall, so a power event at the GPU box takes out the management plane
|
|
with it -- which is precisely why this incident needs a drive rather than a console
|
|
session. Separate the two and a repeat leaves a live firewall, a live BMC, and
|
|
remote eyes on a dark chassis.
|
|
2. **A 1500 VA unit was never going to hold this box.** It has four cards, not the two
|
|
every record claimed until 2026-09-12.
|
|
|
|
⚠ **Do NOT use a UPS's surge-only outlets to get around its rating.** Both outlet banks
|
|
sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA
|
|
unit that is a single NEMA 5-15P rated **12 A at maximum load**, total across all
|
|
outlets. The surge bank bypasses the inverter, not the current rating. Overloading the
|
|
inverter trips or kills the unit; overloading the cord is a thermal problem in an
|
|
unattended rack. Bypass the UPS entirely instead.
|
|
|
|
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
|
|
probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
|
|
|
## Afterwards
|
|
|
|
- ⭐ **The OOB design gap this exposes.** The cutover chose OPNsense-as-subnet-router
|
|
specifically so the BMC stays reachable when the GPU box is down. That works for
|
|
box-down/gateway-up. It does **nothing** for a site-wide power or gateway loss —
|
|
exactly what happened — because the BMC's only path to the fleet is through that
|
|
gateway. A genuine OOB path at FV needs something the FV circuit cannot take down:
|
|
an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
|
|
- Get the actual circuit rating and the box's real peak draw, now that it is known to
|
|
have four cards and not two. Until then, treat concurrent multi-card load at FV as
|
|
unproven rather than safe.
|
|
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
|
|
up to the cut. Recover it after boot — it is the only measurement of what the load
|
|
actually drew, and it survives on `/tank`, not in the container.
|
|
|
|
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
|
|
|
|
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
|
|
> power limited to 200w"
|
|
>
|
|
> **CLARIFIED BY OPERATOR 2026-09-13 — the two boxes are different hardware:**
|
|
>
|
|
> | box | cards | TGP each | VRAM total | status |
|
|
> |---|---|---|---|---|
|
|
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
|
|
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
|
|
>
|
|
> **CAPS (operator, 2026-09-13):** **fv-ml1 → 250 W** per card (83% of TGP, ~5%
|
|
> throughput). **ana-ml3 → 200 W** per card (67% of TGP, ~10-15%). See the plug-side
|
|
> arithmetic below — 250 W is marginal on a 15 A circuit once PSU efficiency is counted,
|
|
> and must be verified with the ammeter rather than assumed.
|
|
|
|
The generalised lesson from this outage: **decide the power envelope first and size the
|
|
cards into it**, rather than installing cards and discovering the constraint by tripping
|
|
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
|
|
|
|
Three things to settle before that is a plan:
|
|
|
|
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
|
|
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
|
|
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
|
|
is scripted, refused quietly. **First command on the new hardware:**
|
|
|
|
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
|
|
|
|
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
|
|
feed, not from the cap.
|
|
|
|
2. ✅ **RESOLVED — both card types are 300 W.** So 200 W is a cap to **67% of TGP**, the
|
|
favourable part of the concave curve, not the severe 33% cap a 600 W part would have
|
|
implied. The enforceable-floor concern largely goes away too: 200 W was borderline
|
|
against a 600 W card's minimum and is very unlikely to sit below a 300 W card's. Worth
|
|
the one command; expect it to take.
|
|
|
|
**ana-ml3: 2 x 200 W = 400 W of card.** Modest, and pointed — ana-ml3 lands in the
|
|
**Anaheim** rack whose breaker tripped on 2026-08-26 and 2026-09-11, one of those
|
|
caused by this very chassis before it relocated. The cap there is remediation of a
|
|
circuit with a track record, not precaution.
|
|
|
|
### ⭐⭐ The outage arithmetic, now that the TGP is known
|
|
|
|
2 x Blackwell Max-Q @ 300 W ~ 600 W of card under load
|
|
host (board, 566 GB RAM, drives,
|
|
fans, PSU conversion loss) ~ 200-350 W <-- UNMEASURED, the gap
|
|
-----------
|
|
~ 800-950 W
|
|
|
|
Eaton 1500 VA real watt rating ~ 900-1200 W depending on model
|
|
|
|
**At or just over the line** — and this is what a vague "undersized" could not explain:
|
|
why it ran a full day on one card (~500-650 W, comfortably inside) and died minutes into
|
|
the second (~800-950 W, at or past the rating). The host term is the only one being
|
|
guessed at, and idle-at-the-plug with all seats down measures it directly.
|
|
|
|
### ⚠⚠ FOUR cards is a BREAKER problem, not a UPS problem
|
|
|
|
4 x 300 W card + ~300 W host ~ 1500 W
|
|
15 A circuit, 80% continuous = 1440 W
|
|
|
|
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all**, so
|
|
capping belongs at fv-ml1 too. Note today's incident only ever had TWO of the four cards
|
|
working; nobody has loaded all four.
|
|
|
|
### ⚠⚠⚠ AND `nvidia-smi -pl` CAPS BOARD POWER, NOT WALL POWER
|
|
|
|
This term is easy to drop and it is the one that decides whether 250 W clears a 15 A
|
|
feed:
|
|
|
|
4 x 250 W board = 1000 W
|
|
host components (board, 566 GB RAM,
|
|
drives, fans) ~ 180-300 W
|
|
-----------
|
|
component total ~ 1180-1300 W
|
|
/ PSU efficiency (~0.90) -> AT THE PLUG ~ 1310-1445 W
|
|
|
|
15 A circuit, NEC 80% continuous derating = 1440 W
|
|
|
|
**250 W lands ON the limit, not under it.** An inference box serving all day is a
|
|
continuous load, so 1440 W is the design figure, not 1800.
|
|
|
|
At **200 W** the same arithmetic gives **~1090-1220 W at the plug** — comfortable, 200+ W
|
|
of margin.
|
|
|
|
⚠ **So 250 W is PROBABLY fine and POSSIBLY not, and the deciding term is the host draw,
|
|
which is still an estimate.** Procedure: set 250 W, then **verify at the plug under
|
|
four-card load** before calling it done; fall back to 200 W if the reading comes in near
|
|
1440 W. A cap is a claim; the ammeter is the verification.
|
|
|
|
⚠ **Power limits cap SUSTAINED draw, not transients.** The enforcement window is short
|
|
but not instantaneous, so four cards at 250 W can momentarily exceed 1000 W of board.
|
|
A breaker tolerates that (thermal-magnetic curves are forgiving of brief overload); a
|
|
UPS's overload protection is not. Which means **250 W implicitly commits the chassis to
|
|
the PDU rather than behind the 1500 VA unit** — even capped, 4 x 250 W + host exceeds
|
|
that UPS's real rating.
|
|
|
|
⚠ **200 W on ana-ml3's TWO cards is deliberately conservative** (2 x 200 = 400 W is
|
|
trivial on any circuit). Relaxable later if Anaheim's measured headroom beats its trip
|
|
history; not a permanent figure.
|
|
|
|
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
|
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
|
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
|
|
33% of TGP is deeper into the steep region — the cost is real and should be measured
|
|
on the first card rather than predicted, and it will hurt a prefill-heavy or training
|
|
workload considerably more than a serving seat.
|
|
|
|
### ⚠ Two placement consequences of Ada, independent of power
|
|
|
|
- **sm_89 has native FP8 but NOT NVFP4** (Blackwell-only, sm_100/sm_120). Most of our
|
|
in-house quants are NVFP4, so **they will not run accelerated on that colo's cards.**
|
|
Its seats want FP8 W8A8 builds, or the NVFP4 checkpoints stay on fv-ml1. Same class of
|
|
constraint as the Ampere finding for irv-ml1, one generation up.
|
|
- ⭐ **It unparks the triton-backend item.** That is a hard no on Ampere — crashes every
|
|
render on the A6000, `fp8e4nv` unsupported on sm_86 — and was explicitly deferred TO
|
|
Ada. sm_89 has the FP8 support it needs, so it becomes testable on this hardware.
|
|
- **VRAM:** 4 x 48 GB = 192 GB, against fv-ml1's 4 x 96 = 391 GB. Big-model placement
|
|
stays at FV. The Flash-Next seat needs 74 GiB resident on ONE card and would not fit a
|
|
48 GB Ada card even with the n-gram table offloaded — the offload moves the *table*,
|
|
not the experts.
|
|
|
|
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
|
|
stops holding — the worst possible failure shape, because the thing that reboots the box
|
|
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
|
|
mode, ordered before Docker starts.
|
|
|
|
## Can the whole BANK be capped at 1000 W instead of per-card? (operator question)
|
|
|
|
> "is it possible to cap the ENTIRE bank to 1000w? meaning that each card can go to max
|
|
> until they're all loaded?"
|
|
|
|
**The concept is first-class in DCGM, the dynamic behaviour is not free, and the static
|
|
cap already equals the bank budget.**
|
|
|
|
### What exists
|
|
|
|
`dcgmConfigPowerLimitType_enum` (DCGM API) carries exactly this distinction:
|
|
|
|
DCGM_CONFIG_POWER_CAP_INDIVIDUAL "the power cap to be applied for each member of the group"
|
|
DCGM_CONFIG_POWER_BUDGET_GROUP "the power budget for the entire group"
|
|
|
|
⚠ **The documentation does not state how a group budget is distributed.** Deduction, not
|
|
a quote: the only enforcement primitive underneath is NVML's per-GPU
|
|
`nvmlDeviceSetPowerManagementLimit` — **there is no bank-level register** — so any group
|
|
budget ultimately resolves to N per-GPU writes. Static even division needs one write
|
|
each; "each card free until they are all loaded" needs **continuous re-writing**, i.e. a
|
|
control loop rather than a hardware feature.
|
|
|
|
**What DCGM is:** NVIDIA's own **Data Center GPU Manager** — first-party, open source
|
|
(Apache 2.0, `NVIDIA/DCGM`), packaged as `datacenter-gpu-manager`. It layers above NVML:
|
|
|
|
nvidia-smi CLI, thin wrapper over NVML
|
|
NVML low-level C library, PER-GPU primitives (what -pl actually calls)
|
|
DCGM daemon (nv-hostengine) + dcgmi, ABOVE NVML — health, diagnostics,
|
|
config enforcement, policy, group abstractions; dcgm-exporter is its
|
|
Prometheus sidecar
|
|
|
|
Its "group" notion is therefore a management-layer abstraction over per-GPU NVML calls,
|
|
which is why the bank budget still resolves to N per-GPU writes underneath.
|
|
|
|
✅ **VERIFIED 2026-09-13 — DCGM DOES SUPPORT OUR CARDS, and an earlier caveat in this
|
|
runbook claiming otherwise was WRONG and has been removed.**
|
|
|
|
Supported platforms, quoted: *"All NVIDIA Maxwell™ and newer **non-datacenter** (e.g.
|
|
NVIDIA® GeForce® or NVIDIA® Quadro®) GPUs"* — plus *"Starting with v1.3, limited DCGM
|
|
functionality is available on non-datacenter GPUs."*
|
|
|
|
And the feature-overview table settles what "limited" excludes — **not** configuration:
|
|
|
|
Feature Group Tesla Titan Quadro GeForce
|
|
Configuration Management X X X X
|
|
|
|
Configuration Management explicitly includes *"Power Limit: Set the maximum allowed power
|
|
consumption."* The thing actually gated on non-datacenter cards is **diagnostics**:
|
|
|
|
GPU Diagnostics (Levels 1,2,3): All Levels [Tesla]; Level 1 [Titan/Quadro/GeForce]
|
|
|
|
⚠ **One soft edge:** the table says "Quadro", the former name for the professional line.
|
|
RTX 6000 Ada and RTX PRO 6000 are its successors and should fall in that column, but the
|
|
table predates the rename — so that last step is inference, settled by one command on the
|
|
box.
|
|
|
|
⭐ **So the group-budget test is worth actually running**, not a curiosity. What it does
|
|
NOT settle is *distribution*: the docs still never say how a group budget is divided, and
|
|
the NVML argument is untouched — no bank-level register means per-GPU writes either way,
|
|
so the likely finding is static even division (= 4 x 250 W).
|
|
|
|
✅ **The static cap needs none of this.** `nvidia-smi -pl 250` is plain NVML and works on
|
|
these cards. DCGM would only buy the group-budget experiment and richer telemetry, and is
|
|
probably not even installed — `beszel-agent-nvidia` shells out to `nvidia-smi`.
|
|
|
|
### ⚠⚠ If a loop is written, its failure direction matters more than its logic
|
|
|
|
Power readings lag and `-pl` application takes tens of ms, so a reactive daemon
|
|
overshoots during a load RAMP — and the ramp is exactly the dangerous moment, because it
|
|
is the all-four-cards-loading-at-once case (the same shape as this box's ten
|
|
`restart: unless-stopped` containers starting together).
|
|
|
|
**Therefore: safe-by-default, opportunistic upward.** Boot every card at budget/N and
|
|
only ever RAISE a card's cap after observing idle neighbours. Never boot high and react
|
|
down. Failure mode then becomes "slower than it could have been" instead of "tripped the
|
|
breaker." Inverted, it is a thing that works for weeks and then fails on precisely the
|
|
event it existed to prevent.
|
|
|
|
### Why not yet
|
|
|
|
**4 x 250 W = 1000 W — the static cap IS the bank budget**, and it is the conservative
|
|
floor of the dynamic scheme rather than an alternative to it. The daemon's entire
|
|
contribution is the one-card-busy case: ~300 W instead of ~250 W on a single card, ~17%
|
|
more board power, which on a concave perf/watt curve is perhaps ~5% throughput. That is
|
|
the whole prize, against a control loop whose failure mode points at a breaker.
|
|
|
|
And whether that case is even common depends on workload mix. A **serving** fleet spreads
|
|
across cards by construction — one seat per card, gateway traffic split across aliases —
|
|
so single-card-busy is rare. A **training window** is the opposite: one card hammered,
|
|
three idle, which is where the dynamic scheme actually pays.
|
|
|
|
**Recommendation: static 250 W now, measure at the plug, build the loop only if the
|
|
measurements show single-card-busy is the common case.** It is a pure optimization on
|
|
top; adding it later re-architects nothing and would be built against real numbers.
|