Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md
T
vh 838132cd6b memory: snapshot — FV cross-site routing fixed, fleet conventions pinned
Session captured: the FV outbound-NAT root cause and its diagnostic signature,
the fv-ml1 dead man's switch, fleet identity/group/path conventions and the
root:docker normalization, nh3-dev's ts-input reachability fix, ESPHome
modernisation and the kb KB-search tool, and the Hermes bearer rotation
release. Six new detail files.

Tried-and-abandoned gains three: probing OPNsense endpoints by POSTing at them
(which rebooted the FV firewall), advertising a /32 from nh3-dev, and the
nh3-scale remote-site masquerade rules that fired but were not the fix.

Housekeeping: 8 Recent-decisions entries archived to archival-memory.md, and 21
oversized inline entries split into detail files per the two-tier rule -- they
had been sitting fully inline in the index, which is what the split exists to
prevent. Two pointers to a detail file archived this run were repointed at
archival-memory.md.

The index is 389 lines, still over the ~300 soft cap. The archival guards stop
it there: only 4 further entries are old enough to move and every one carries an
open deferred-work pointer. An over-cap file that keeps live decisions beats a
scannable one that lost a deferred call.
2026-09-15 00:53:48 -07:00

5.9 KiB

[2026-09-13] STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.

⭐ STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. ⭐ Two DIFFERENT boxes, clarified by operator 2026-09-13: fv-ml1 = 4x Blackwell RTX PRO 6000 Max-Q @ 300 W (Max-Q is the reduced-TGP SKU; Workstation Edition is 600 W), 391 GB VRAM, deployed. ana-ml3 = 2x Ada Generation RTX 6000 @ 300 W, 96 GB VRAM, NOT YET DEPLOYED. CAPS: fv-ml1 → 250 W/card (83% of TGP, ~5% throughput cost); ana-ml3 → 200 W/card (67%, ~10-15%). ⚠⚠⚠ nvidia-smi -pl caps BOARD power, not WALL power — 4x250 = 1000 W board + ~180-300 W host components = 1180-1300 W, ÷ ~0.90 PSU efficiency = ~1310-1445 W AT THE PLUG vs a 15 A circuit's 1440 W NEC continuous limit. 250 W lands ON the line, not under it (200 W would give ~1090-1220 W, comfortable). Procedure: set 250 W, then VERIFY at the plug under four-card load; fall back to 200 W if it reads near 1440 W. ⚠ Caps bound SUSTAINED draw, not transients — breakers tolerate brief overload, UPS overload protection does not, so 250 W implicitly commits the fv-ml1 chassis to the PDU rather than behind the 1500 VA unit (even capped it exceeds that UPS). ⭐ BANK-level cap asked about and answered: DCGM_CONFIG_POWER_BUDGET_GROUP ("power budget for the entire group") exists alongside DCGM_CONFIG_POWER_CAP_INDIVIDUAL, but the docs do not say how it distributes — and the only primitive underneath is NVML's PER-GPU nvmlDeviceSetPowerManagementLimit (no bank-level register), so any group budget resolves to N per-GPU writes. "Each card free until all are loaded" is therefore a CONTROL LOOP, not a hardware feature. ⚠⚠ If one is written it must be safe-by-default, opportunistic upward — boot at budget/N and only RAISE after observing idle neighbours; a reactive loop overshoots during a load RAMP, which is the all-cards-at-once case it exists to prevent. Not yet worth it: 4x250 = 1000 W IS the bank budget, so the daemon's whole prize is the one-card-busy case (~17% more board power ≈ ~5% throughput) — rare for a SERVING fleet (one seat per card), valuable for a TRAINING window. DCGM = NVIDIA's first-party Data Center GPU Manager (Apache-2.0), layered ABOVE NVML: nvidia-smi → NVML per-GPU primitives → DCGM daemon/dcgmi for health, diagnostics, config enforcement, groups. ✅ VERIFIED 2026-09-13 (I had claimed the opposite and was WRONG): DCGM SUPPORTS our cards and power config works on them. Supported platforms cover "All NVIDIA Maxwell and newer NON-DATACENTER (e.g. GeForce or Quadro) GPUs", and the feature-overview table marks Configuration Management ✓ for Tesla/Titan/Quadro/GeForce alike — including "Power Limit: set the maximum allowed power consumption". What IS gated on non-datacenter cards is diagnostics (Level 1 only vs All Levels on Tesla), not config. ⚠ Soft edge: the table says "Quadro", the former professional-line name; RTX 6000 Ada / RTX PRO 6000 are its successors and should fall in that column, but the table predates the rename — one command on the box settles it. ⭐ So the group-budget test IS worth running; what remains unsettled is distribution, not availability. ✅ The static cap needs none of it: nvidia-smi -pl 250 is plain NVML and works on these cards; DCGM is probably not even installed (beszel-agent-nvidia shells out to nvidia-smi). ⚠ ana-ml3's 200 W on 2 cards (2x200 = 400 W) — and ⭐ ana-ml3 lands in the Anaheim rack whose breaker tripped 2026-08-26 and 2026-09-11, one of those caused by this very chassis before it relocated, so the cap there is remediation of a known-bad circuit, not precaution. 4x200 W = 800 W of card, which fits a real circuit with a real UPS. This is the generalised lesson of the FV outage: decide the power envelope first and size the cards into it. At 300 W TGP, 200 W is a 67% cap — the favourable part of the concave perf/watt curve, ~10-15% throughput cost, not the severe 33% cap a 600 W part would have meant; and 200 W is very unlikely to sit below a 300 W card's enforceable floor (still confirm with nvidia-smi -q -d POWER | grep -iE 'power limit|default'). ⭐⭐ The outage arithmetic now has numbers: 2x Blackwell Max-Q @300 W ≈ 600 W of card + host (566 GB RAM, drives, fans, PSU losses) ≈ 200-350 W = ~800-950 W against a 1500 VA Eaton's real ~900-1200 W rating — at or just over the line, which is what explains a full day on ONE card (~500-650 W, inside) and death minutes into the SECOND. The host term is the only guess; idle-at-the-plug measures it. ⚠⚠ And FOUR cards is a BREAKER problem, not a UPS problem: 4x300 + ~300 host ≈ 1500 W vs a 15 A circuit's 1440 W continuous (80%) derating — so capping belongs at fv-ml1 too, or it needs a 20 A feed, before anyone loads all four cards again. Today's incident only ever had TWO cards working. ⭐ Decode tolerates caps far better than training (memory-bandwidth-bound, concave curve). ⚠⚠ ana-ml3's Ada is sm_89: native FP8 but NO NVFP4 (Blackwell-only) — most of our in-house quants are NVFP4 and will NOT run accelerated there; ana-ml3's seats want FP8 W8A8, or NVFP4 checkpoints stay on fv-ml1. ⭐ This unparks parked_triton_backend_ampere_fp8 — a hard no on Ampere (fp8e4nv unsupported sm_86), explicitly deferred TO Ada, and sm_89 has the FP8 support it needs. ana-ml3 VRAM is 2x48 = 96 GB vs fv-ml1's 391 GB, so big-model placement stays at FV (Flash-Next needs 74 GiB resident on ONE card — the offload moves the table, not the experts — so it cannot run on a 48 GB Ada card at all). ⚠ PERSIST the cap (systemd unit + persistence mode, ordered before Docker): a hand-set limit stops holding at the next reboot, which is likely to be the very power event it existed to prevent. → docs/runbooks/fv-site-dark-20260913.md