90e71ca75e7572fa1a625f85525de060a8a91f7d
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5a5f5c267e |
power: RETRACT the DCGM caveat — config management and power limits ARE supported on our cards
I claimed in
|
||
|
|
100670eed1 |
power: what DCGM is, and why not to plan around it on workstation-SKU cards
DCGM is NVIDIA's own Data Center GPU Manager -- first-party, Apache-2.0, packaged as datacenter-gpu-manager -- and it layers above NVML rather than beside it: nvidia-smi is a thin CLI over NVML's per-GPU primitives, and DCGM is a daemon plus dcgmi adding health, diagnostics, config enforcement, policy and group abstractions on top. Which is why its group notion still resolves to N per-GPU writes underneath. The caveat that matters, and it undercuts the experiment suggested in the previous commit: DCGM is datacenter-oriented and parts of it are gated to datacenter SKUs of the Tesla/A100/H100 class. Our cards are professional/workstation parts -- RTX PRO 6000 Blackwell Max-Q and RTX 6000 Ada -- and several DCGM capabilities are unsupported or degraded outside that line, plausibly including config enforcement, which is precisely the power path. So DCGM_CONFIG_POWER_BUDGET_GROUP may return 'unsupported on this device'. Downgraded from 'worth testing' to five minutes of curiosity after the real work, and explicitly not a planning assumption. None of which touches the plan: nvidia-smi -pl 250 is plain NVML and works on these cards. DCGM would only have bought the group-budget experiment and nicer telemetry, and is probably not installed anyway since beszel-agent-nvidia shells out to nvidia-smi. |
||
|
|
94fb7b7208 |
power: answer the bank-budget question — DCGM has the concept, the dynamic part is a control loop, and 4x250 already is 1000 W
DCGM_CONFIG_POWER_BUDGET_GROUP ('the power budget for the entire group') exists
alongside DCGM_CONFIG_POWER_CAP_INDIVIDUAL, so the concept is first-class. The docs do
not state how a group budget is distributed, and the deduction is that it cannot be
anything exotic: the only enforcement primitive underneath is NVML's per-GPU
nvmlDeviceSetPowerManagementLimit and there is no bank-level register, so any group
budget resolves to N per-GPU writes. Static even division is one write each; 'each card
free until they are all loaded' requires continuous re-writing, which is a control loop
rather than a hardware feature. Worth a ten-minute test when the box returns, in case
NVIDIA already runs that loop.
Records the design constraint that matters more than the logic: power readings lag and
-pl application takes tens of milliseconds, so a reactive daemon overshoots during a load
ramp -- and the ramp is the dangerous moment, being the same all-cards-at-once shape as
this box's ten restart:unless-stopped containers starting together. So any such loop must
be safe-by-default and opportunistic upward: boot at budget/N, only ever raise after
observing idle neighbours. Inverted, it works for weeks and then fails on precisely the
event it existed to prevent.
And the reason to defer it: 4x250 W is already 1000 W, so the static cap is the
conservative floor of the dynamic scheme rather than an alternative. The daemon's entire
contribution is the one-card-busy case, worth perhaps 5% throughput, which is rare for a
serving fleet that puts one seat per card and common only for a training window.
|
||
|
|
b538fde6f0 |
caps: fv-ml1 250W / ana-ml3 200W — and nvidia-smi -pl caps BOARD power, not wall power
Operator set fv-ml1 at 250 W per card (83% of TGP, ~5% throughput) and ana-ml3 at 200 W (67%, ~10-15%). Records the term that decides whether 250 W actually clears a 15 A feed, because it is easy to drop: a power limit bounds BOARD power, and the wall sees that divided by PSU efficiency. Four cards at 250 W is 1000 W of board; add 180-300 W of host components and divide by ~0.90 and the plug sees ~1310-1445 W, against a 15 A circuit's 1440 W NEC continuous derating -- an inference box serving all day being a continuous load. So 250 W lands ON the limit rather than under it, where 200 W would give ~1090-1220 W with real margin. The deciding term is the host draw, which is still an estimate, so the procedure is: set 250 W, verify at the plug under four-card load, fall back to 200 W if it reads near 1440 W. A cap is a claim; the ammeter is the verification. Two consequences recorded alongside. Caps bound sustained draw and not transients -- the enforcement window is short but not instantaneous -- and while a breaker's thermal-magnetic curve forgives brief overload, a UPS's overload protection does not. So 250 W implicitly commits the fv-ml1 chassis to the PDU rather than behind the 1500 VA unit, which it exceeds even capped. And ana-ml3's 200 W across only two cards is deliberately conservative at 400 W total, relaxable if Anaheim's measured headroom beats its trip history. |
||
|
|
2da0c76d99 |
correct the hardware: fv-ml1 is 4x Blackwell Max-Q 300W, ana-ml3 is 2x Ada RTX 6000 — and four cards is a breaker problem
Operator clarification, and it separates two boxes I had been conflating. fv-ml1 is 4x Blackwell RTX PRO 6000 Max-Q at 300 W each (Max-Q being the reduced-TGP SKU; the Workstation Edition is the 600 W part), 391 GB VRAM, deployed and currently dark. ana-ml3 is 2x Ada Generation RTX 6000 at 300 W, 96 GB VRAM, not yet deployed. The 200 W cap directive is ana-ml3's. With the TGP known, the outage stops being a vague 'undersized' and acquires a mechanism: two Max-Q cards at 300 W is ~600 W of card, plus a host carrying 566 GB of RAM, drives, fans and PSU conversion loss at perhaps 200-350 W, against an Eaton 1500 VA's real ~900-1200 W. That lands at or just over the rating, which is precisely what explains a full day of service on one card and failure minutes into the second. The host term is the only one being guessed; idle-at-the-plug measures it directly. It also surfaces something that is not a UPS question at all. Four cards at 300 W plus ~300 W of host is ~1500 W against a 15 A circuit's 1440 W continuous derating, so four cards uncapped is marginal on the breaker with no UPS in the path. Capping therefore belongs at fv-ml1 as well as ana-ml3, or fv-ml1 needs a 20 A feed -- and worth noting today's incident only ever had two of the four cards working. ana-ml3's placement constraints sharpen too: sm_89 has native FP8 but no NVFP4, so the in-house NVFP4 quants stay at FV, and at 96 GB total it cannot host the Flash-Next seat at all -- that needs 74 GiB resident on a single card, and the offload moves the n-gram table rather than the experts. |
||
|
|
8fcc26e2c9 |
policy(gpu-power): cards are RTX 6000 Ada at 300 W — 200 W is a mild cap, plus two sm_89 placement consequences
Corrects the SKU: RTX 6000 Ada, 300 W, not the ~600 W initially recalled. That makes 200 W a cap to 67% of TGP -- the favourable part of the concave perf/watt curve, roughly 10-15% of throughput -- rather than the severe 33% cap a 600 W part would have implied, and it very likely sits above the card's enforceable floor, so the check becomes a formality rather than a gate. The protective value is worth stating: four cards at 300 W uncapped is ~1200 W, which is roughly the neighbourhood that overwhelmed a 1500 VA unit at FV with only TWO Blackwell cards drawing. Capping to 800 W makes a repeat of today a non-event. Two consequences that follow from Ada independent of power, and both are placement constraints rather than details. sm_89 has native FP8 but NOT NVFP4, which is Blackwell-only -- so the in-house NVFP4 quants that most of this fleet runs will not be accelerated on that colo's cards, and its seats want FP8 W8A8 builds or the NVFP4 checkpoints stay at FV. And it unparks the triton-backend item, which is a hard no on Ampere because fp8e4nv is unsupported on sm_86 and was explicitly deferred to Ada; sm_89 has what it needs. VRAM is 4x48 = 192 GB against fv-ml1's 391 GB, so big-model placement stays at FV. The Flash-Next seat needs 74 GiB resident on one card and would not fit a 48 GB Ada card even with the n-gram table offloaded -- the offload moves the table, not the experts. |
||
|
|
3e61d7d4e0 |
policy: cap GPU power limits at build time — 200 W for the other colo's cards
Operator directive, and the right generalisation of the FV outage: decide the power envelope first and size the cards into it, rather than installing cards and discovering the constraint by tripping it. Four cards at 200 W is 800 W, which fits a real circuit with a real UPS and headroom. Records three things to settle before it is a plan. First, 200 W may sit below the card's enforceable floor: nvidia-smi -pl is bounded by Min Power Limit, often around half of TGP on a high-TGP part, and a sub-floor request is refused -- quietly, depending on how it is scripted. Run nvidia-smi -q -d POWER before any build planning depends on the number. Second, the 600 W figure wants confirming against the actual SKU. The Ada parts do not land there -- RTX 6000 Ada is 300 W, L40/L40S 300/350 W, 4090 450 W -- while 600 W is Blackwell RTX PRO 6000 Workstation territory, so these may be Blackwell or the figure may be a two-card total. Read it off the device rather than a spec sheet. Third, the workload asymmetry is in this fleet's favour: decode is memory-bandwidth-bound and tolerates a cap far better than training does, with a concave perf/watt curve where 60-70% of TGP costs roughly 10-15% of throughput. A cap to a third of TGP is deeper into the steep region; measure it on the first card rather than predicting, and expect prefill-heavy and training work to pay more than a serving seat. And persist the cap. A hand-set limit holds until the next reboot and then silently stops holding, which is the worst shape available given that the thing rebooting the box is likely to be the power event the cap existed to prevent. |
||
|
|
b6335bf6ad |
runbook(fv-outage): the circuit case — split power survives a trip on battery, but only if the colo handoff does
Operator: 'unless of course the thing trips the circuit anyway.' Correct, and it splits into two halves with different answers. A breaker trip is the event the split-power proposal survives: firewall + BMC is 25-40 W on a 1500 VA unit, which is hours of battery, and on a trip the UPS stops being a load-bearing supply and goes back to being what it is for. What it does NOT cover is the colo's own handoff -- their switch, ONT or demarc. If that sits on the circuit we just tripped, the outcome is a firewall running on battery with nothing upstream to talk to and the drive happens anyway. Added as a question for the facility, because it decides whether split power delivers remote diagnosis or merely feels like it does. Records the case where none of it matters: removing an undersized UPS does not remove the constraint, it promotes the next one -- UPS ~900-1200 W to circuit ~1800 W at 15 A or ~2400 W at 20 A. Which side the four-card figure lands on decides everything, which is what makes that single ammeter reading the load-bearing measurement of the visit. Surfaces the lever that may avoid an electrician entirely: nvidia-smi -pl caps per-card TGP, so the box can be made to fit its feed at a throughput cost rather than a rewiring cost. Read nvidia-smi -q -d POWER for the enforced range before assuming how much room the dial has, and persist any cap -- one that evaporates on reboot will hold right up until the next power event and then silently stop holding. |
||
|
|
00b842bb9b |
runbook(fv-outage): operator ruling — undersized UPS; NAT demoted; ammeter protocol for the visit
Operator's reasoning, accepted and better than the hypothesis-space argument it replaces: the NAT change went effective, was verified bidirectional, and then ran correctly for twenty minutes before the site died the moment GPU load was applied. A working config change does not spontaneously fail under an unrelated physical variable. The load correlation is tight; the NAT correlation is merely adjacent in time. Undersized UPS is the only candidate that explains the trigger. NAT material retained as record, and the power.log/uptime check demoted from decision point to free confirmation. Adds the measurement protocol, since the operator is bringing a PDU and an ammeter. The load-bearing caveat: power.log is GPU-ONLY -- nvidia-smi per-card, excluding CPU, 566 GB of RAM, drives, fans and PSU conversion losses -- so the ammeter at the plug is the primary instrument and power.log only cross-checks the GPU share. Four states to capture (idle, one card, two cards, four cards), and capture PEAK rather than average: UPS overload protection responds to short-term overload, so an average-only reading that hides transients will mis-size the replacement exactly the way the present unit got mis-sized, and must be recorded as a floor rather than as the draw. The four-card figure is earmarked for servers/fv-ml1/README.md, because it closes the cutover's own open question -- that the FV circuit was likely specced against half the real draw, back when every record still said the box had two GPUs. |
||
|
|
59ddedd980 |
runbook(fv-outage): a NAT change 34 min earlier means power is not established — and power.log settles it for free
Another session applied a scoped Tailscale SNAT rule to the FV gateway at ~06:22Z, 34 minutes before the site went dark (docs/runbooks/fv-to-ana-nat.md, not my work, left uncommitted). That makes the UPS-overload theory a hypothesis rather than a finding, and nobody should buy hardware on it until the discriminator below has been read. On the evidence that change is the wrong shape to have caused this, and it is recorded as such so the visit is not wasted chasing it: one OUTBOUND SNAT rule scoped to a single source /32 and a single destination /16 cannot stop the gateway, the BMC or the public WAN address from answering inbound; no routes, filter rules, WAN settings or subnet advertisements were touched; pfctl -sr came back byte-identical; and it was verified bidirectional afterwards including ANA->FV SSH with Beszel 18/18 up. Their BMC datapoint used 10.251.50.50, which is not the BMC -- that is 10.251.250.50, a different subnet. They correctly declined to claim BMC health, but the observation is void rather than negative and should not be reasoned from. The discriminator costs nothing and is already on disk: power.log is written locally to /tank every 10 s by a shell loop on the box and does not depend on the network. Entries past 06:56Z mean the machine never lost power, which makes this a routing fault and the UPS innocent; entries stopping at 06:56Z confirm power. Cross-check with uptime and journalctl --list-boots -- continuous uptime across 06:56Z kills the UPS theory outright. So the first action on site is now to READ, not to fix. The two hypotheses lead to completely different remediations and only one of them needs a new UPS. |
||
|
|
312725ddfb |
memory: snapshot — Flash-Next seat on one card, and the FV outage that followed
Durable capture so tomorrow's session does not have to reconstruct either half. Built and verified before the power failed: Qwen3.8-Flash-Next serving on a single RTX PRO 6000 with its 51B n-gram table pinned in host RAM and read over CUDA UVA -- 74.36 GiB weights resident, 14.00 GiB KV for 560,654 tokens at the full 262,144 context, 67 GiB host RSS -- plus a gen-large gateway alias verified end to end. The five findings worth carrying: the offload is #54371 (UVA, merged) which supersedes the paused worker-based #53899 and designs out its entire bug family; text_config.ple_embedding_dtype is the load-or-fail discriminator for any community build; --kv-cache-memory makes vLLM SKIP memory profiling and ignore gpu-memory-utilization, which inverts the usual pin-bytes advice and let a 16 GiB pin nearly OOM with no visible failure; MTP is off pending measurement here rather than written off, because the recipe's number is cross-harness and tested k=3 only while the head is one layer run autoregressively; and a container once reported (healthy) with no published port at all, because the healthcheck runs inside the boundary it was trusted to validate. Then the outage. Records it as will-not-self-recover, so no session wastes effort polling a dead site, and carries the three things that change the visit: bypass the UPS rather than using its surge-only bank (both banks share one 12 A inlet -- the surge bank bypasses the inverter, not the current rating), recover power.log before anything else because it is the only load measurement that exists anywhere, and bring seats up one at a time because ten restart:unless-stopped containers loading at once is the largest transient the box can make into whatever just failed. Also records what is still half-done: the stale homepage labels on the 10 containers that died before they could be recreated, which the staged bring-up fixes as a side effect, and the eight drifted stacks plus three untracked host-only stacks that were deliberately left for a deliberate reconciliation. |
||
|
|
d79f10457a |
runbook(fv-outage): UPS overload as leading hypothesis, site-visit bring-list, no-local-fallback correction
Operator's read is that the UPS the box was plugged into overloaded and died, and it fits better than the breaker-trip theory: a UPS's output rating sits far below the circuit's, so it is the first protective device to give -- which explains why the site let go at TWO cards loaded rather than four, and why the ~25 W firewall died with it. Records the operationally important consequence: a tripped UPS resets, an overloaded one can kill its output stage permanently. If it is dead, nothing on site can be reset back to life, so the visit needs the means to BYPASS the UPS or it is wasted. Elevates recovery of /tank/.../power.log to the first action on site. It sampled all four cards every 10 s up to the cut, lives on /tank rather than in a container, and is the only measurement of what the load actually drew -- without it a replacement UPS gets sized by guesswork. Also states that no load figure exists yet, only idle. Corrects an earlier claim of mine in this session: there is NO local fallback for the 19 dark aliases. Probed -- every free local model is on fv-ml1, and irv-ml1 runs no chat seat at all, only TTS/ComfyUI/arbo/clipper work on two partly-occupied Ampere cards. The only non-fv chat backends are paid. Any paid coverage must go under a new opt-in alias name rather than a silent repoint of summarizer/gen/classifier. |
||
|
|
969a1b64a2 |
runbook: FV site dark 2026-09-13 — outage facts, blast radius, staged recovery, OOB design gap
Written while the site is down so recovery does not have to be reconstructed later. Records what was measured rather than what is suspected: every FV address including the BMC is unreachable while all three other sites answer, the campaign's last log line was off_A rep 2 at 06:56:04Z, and the site was dark by 06:58:40Z. Names three candidate causes with the evidence that would distinguish them, because the instrument that could have settled it -- the per-card power log -- died with the box. The two-card-load hypothesis fits the timing and the two prior Anaheim breaker trips on this same chassis, but it is circumstantial and is recorded as such. Carries the recovery hazard that matters: every seat on the box is restart:unless-stopped, so resetting power alone brings ten vLLM containers up loading at once -- the largest transient the box can produce, into a circuit that may have just tripped. Staged sequence given, gen first and flash-next last. Also records the OOB gap the outage exposes: OPNsense-as-subnet-router protects against box-down/gateway-up, and not at all against the site-wide loss that actually happened, because the BMC's only path out is through that same gateway. |