caps: fv-ml1 250W / ana-ml3 200W — and nvidia-smi -pl caps BOARD power, not wall power
Operator set fv-ml1 at 250 W per card (83% of TGP, ~5% throughput) and ana-ml3 at 200 W (67%, ~10-15%). Records the term that decides whether 250 W actually clears a 15 A feed, because it is easy to drop: a power limit bounds BOARD power, and the wall sees that divided by PSU efficiency. Four cards at 250 W is 1000 W of board; add 180-300 W of host components and divide by ~0.90 and the plug sees ~1310-1445 W, against a 15 A circuit's 1440 W NEC continuous derating -- an inference box serving all day being a continuous load. So 250 W lands ON the limit rather than under it, where 200 W would give ~1090-1220 W with real margin. The deciding term is the host draw, which is still an estimate, so the procedure is: set 250 W, verify at the plug under four-card load, fall back to 200 W if it reads near 1440 W. A cap is a claim; the ammeter is the verification. Two consequences recorded alongside. Caps bound sustained draw and not transients -- the enforcement window is short but not instantaneous -- and while a breaker's thermal-magnetic curve forgives brief overload, a UPS's overload protection does not. So 250 W implicitly commits the fv-ml1 chassis to the PDU rather than behind the 1500 VA unit, which it exceeds even capped. And ana-ml3's 200 W across only two cards is deliberately conservative at 400 W total, relaxable if Anaheim's measured headroom beats its trip history.
This commit is contained in:
@@ -358,8 +358,10 @@ probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
||||
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
|
||||
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
|
||||
>
|
||||
> The 200 W cap directive applies to **ana-ml3**. It very likely wants applying to
|
||||
> fv-ml1 as well — see the four-card circuit arithmetic below.
|
||||
> **CAPS (operator, 2026-09-13):** **fv-ml1 → 250 W** per card (83% of TGP, ~5%
|
||||
> throughput). **ana-ml3 → 200 W** per card (67% of TGP, ~10-15%). See the plug-side
|
||||
> arithmetic below — 250 W is marginal on a 15 A circuit once PSU efficiency is counted,
|
||||
> and must be verified with the ammeter rather than assumed.
|
||||
|
||||
The generalised lesson from this outage: **decide the power envelope first and size the
|
||||
cards into it**, rather than installing cards and discovering the constraint by tripping
|
||||
@@ -408,10 +410,45 @@ guessed at, and idle-at-the-plug with all seats down measures it directly.
|
||||
4 x 300 W card + ~300 W host ~ 1500 W
|
||||
15 A circuit, 80% continuous = 1440 W
|
||||
|
||||
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all.** So
|
||||
capping belongs at **fv-ml1 too**, not only ana-ml3. If the four-card ammeter reading
|
||||
confirms it, fv-ml1 needs a per-card cap or a 20 A feed before anyone loads all four
|
||||
again — and note that today's incident only ever had TWO cards working.
|
||||
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all**, so
|
||||
capping belongs at fv-ml1 too. Note today's incident only ever had TWO of the four cards
|
||||
working; nobody has loaded all four.
|
||||
|
||||
### ⚠⚠⚠ AND `nvidia-smi -pl` CAPS BOARD POWER, NOT WALL POWER
|
||||
|
||||
This term is easy to drop and it is the one that decides whether 250 W clears a 15 A
|
||||
feed:
|
||||
|
||||
4 x 250 W board = 1000 W
|
||||
host components (board, 566 GB RAM,
|
||||
drives, fans) ~ 180-300 W
|
||||
-----------
|
||||
component total ~ 1180-1300 W
|
||||
/ PSU efficiency (~0.90) -> AT THE PLUG ~ 1310-1445 W
|
||||
|
||||
15 A circuit, NEC 80% continuous derating = 1440 W
|
||||
|
||||
**250 W lands ON the limit, not under it.** An inference box serving all day is a
|
||||
continuous load, so 1440 W is the design figure, not 1800.
|
||||
|
||||
At **200 W** the same arithmetic gives **~1090-1220 W at the plug** — comfortable, 200+ W
|
||||
of margin.
|
||||
|
||||
⚠ **So 250 W is PROBABLY fine and POSSIBLY not, and the deciding term is the host draw,
|
||||
which is still an estimate.** Procedure: set 250 W, then **verify at the plug under
|
||||
four-card load** before calling it done; fall back to 200 W if the reading comes in near
|
||||
1440 W. A cap is a claim; the ammeter is the verification.
|
||||
|
||||
⚠ **Power limits cap SUSTAINED draw, not transients.** The enforcement window is short
|
||||
but not instantaneous, so four cards at 250 W can momentarily exceed 1000 W of board.
|
||||
A breaker tolerates that (thermal-magnetic curves are forgiving of brief overload); a
|
||||
UPS's overload protection is not. Which means **250 W implicitly commits the chassis to
|
||||
the PDU rather than behind the 1500 VA unit** — even capped, 4 x 250 W + host exceeds
|
||||
that UPS's real rating.
|
||||
|
||||
⚠ **200 W on ana-ml3's TWO cards is deliberately conservative** (2 x 200 = 400 W is
|
||||
trivial on any circuit). Relaxable later if Anaheim's measured headroom beats its trip
|
||||
history; not a permanent figure.
|
||||
|
||||
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
||||
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
||||
|
||||
@@ -207,7 +207,7 @@ last, and do NOT restart the MTP campaign.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-13]` ⭐ **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** ⭐ **Two DIFFERENT boxes, clarified by operator 2026-09-13:** **fv-ml1** = 4x **Blackwell** RTX PRO 6000 **Max-Q @ 300 W** (Max-Q is the reduced-TGP SKU; Workstation Edition is 600 W), 391 GB VRAM, deployed. **ana-ml3** = 2x **Ada Generation** RTX 6000 @ 300 W, 96 GB VRAM, **NOT YET DEPLOYED**. The 200 W cap directive is for **ana-ml3** (2x200 = 400 W) — and ⭐ ana-ml3 lands in the **Anaheim** rack whose breaker tripped 2026-08-26 and 2026-09-11, one of those caused by this very chassis before it relocated, so the cap there is remediation of a known-bad circuit, not precaution. 4x200 W = 800 W of card, which fits a real circuit with a real UPS. This is the generalised lesson of the FV outage: decide the power envelope first and size the cards into it. At 300 W TGP, 200 W is a **67% cap — the favourable part of the concave perf/watt curve, ~10-15% throughput cost**, not the severe 33% cap a 600 W part would have meant; and 200 W is very unlikely to sit below a 300 W card's enforceable floor (still confirm with `nvidia-smi -q -d POWER | grep -iE 'power limit|default'`). ⭐⭐ **The outage arithmetic now has numbers:** 2x Blackwell Max-Q @300 W ≈ 600 W of card + host (566 GB RAM, drives, fans, PSU losses) ≈ 200-350 W = **~800-950 W against a 1500 VA Eaton's real ~900-1200 W rating** — at or just over the line, which is what explains a full day on ONE card (~500-650 W, inside) and death minutes into the SECOND. The host term is the only guess; idle-at-the-plug measures it. ⚠⚠ **And FOUR cards is a BREAKER problem, not a UPS problem:** 4x300 + ~300 host ≈ **1500 W vs a 15 A circuit's 1440 W continuous (80%) derating** — so **capping belongs at fv-ml1 too**, or it needs a 20 A feed, before anyone loads all four cards again. Today's incident only ever had TWO cards working. ⭐ **Decode tolerates caps far better than training** (memory-bandwidth-bound, concave curve). ⚠⚠ **ana-ml3's Ada is sm_89: native FP8 but NO NVFP4** (Blackwell-only) — most of our in-house quants are NVFP4 and will NOT run accelerated there; ana-ml3's seats want FP8 W8A8, or NVFP4 checkpoints stay on fv-ml1. ⭐ **This unparks [[parked_triton_backend_ampere_fp8]]** — a hard no on Ampere (fp8e4nv unsupported sm_86), explicitly deferred TO Ada, and sm_89 has the FP8 support it needs. **ana-ml3 VRAM is 2x48 = 96 GB** vs fv-ml1's 391 GB, so big-model placement stays at FV (Flash-Next needs 74 GiB resident on ONE card — the offload moves the table, not the experts — so it cannot run on a 48 GB Ada card at all). ⚠ **PERSIST the cap** (systemd unit + persistence mode, ordered before Docker): a hand-set limit stops holding at the next reboot, which is likely to be the very power event it existed to prevent. → `docs/runbooks/fv-site-dark-20260913.md`
|
||||
- `[2026-09-13]` ⭐ **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** ⭐ **Two DIFFERENT boxes, clarified by operator 2026-09-13:** **fv-ml1** = 4x **Blackwell** RTX PRO 6000 **Max-Q @ 300 W** (Max-Q is the reduced-TGP SKU; Workstation Edition is 600 W), 391 GB VRAM, deployed. **ana-ml3** = 2x **Ada Generation** RTX 6000 @ 300 W, 96 GB VRAM, **NOT YET DEPLOYED**. **CAPS: fv-ml1 → 250 W/card (83% of TGP, ~5% throughput cost); ana-ml3 → 200 W/card (67%, ~10-15%).** ⚠⚠⚠ **`nvidia-smi -pl` caps BOARD power, not WALL power** — 4x250 = 1000 W board + ~180-300 W host components = 1180-1300 W, ÷ ~0.90 PSU efficiency = **~1310-1445 W AT THE PLUG vs a 15 A circuit's 1440 W NEC continuous limit. 250 W lands ON the line, not under it** (200 W would give ~1090-1220 W, comfortable). **Procedure: set 250 W, then VERIFY at the plug under four-card load; fall back to 200 W if it reads near 1440 W.** ⚠ Caps bound SUSTAINED draw, not transients — breakers tolerate brief overload, UPS overload protection does not, so 250 W implicitly commits the fv-ml1 chassis to the PDU rather than behind the 1500 VA unit (even capped it exceeds that UPS). ⚠ ana-ml3's 200 W on 2 cards (2x200 = 400 W) — and ⭐ ana-ml3 lands in the **Anaheim** rack whose breaker tripped 2026-08-26 and 2026-09-11, one of those caused by this very chassis before it relocated, so the cap there is remediation of a known-bad circuit, not precaution. 4x200 W = 800 W of card, which fits a real circuit with a real UPS. This is the generalised lesson of the FV outage: decide the power envelope first and size the cards into it. At 300 W TGP, 200 W is a **67% cap — the favourable part of the concave perf/watt curve, ~10-15% throughput cost**, not the severe 33% cap a 600 W part would have meant; and 200 W is very unlikely to sit below a 300 W card's enforceable floor (still confirm with `nvidia-smi -q -d POWER | grep -iE 'power limit|default'`). ⭐⭐ **The outage arithmetic now has numbers:** 2x Blackwell Max-Q @300 W ≈ 600 W of card + host (566 GB RAM, drives, fans, PSU losses) ≈ 200-350 W = **~800-950 W against a 1500 VA Eaton's real ~900-1200 W rating** — at or just over the line, which is what explains a full day on ONE card (~500-650 W, inside) and death minutes into the SECOND. The host term is the only guess; idle-at-the-plug measures it. ⚠⚠ **And FOUR cards is a BREAKER problem, not a UPS problem:** 4x300 + ~300 host ≈ **1500 W vs a 15 A circuit's 1440 W continuous (80%) derating** — so **capping belongs at fv-ml1 too**, or it needs a 20 A feed, before anyone loads all four cards again. Today's incident only ever had TWO cards working. ⭐ **Decode tolerates caps far better than training** (memory-bandwidth-bound, concave curve). ⚠⚠ **ana-ml3's Ada is sm_89: native FP8 but NO NVFP4** (Blackwell-only) — most of our in-house quants are NVFP4 and will NOT run accelerated there; ana-ml3's seats want FP8 W8A8, or NVFP4 checkpoints stay on fv-ml1. ⭐ **This unparks [[parked_triton_backend_ampere_fp8]]** — a hard no on Ampere (fp8e4nv unsupported sm_86), explicitly deferred TO Ada, and sm_89 has the FP8 support it needs. **ana-ml3 VRAM is 2x48 = 96 GB** vs fv-ml1's 391 GB, so big-model placement stays at FV (Flash-Next needs 74 GiB resident on ONE card — the offload moves the table, not the experts — so it cannot run on a 48 GB Ada card at all). ⚠ **PERSIST the cap** (systemd unit + persistence mode, ordered before Docker): a hand-set limit stops holding at the next reboot, which is likely to be the very power event it existed to prevent. → `docs/runbooks/fv-site-dark-20260913.md`
|
||||
|
||||
- `[2026-09-13]` ⚠⚠⚠ **FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load test; all other sites healthy. Operator's leading hypothesis: the 1500 VA Eaton UPS overloaded and DIED.** It fits better than a breaker trip because a UPS's output rating sits far below the circuit's, making it the first protective device to give — which explains why the site let go at **two** cards loaded rather than four, and why the ~25 W OPNsense box died with it. ⚠ **Will not self-recover** (tripped needs a human, dead needs replacing) — do NOT poll FV. ⚠ **Do NOT use surge-only outlets to exceed a UPS rating**: both banks share one NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current limit. ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST** — all four cards every 10 s to the cut, on `/tank` not in a container, and the ONLY load measurement that exists. ⚠ **19 of 30 gateway aliases dark and NO local fallback** — every free local model was on fv-ml1, irv-ml1 runs no chat seat at all; the only non-fv chat backends are paid, and any coverage must be a NEW opt-in alias, never a silent repoint. ⭐ **OOB design gap**: OPNsense-as-subnet-router covers box-down/gateway-up and nothing for a site-wide loss, since the BMC's only route out is that gateway. ⚠ Recovery hazard: ten `restart: unless-stopped` vLLM containers will all load at once on power-up — mask Docker first, then `compose up -d` seat by seat (which also finishes the stale-homepage-label fix, since labels attach only at creation). → `docs/runbooks/fv-site-dark-20260913.md`, `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
|
||||
- `[2026-09-13]` ⭐⭐ **Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.** `stacks/flash-next-seat/`, fv-ml1 GPU 2 `:8022`, plus a `gen-large` LiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM **#54371 (UVA, merged 2026-09-09)** which **supersedes the paused #53899** — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960, `pidfd_getfd`/ptrace gate, stale-output-under-graphs) is designed out; in `v0.29.1rc0`, **not** `v0.29.0`. ⚠ **`text_config.ple_embedding_dtype` is the load-or-fail discriminator** for any community build. ⚠⚠ **`--kv-cache-memory` makes vLLM SKIP MEMORY PROFILING and ignore `--gpu-memory-utilization`** — 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off **pending measurement here, not written off** — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (`services/flash-next-mtp-bench/`, one `off_A` rep banked before the outage). ⚠ A container once ran `(healthy)` with `PORTS=[]` — verify `docker port`, not the healthcheck. → `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
|
||||
|
||||
Reference in New Issue
Block a user