Operator clarification, and it separates two boxes I had been conflating. fv-ml1 is 4x Blackwell RTX PRO 6000 Max-Q at 300 W each (Max-Q being the reduced-TGP SKU; the Workstation Edition is the 600 W part), 391 GB VRAM, deployed and currently dark. ana-ml3 is 2x Ada Generation RTX 6000 at 300 W, 96 GB VRAM, not yet deployed. The 200 W cap directive is ana-ml3's. With the TGP known, the outage stops being a vague 'undersized' and acquires a mechanism: two Max-Q cards at 300 W is ~600 W of card, plus a host carrying 566 GB of RAM, drives, fans and PSU conversion loss at perhaps 200-350 W, against an Eaton 1500 VA's real ~900-1200 W. That lands at or just over the rating, which is precisely what explains a full day of service on one card and failure minutes into the second. The host term is the only one being guessed; idle-at-the-plug measures it directly. It also surfaces something that is not a UPS question at all. Four cards at 300 W plus ~300 W of host is ~1500 W against a 15 A circuit's 1440 W continuous derating, so four cards uncapped is marginal on the breaker with no UPS in the path. Capping therefore belongs at fv-ml1 as well as ana-ml3, or fv-ml1 needs a 20 A feed -- and worth noting today's incident only ever had two of the four cards working. ana-ml3's placement constraints sharpen too: sm_89 has native FP8 but no NVFP4, so the in-house NVFP4 quants stay at FV, and at 96 GB total it cannot host the Flash-Next seat at all -- that needs 74 GiB resident on a single card, and the offload moves the n-gram table rather than the experts.
441 lines
24 KiB
Markdown
441 lines
24 KiB
Markdown
# FV site dark — 2026-09-13 ~06:56Z
|
|
|
|
**Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley.**
|
|
Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one
|
|
needs replacing, so there is no recovery to wait for and no point polling the site.
|
|
Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is
|
|
down so recovery does not have to be reconstructed from memory.
|
|
|
|
## What is down
|
|
|
|
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
|
|
|
|
| target | result |
|
|
|---|---|
|
|
| `fv-ml1` 10.251.50.54 | 100% loss |
|
|
| `fv-ml1` mesh 100.64.0.7 | 100% loss |
|
|
| FV gateway 10.251.50.1 | 100% loss |
|
|
| FV gateway mesh 100.64.0.8 | 100% loss |
|
|
| **fv-ml1-bmc 10.251.250.50** | **100% loss** |
|
|
| `fv.phasefinal.com` (172.83.89.66) | no answer |
|
|
|
|
**My own path is healthy** — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all
|
|
answer with 0% loss, so this is FV-local, not an isolation of the observer.
|
|
|
|
## Timeline
|
|
|
|
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
|
|
the two-card power risk. Production seat already live on GPU 2.
|
|
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
|
|
06:54:39 off_A arm healthy after 211 s
|
|
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
|
|
06:56:04 off_A rep 2 started <-- LAST LOG LINE
|
|
06:58:40 every FV address unreachable, including the BMC
|
|
|
|
So the site went dark inside a ~2.5 minute window, roughly five minutes into
|
|
sustained bench load on GPU 3 with GPU 2 also resident/serving.
|
|
|
|
## ⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
|
|
|
|
Operator's read, and it fits the evidence **better** than the breaker-trip theory
|
|
below, for a reason worth writing down:
|
|
|
|
**A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective
|
|
device to give.** A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit
|
|
carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM,
|
|
plus the firewall, plausibly clears the UPS rating while staying comfortably under the
|
|
breaker. That explains the thing a breaker trip explains poorly: **why it let go at
|
|
TWO cards loaded and not four.** It also explains the OPNsense box dying -- same UPS.
|
|
|
|
⚠ **"Died" is the operationally important word.** A tripped UPS resets; an overloaded
|
|
one can kill its output stage or battery pack permanently. If it is dead rather than
|
|
tripped, **nothing on site can be reset back to life** and the trip is wasted without
|
|
the means to bypass it.
|
|
|
|
### Bring / check list for the site visit
|
|
|
|
- **The means to BYPASS the UPS entirely** -- chassis and firewall straight to the
|
|
PDU or wall. Do this first, get the site back, diagnose the UPS after.
|
|
- **Read off the UPS before moving it:** make/model, VA and W rating, fault LEDs, and
|
|
any LCD or event log entry (many units record "overload" explicitly).
|
|
- **Note which outlets were battery-backed vs surge-only.** Mixed-bank units are
|
|
common, and a GPU chassis on the battery-backed bank overloads soonest.
|
|
- ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST, before any
|
|
bring-up.** It sampled all four cards every 10 s right up to the cut and is the
|
|
ONLY measurement of what the load actually drew. It lives on `/tank`, not in a
|
|
container, so it survived. Without it, the replacement UPS gets sized by guesswork.
|
|
- ⚠ **Size the replacement for FOUR cards under load, not two.** Anything sized to
|
|
today's failure just relocates the trip point to the next person who loads all four.
|
|
|
|
⚠ **No measured load figure exists yet.** Idle draw was measured (GPU0 3.80 W, GPU1
|
|
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
|
|
and is recoverable only from `power.log`. Do not let anyone put a wattage in a
|
|
purchasing decision until that file has been read.
|
|
|
|
## ⭐ OPERATOR RULING 2026-09-13: undersized UPS. NAT hypothesis DEMOTED.
|
|
|
|
> "highly doubt the nat hypothesis, it went effective (connectivity confirmed for
|
|
> previously dark path) -- then 20 minutes later, during load, site went dark. that's
|
|
> pretty unlikely to be the cause. for sure I think the ups was undersized."
|
|
|
|
Accepted, and the reasoning is better than the hypothesis-space argument below: **a
|
|
config change that went effective, was verified bidirectional, and then ran correctly
|
|
for twenty minutes does not spontaneously fail when an unrelated physical variable --
|
|
someone else's GPU load -- is introduced.** The load correlation is tight; the NAT
|
|
correlation is merely adjacent in time. Undersized UPS is the only candidate that
|
|
explains the *trigger*.
|
|
|
|
The NAT material below is retained as record, not as a live competing theory, and the
|
|
`power.log` / `uptime` check is retained as **free confirmation** rather than as a
|
|
decision point.
|
|
|
|
## Measurement plan for the site visit (operator bringing a PDU + ammeter)
|
|
|
|
⚠ **`power.log` is GPU-ONLY.** It samples `nvidia-smi` per-card draw and does **not**
|
|
include the host: CPU, 566 GB of RAM, drives, fans, or PSU conversion losses. The
|
|
number that matters against a UPS rating is the whole chassis **at the plug**. The
|
|
ammeter is therefore the primary instrument and `power.log` is a cross-check on the
|
|
GPU share.
|
|
|
|
Capture four states -- this is the first real sizing data that has ever existed for
|
|
this box:
|
|
|
|
| state | why it matters |
|
|
|---|---|
|
|
| all seats down, idle | the floor (GPU idle measured 3.80 / 3.88 / 14.27 / 6.75 W) |
|
|
| one card loaded | the condition that ran fine for a day |
|
|
| **two cards loaded** | the condition that took the site down |
|
|
| **four cards loaded** | the only number that can size a replacement honestly |
|
|
|
|
⚠⚠ **CAPTURE PEAK, NOT AVERAGE.** GPU power has fast transients and UPS overload
|
|
protection responds to short-term overload, so a 1-second-average reading can
|
|
under-read peaks badly. Use max-hold/peak capture if the meter has it. A figure like
|
|
"1100 W average" that hides 1600 W spikes will mis-size the replacement the same way
|
|
the current unit got mis-sized. **If the meter is average-only, record the number as a
|
|
FLOOR, not as the draw.**
|
|
|
|
⭐ **The four-card figure goes into `servers/fv-ml1/README.md` permanently.** The
|
|
cutover notes flagged that the FV circuit was "likely specced against half the real
|
|
draw" while every record still said two GPUs; this closes that with a measurement
|
|
instead of an assumption.
|
|
|
|
## The NAT change, retained as record (DEMOTED — see the ruling above)
|
|
|
|
**Another session applied a Tailscale SNAT rule to the FV gateway at ~06:22Z — 34
|
|
minutes before the site went dark.** See `docs/runbooks/fv-to-ana-nat.md` (uncommitted
|
|
as of this writing; not my work, left alone). So "we overloaded the power" is no longer
|
|
the only live hypothesis, and the UPS should not be replaced on the strength of a theory
|
|
until the one below is run.
|
|
|
|
**On the evidence, that change is the WRONG SHAPE to have caused this**, and I want that
|
|
on the record so nobody wastes the visit chasing it:
|
|
|
|
- It is one **outbound** SNAT rule, source-scoped to `10.251.50.54/32`, destination-
|
|
scoped to `10.250.0.0/16`. An outbound NAT rule cannot stop the gateway, the BMC or
|
|
the public WAN address from answering **inbound**.
|
|
- The runbook states no routes, filter rules, WAN settings or subnet advertisements were
|
|
touched, and that `pfctl -sr` came back byte-identical.
|
|
- It was verified working in **both** directions afterwards: FV→hub HTTP 200, FV→ANA
|
|
TCP 5432, FV internet HTTPS 200, **ANA→FV SSH reachable**, gateway management intact,
|
|
Beszel **18/18 up**.
|
|
|
|
⚠ Note their BMC observation used **`10.251.50.50`**, which is not the BMC — the BMC is
|
|
**`10.251.250.50`**, a different subnet. They correctly declined to claim BMC health, but
|
|
the datapoint is *void*, not negative. Do not reason from it either way.
|
|
|
|
### ⭐⭐ THE DISCRIMINATOR — run this before forming any conclusion
|
|
|
|
`power.log` is written **locally on `/tank`, every 10 seconds, by a shell loop on the
|
|
box.** It does not depend on the network. So:
|
|
|
|
| `power.log` last entry | what it means |
|
|
|---|---|
|
|
| **past 06:56Z** | the box **never lost power**. This is a routing/gateway fault, and the UPS is innocent. |
|
|
| **stops at ~06:56Z** | the box lost power. UPS/circuit confirmed. |
|
|
|
|
Cross-check with `uptime` and `journalctl --list-boots` the moment there is a console:
|
|
**continuous uptime across 06:56Z kills the UPS theory outright.**
|
|
|
|
⚠ **So the FIRST action on site is to read, not to fix.** `uptime`,
|
|
`journalctl --list-boots`, then `tail power.log`. Establish whether the machine ever
|
|
went down before anyone buys hardware or flips anything — the two hypotheses lead to
|
|
completely different remediations and only one of them needs a new UPS.
|
|
|
|
## Three further candidate causes, and what distinguishes them
|
|
|
|
Cannot be distinguished remotely, because every FV path — including the BMC —
|
|
traverses the OPNsense gateway, and the gateway is also dark.
|
|
|
|
1. **Circuit tripped under two-card load.** Fits the timing and is the predicted
|
|
failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this
|
|
same chassis, and the FV circuit was specced while every record still said the box
|
|
had **two** GPUs rather than four. ⭐ Distinguishing evidence: **the breaker is
|
|
visibly tripped**, and the OPNsense box is dark too (it draws ~20-30 W and would
|
|
survive anything short of a circuit/utility loss).
|
|
2. **OPNsense gateway crashed or rebooted**, taking all FV routing with it while the
|
|
GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and
|
|
the OPNsense box is the only dead thing.
|
|
3. **Upstream utility or colo-side power/network loss**, unrelated to us.
|
|
Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has
|
|
power; neighbouring equipment is also dark.
|
|
|
|
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong
|
|
circumstantial evidence, not a measurement — and the instrument that would have
|
|
measured it (the power log on fv-ml1) died with the box.
|
|
|
|
## Blast radius
|
|
|
|
**19 of 30 LiteLLM aliases are dark** — fv-ml1 backs most of the fleet's inference:
|
|
|
|
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
|
|
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
|
|
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
|
|
summarizer-large
|
|
|
|
⚠⚠ **THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied
|
|
there was.** Probed 2026-09-13: every free local model on the gateway lives on
|
|
fv-ml1. irv-ml1 runs **no LLM chat seat at all** -- it carries TTS (tts-gateway,
|
|
breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and
|
|
waterland-studio, and its two Ampere cards are partly occupied by them. The only
|
|
non-fv chat backends on the gateway are **paid**: api.z.ai (9 aliases) and
|
|
Moonshot/Kimi (2).
|
|
|
|
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
|
|
vendor credits. ⚠ **If credits are spent, it must be under a NEW alias name that
|
|
callers opt into** -- never by silently repointing `summarizer`/`gen`/`classifier` at
|
|
GLM. Silent model substitution behind a familiar name is a standing prohibition here
|
|
and has already been violated twice; an outage is not an exemption.
|
|
|
|
The gateway itself on ana-docker is healthy — it is the backends that are gone.
|
|
|
|
## ⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
|
|
|
|
Every vLLM seat on fv-ml1 carries `restart: unless-stopped`. On boot **all of them
|
|
start loading simultaneously** — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
|
|
embed, rerank, reward, coder, scriberr — which is the single largest power transient
|
|
the box can produce, fed straight into a circuit that may have just tripped. That is a
|
|
re-trip, and a re-trip during model load can leave a half-written page cache and a
|
|
much longer recovery.
|
|
|
|
**Preferred sequence:**
|
|
|
|
1. Power the chassis on with **Docker masked**, so nothing auto-starts:
|
|
at the BMC/console, boot to the OS and before the network comes up run
|
|
`systemctl mask docker containerd` — or if the box is already up and loading,
|
|
`systemctl stop docker` immediately.
|
|
2. Confirm `nvidia-smi` sees all four cards and `zpool status tank` is ONLINE.
|
|
3. Unmask, then bring seats up **one at a time**, waiting for each to report healthy:
|
|
`gen` first (19 aliases depend on it), then embed/rerank/reward/coder, then
|
|
mog-sec, then the rest. `flash-next` LAST — it is the newest and least depended-on.
|
|
4. **Do not restart the MTP campaign.** It is the prime suspect.
|
|
5. Watch power while seats come up: `nvidia-smi --query-gpu=index,power.draw --format=csv`.
|
|
|
|
### ⭐ The staged bring-up also finishes the homepage-label fix, for free
|
|
|
|
The 2026-09-13 renumber (commit `3132a16`) repaired 25 compose files on this box but
|
|
the 10 RUNNING containers were never recreated, so their labels still carried the dead
|
|
10.250.50.54. Those containers are gone with the power loss.
|
|
|
|
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. `restart: unless-stopped` restarts the existing
|
|
container with its existing labels; labels only attach at container CREATION. But the
|
|
staged `docker compose up -d <svc>` sequence above **is** a recreate, and the compose
|
|
files on disk are already corrected — so bringing seats up that way applies the new
|
|
labels as a side effect and the dashboard comes back correct. Bring them up with
|
|
`compose up -d`, not by letting Docker restore the old containers.
|
|
|
|
Afterwards, confirm with:
|
|
|
|
curl -s http://10.0.50.45:5100/api/services | \
|
|
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
|
|
|
|
Expect `[]`. Before the outage that query returned 16 entries.
|
|
|
|
**Do not reboot the OPNsense firewall** (standing operator directive; its reboot API
|
|
403s anyway).
|
|
|
|
## ⚠ THE CIRCUIT CASE — what split power does and does not buy (operator, 2026-09-13)
|
|
|
|
> "unless of course the thing trips the circuit anyway."
|
|
|
|
**It still helps, but only halfway, and the halfway matters.**
|
|
|
|
- ✅ **A breaker trip is exactly what the split survives.** Firewall + BMC on the UPS is
|
|
~25-40 W of load on a 1500 VA unit — hours of battery, not minutes. On a trip the UPS
|
|
stops being a load-bearing supply and goes back to being what it is for.
|
|
- ❌ **A live firewall is useless if the path OUT of the site is dead.** Our UPS covers
|
|
our gear; it does not cover the **colo's handoff** — their switch, ONT or demarc. If
|
|
that sits on the circuit we just tripped, the result is a firewall running happily on
|
|
battery with nothing upstream to talk to, and the drive happens anyway.
|
|
⭐ **ASK THE FACILITY: is the network handoff on our circuit or theirs, and is theirs
|
|
on facility UPS?** This is the question that decides whether split power actually
|
|
delivers remote diagnosis or merely feels like it does.
|
|
|
|
### ⚠⚠ And the case where none of the above matters
|
|
|
|
**If four cards plus host exceeds the circuit, no UPS arrangement helps** — the box does
|
|
not fit its feed. Removing an undersized UPS does not remove the constraint, it promotes
|
|
the next one:
|
|
|
|
UPS ~900-1200 W (the one that just gave way)
|
|
circuit ~1800 W @ 15 A / ~2400 W @ 20 A
|
|
|
|
Which side of those the four-card figure lands on decides everything, which is why that
|
|
single ammeter reading is the load-bearing measurement of the visit.
|
|
|
|
### ⭐ The lever that may avoid an electrician: per-card power limits
|
|
|
|
`nvidia-smi -pl <watts>` caps TGP per card. The box can be made to fit whatever the feed
|
|
turns out to be, at a **throughput** cost rather than a **rewiring** cost — four capped
|
|
cards on a 15 A circuit is a dial we control today, where a 20 A drop is a ticket and a
|
|
site visit.
|
|
|
|
- Read `nvidia-smi -q -d POWER` first for the enforced min/max range per card; do not
|
|
assume how much room the dial has.
|
|
- ⚠ **If capping is the answer it MUST be persisted** (systemd unit, or an `if-up`
|
|
equivalent). A limit that evaporates on reboot is worse than no limit, because it will
|
|
hold right up until the next power event and then silently stop holding.
|
|
|
|
### Three questions for the site visit
|
|
|
|
1. What is the **breaker rating** on that circuit?
|
|
2. Is the circuit **dedicated** to us, or shared with other racks/tenants?
|
|
3. Is the **network handoff** on our circuit or the facility's, and is the facility's on
|
|
their UPS?
|
|
|
|
## ⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
|
|
|
|
**This is a recommendation awaiting the operator's call, not settled intent.** Written
|
|
down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring
|
|
the arrangement that just failed by default.
|
|
|
|
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
|
|
PDU / wall <- GPU chassis (no UPS in series)
|
|
|
|
Two reasons:
|
|
|
|
1. **It fixes the OOB gap this outage exposed.** The BMC's only route to the fleet is
|
|
through the firewall, so a power event at the GPU box takes out the management plane
|
|
with it -- which is precisely why this incident needs a drive rather than a console
|
|
session. Separate the two and a repeat leaves a live firewall, a live BMC, and
|
|
remote eyes on a dark chassis.
|
|
2. **A 1500 VA unit was never going to hold this box.** It has four cards, not the two
|
|
every record claimed until 2026-09-12.
|
|
|
|
⚠ **Do NOT use a UPS's surge-only outlets to get around its rating.** Both outlet banks
|
|
sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA
|
|
unit that is a single NEMA 5-15P rated **12 A at maximum load**, total across all
|
|
outlets. The surge bank bypasses the inverter, not the current rating. Overloading the
|
|
inverter trips or kills the unit; overloading the cord is a thermal problem in an
|
|
unattended rack. Bypass the UPS entirely instead.
|
|
|
|
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
|
|
probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
|
|
|
|
## Afterwards
|
|
|
|
- ⭐ **The OOB design gap this exposes.** The cutover chose OPNsense-as-subnet-router
|
|
specifically so the BMC stays reachable when the GPU box is down. That works for
|
|
box-down/gateway-up. It does **nothing** for a site-wide power or gateway loss —
|
|
exactly what happened — because the BMC's only path to the fleet is through that
|
|
gateway. A genuine OOB path at FV needs something the FV circuit cannot take down:
|
|
an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
|
|
- Get the actual circuit rating and the box's real peak draw, now that it is known to
|
|
have four cards and not two. Until then, treat concurrent multi-card load at FV as
|
|
unproven rather than safe.
|
|
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
|
|
up to the cut. Recover it after boot — it is the only measurement of what the load
|
|
actually drew, and it survives on `/tank`, not in the container.
|
|
|
|
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
|
|
|
|
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
|
|
> power limited to 200w"
|
|
>
|
|
> **CLARIFIED BY OPERATOR 2026-09-13 — the two boxes are different hardware:**
|
|
>
|
|
> | box | cards | TGP each | VRAM total | status |
|
|
> |---|---|---|---|---|
|
|
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
|
|
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
|
|
>
|
|
> The 200 W cap directive applies to **ana-ml3**. It very likely wants applying to
|
|
> fv-ml1 as well — see the four-card circuit arithmetic below.
|
|
|
|
The generalised lesson from this outage: **decide the power envelope first and size the
|
|
cards into it**, rather than installing cards and discovering the constraint by tripping
|
|
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
|
|
|
|
Three things to settle before that is a plan:
|
|
|
|
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
|
|
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
|
|
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
|
|
is scripted, refused quietly. **First command on the new hardware:**
|
|
|
|
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
|
|
|
|
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
|
|
feed, not from the cap.
|
|
|
|
2. ✅ **RESOLVED — both card types are 300 W.** So 200 W is a cap to **67% of TGP**, the
|
|
favourable part of the concave curve, not the severe 33% cap a 600 W part would have
|
|
implied. The enforceable-floor concern largely goes away too: 200 W was borderline
|
|
against a 600 W card's minimum and is very unlikely to sit below a 300 W card's. Worth
|
|
the one command; expect it to take.
|
|
|
|
**ana-ml3: 2 x 200 W = 400 W of card.** Modest, and pointed — ana-ml3 lands in the
|
|
**Anaheim** rack whose breaker tripped on 2026-08-26 and 2026-09-11, one of those
|
|
caused by this very chassis before it relocated. The cap there is remediation of a
|
|
circuit with a track record, not precaution.
|
|
|
|
### ⭐⭐ The outage arithmetic, now that the TGP is known
|
|
|
|
2 x Blackwell Max-Q @ 300 W ~ 600 W of card under load
|
|
host (board, 566 GB RAM, drives,
|
|
fans, PSU conversion loss) ~ 200-350 W <-- UNMEASURED, the gap
|
|
-----------
|
|
~ 800-950 W
|
|
|
|
Eaton 1500 VA real watt rating ~ 900-1200 W depending on model
|
|
|
|
**At or just over the line** — and this is what a vague "undersized" could not explain:
|
|
why it ran a full day on one card (~500-650 W, comfortably inside) and died minutes into
|
|
the second (~800-950 W, at or past the rating). The host term is the only one being
|
|
guessed at, and idle-at-the-plug with all seats down measures it directly.
|
|
|
|
### ⚠⚠ FOUR cards is a BREAKER problem, not a UPS problem
|
|
|
|
4 x 300 W card + ~300 W host ~ 1500 W
|
|
15 A circuit, 80% continuous = 1440 W
|
|
|
|
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all.** So
|
|
capping belongs at **fv-ml1 too**, not only ana-ml3. If the four-card ammeter reading
|
|
confirms it, fv-ml1 needs a per-card cap or a 20 A feed before anyone loads all four
|
|
again — and note that today's incident only ever had TWO cards working.
|
|
|
|
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
|
|
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
|
|
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
|
|
33% of TGP is deeper into the steep region — the cost is real and should be measured
|
|
on the first card rather than predicted, and it will hurt a prefill-heavy or training
|
|
workload considerably more than a serving seat.
|
|
|
|
### ⚠ Two placement consequences of Ada, independent of power
|
|
|
|
- **sm_89 has native FP8 but NOT NVFP4** (Blackwell-only, sm_100/sm_120). Most of our
|
|
in-house quants are NVFP4, so **they will not run accelerated on that colo's cards.**
|
|
Its seats want FP8 W8A8 builds, or the NVFP4 checkpoints stay on fv-ml1. Same class of
|
|
constraint as the Ampere finding for irv-ml1, one generation up.
|
|
- ⭐ **It unparks the triton-backend item.** That is a hard no on Ampere — crashes every
|
|
render on the A6000, `fp8e4nv` unsupported on sm_86 — and was explicitly deferred TO
|
|
Ada. sm_89 has the FP8 support it needs, so it becomes testable on this hardware.
|
|
- **VRAM:** 4 x 48 GB = 192 GB, against fv-ml1's 4 x 96 = 391 GB. Big-model placement
|
|
stays at FV. The Flash-Next seat needs 74 GiB resident on ONE card and would not fit a
|
|
48 GB Ada card even with the n-gram table offloaded — the offload moves the *table*,
|
|
not the experts.
|
|
|
|
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
|
|
stops holding — the worst possible failure shape, because the thing that reboots the box
|
|
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
|
|
mode, ordered before Docker starts.
|