Operator's reasoning, accepted and better than the hypothesis-space argument it replaces: the NAT change went effective, was verified bidirectional, and then ran correctly for twenty minutes before the site died the moment GPU load was applied. A working config change does not spontaneously fail under an unrelated physical variable. The load correlation is tight; the NAT correlation is merely adjacent in time. Undersized UPS is the only candidate that explains the trigger. NAT material retained as record, and the power.log/uptime check demoted from decision point to free confirmation. Adds the measurement protocol, since the operator is bringing a PDU and an ammeter. The load-bearing caveat: power.log is GPU-ONLY -- nvidia-smi per-card, excluding CPU, 566 GB of RAM, drives, fans and PSU conversion losses -- so the ammeter at the plug is the primary instrument and power.log only cross-checks the GPU share. Four states to capture (idle, one card, two cards, four cards), and capture PEAK rather than average: UPS overload protection responds to short-term overload, so an average-only reading that hides transients will mis-size the replacement exactly the way the present unit got mis-sized, and must be recorded as a floor rather than as the draw. The four-card figure is earmarked for servers/fv-ml1/README.md, because it closes the cutover's own open question -- that the FV circuit was likely specced against half the real draw, back when every record still said the box had two GPUs.
16 KiB
FV site dark — 2026-09-13 ~06:56Z
Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley. Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one needs replacing, so there is no recovery to wait for and no point polling the site. Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is down so recovery does not have to be reconstructed from memory.
What is down
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
| target | result |
|---|---|
fv-ml1 10.251.50.54 |
100% loss |
fv-ml1 mesh 100.64.0.7 |
100% loss |
| FV gateway 10.251.50.1 | 100% loss |
| FV gateway mesh 100.64.0.8 | 100% loss |
| fv-ml1-bmc 10.251.250.50 | 100% loss |
fv.phasefinal.com (172.83.89.66) |
no answer |
My own path is healthy — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all answer with 0% loss, so this is FV-local, not an isolation of the observer.
Timeline
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
the two-card power risk. Production seat already live on GPU 2.
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
06:54:39 off_A arm healthy after 211 s
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
06:56:04 off_A rep 2 started <-- LAST LOG LINE
06:58:40 every FV address unreachable, including the BMC
So the site went dark inside a ~2.5 minute window, roughly five minutes into sustained bench load on GPU 3 with GPU 2 also resident/serving.
⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
Operator's read, and it fits the evidence better than the breaker-trip theory below, for a reason worth writing down:
A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective device to give. A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM, plus the firewall, plausibly clears the UPS rating while staying comfortably under the breaker. That explains the thing a breaker trip explains poorly: why it let go at TWO cards loaded and not four. It also explains the OPNsense box dying -- same UPS.
⚠ "Died" is the operationally important word. A tripped UPS resets; an overloaded one can kill its output stage or battery pack permanently. If it is dead rather than tripped, nothing on site can be reset back to life and the trip is wasted without the means to bypass it.
Bring / check list for the site visit
- The means to BYPASS the UPS entirely -- chassis and firewall straight to the PDU or wall. Do this first, get the site back, diagnose the UPS after.
- Read off the UPS before moving it: make/model, VA and W rating, fault LEDs, and any LCD or event log entry (many units record "overload" explicitly).
- Note which outlets were battery-backed vs surge-only. Mixed-bank units are common, and a GPU chassis on the battery-backed bank overloads soonest.
- ⭐ Recover
/tank/aimodels/flash-next-mtp-bench/power.logFIRST, before any bring-up. It sampled all four cards every 10 s right up to the cut and is the ONLY measurement of what the load actually drew. It lives on/tank, not in a container, so it survived. Without it, the replacement UPS gets sized by guesswork. - ⚠ Size the replacement for FOUR cards under load, not two. Anything sized to today's failure just relocates the trip point to the next person who loads all four.
⚠ No measured load figure exists yet. Idle draw was measured (GPU0 3.80 W, GPU1
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
and is recoverable only from power.log. Do not let anyone put a wattage in a
purchasing decision until that file has been read.
⭐ OPERATOR RULING 2026-09-13: undersized UPS. NAT hypothesis DEMOTED.
"highly doubt the nat hypothesis, it went effective (connectivity confirmed for previously dark path) -- then 20 minutes later, during load, site went dark. that's pretty unlikely to be the cause. for sure I think the ups was undersized."
Accepted, and the reasoning is better than the hypothesis-space argument below: a config change that went effective, was verified bidirectional, and then ran correctly for twenty minutes does not spontaneously fail when an unrelated physical variable -- someone else's GPU load -- is introduced. The load correlation is tight; the NAT correlation is merely adjacent in time. Undersized UPS is the only candidate that explains the trigger.
The NAT material below is retained as record, not as a live competing theory, and the
power.log / uptime check is retained as free confirmation rather than as a
decision point.
Measurement plan for the site visit (operator bringing a PDU + ammeter)
⚠ power.log is GPU-ONLY. It samples nvidia-smi per-card draw and does not
include the host: CPU, 566 GB of RAM, drives, fans, or PSU conversion losses. The
number that matters against a UPS rating is the whole chassis at the plug. The
ammeter is therefore the primary instrument and power.log is a cross-check on the
GPU share.
Capture four states -- this is the first real sizing data that has ever existed for this box:
| state | why it matters |
|---|---|
| all seats down, idle | the floor (GPU idle measured 3.80 / 3.88 / 14.27 / 6.75 W) |
| one card loaded | the condition that ran fine for a day |
| two cards loaded | the condition that took the site down |
| four cards loaded | the only number that can size a replacement honestly |
⚠⚠ CAPTURE PEAK, NOT AVERAGE. GPU power has fast transients and UPS overload protection responds to short-term overload, so a 1-second-average reading can under-read peaks badly. Use max-hold/peak capture if the meter has it. A figure like "1100 W average" that hides 1600 W spikes will mis-size the replacement the same way the current unit got mis-sized. If the meter is average-only, record the number as a FLOOR, not as the draw.
⭐ The four-card figure goes into servers/fv-ml1/README.md permanently. The
cutover notes flagged that the FV circuit was "likely specced against half the real
draw" while every record still said two GPUs; this closes that with a measurement
instead of an assumption.
The NAT change, retained as record (DEMOTED — see the ruling above)
Another session applied a Tailscale SNAT rule to the FV gateway at ~06:22Z — 34
minutes before the site went dark. See docs/runbooks/fv-to-ana-nat.md (uncommitted
as of this writing; not my work, left alone). So "we overloaded the power" is no longer
the only live hypothesis, and the UPS should not be replaced on the strength of a theory
until the one below is run.
On the evidence, that change is the WRONG SHAPE to have caused this, and I want that on the record so nobody wastes the visit chasing it:
- It is one outbound SNAT rule, source-scoped to
10.251.50.54/32, destination- scoped to10.250.0.0/16. An outbound NAT rule cannot stop the gateway, the BMC or the public WAN address from answering inbound. - The runbook states no routes, filter rules, WAN settings or subnet advertisements were
touched, and that
pfctl -srcame back byte-identical. - It was verified working in both directions afterwards: FV→hub HTTP 200, FV→ANA TCP 5432, FV internet HTTPS 200, ANA→FV SSH reachable, gateway management intact, Beszel 18/18 up.
⚠ Note their BMC observation used 10.251.50.50, which is not the BMC — the BMC is
10.251.250.50, a different subnet. They correctly declined to claim BMC health, but
the datapoint is void, not negative. Do not reason from it either way.
⭐⭐ THE DISCRIMINATOR — run this before forming any conclusion
power.log is written locally on /tank, every 10 seconds, by a shell loop on the
box. It does not depend on the network. So:
power.log last entry |
what it means |
|---|---|
| past 06:56Z | the box never lost power. This is a routing/gateway fault, and the UPS is innocent. |
| stops at ~06:56Z | the box lost power. UPS/circuit confirmed. |
Cross-check with uptime and journalctl --list-boots the moment there is a console:
continuous uptime across 06:56Z kills the UPS theory outright.
⚠ So the FIRST action on site is to read, not to fix. uptime,
journalctl --list-boots, then tail power.log. Establish whether the machine ever
went down before anyone buys hardware or flips anything — the two hypotheses lead to
completely different remediations and only one of them needs a new UPS.
Three further candidate causes, and what distinguishes them
Cannot be distinguished remotely, because every FV path — including the BMC — traverses the OPNsense gateway, and the gateway is also dark.
- Circuit tripped under two-card load. Fits the timing and is the predicted failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this same chassis, and the FV circuit was specced while every record still said the box had two GPUs rather than four. ⭐ Distinguishing evidence: the breaker is visibly tripped, and the OPNsense box is dark too (it draws ~20-30 W and would survive anything short of a circuit/utility loss).
- OPNsense gateway crashed or rebooted, taking all FV routing with it while the GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and the OPNsense box is the only dead thing.
- Upstream utility or colo-side power/network loss, unrelated to us. Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has power; neighbouring equipment is also dark.
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong circumstantial evidence, not a measurement — and the instrument that would have measured it (the power log on fv-ml1) died with the box.
Blast radius
19 of 30 LiteLLM aliases are dark — fv-ml1 backs most of the fleet's inference:
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
summarizer-large
⚠⚠ THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied there was. Probed 2026-09-13: every free local model on the gateway lives on fv-ml1. irv-ml1 runs no LLM chat seat at all -- it carries TTS (tts-gateway, breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and waterland-studio, and its two Ampere cards are partly occupied by them. The only non-fv chat backends on the gateway are paid: api.z.ai (9 aliases) and Moonshot/Kimi (2).
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
vendor credits. ⚠ If credits are spent, it must be under a NEW alias name that
callers opt into -- never by silently repointing summarizer/gen/classifier at
GLM. Silent model substitution behind a familiar name is a standing prohibition here
and has already been violated twice; an outage is not an exemption.
The gateway itself on ana-docker is healthy — it is the backends that are gone.
⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
Every vLLM seat on fv-ml1 carries restart: unless-stopped. On boot all of them
start loading simultaneously — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
embed, rerank, reward, coder, scriberr — which is the single largest power transient
the box can produce, fed straight into a circuit that may have just tripped. That is a
re-trip, and a re-trip during model load can leave a half-written page cache and a
much longer recovery.
Preferred sequence:
- Power the chassis on with Docker masked, so nothing auto-starts:
at the BMC/console, boot to the OS and before the network comes up run
systemctl mask docker containerd— or if the box is already up and loading,systemctl stop dockerimmediately. - Confirm
nvidia-smisees all four cards andzpool status tankis ONLINE. - Unmask, then bring seats up one at a time, waiting for each to report healthy:
genfirst (19 aliases depend on it), then embed/rerank/reward/coder, then mog-sec, then the rest.flash-nextLAST — it is the newest and least depended-on. - Do not restart the MTP campaign. It is the prime suspect.
- Watch power while seats come up:
nvidia-smi --query-gpu=index,power.draw --format=csv.
⭐ The staged bring-up also finishes the homepage-label fix, for free
The 2026-09-13 renumber (commit 3132a16) repaired 25 compose files on this box but
the 10 RUNNING containers were never recreated, so their labels still carried the dead
10.250.50.54. Those containers are gone with the power loss.
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. restart: unless-stopped restarts the existing
container with its existing labels; labels only attach at container CREATION. But the
staged docker compose up -d <svc> sequence above is a recreate, and the compose
files on disk are already corrected — so bringing seats up that way applies the new
labels as a side effect and the dashboard comes back correct. Bring them up with
compose up -d, not by letting Docker restore the old containers.
Afterwards, confirm with:
curl -s http://10.0.50.45:5100/api/services | \
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
Expect []. Before the outage that query returned 16 entries.
Do not reboot the OPNsense firewall (standing operator directive; its reboot API 403s anyway).
⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
This is a recommendation awaiting the operator's call, not settled intent. Written down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring the arrangement that just failed by default.
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
PDU / wall <- GPU chassis (no UPS in series)
Two reasons:
- It fixes the OOB gap this outage exposed. The BMC's only route to the fleet is through the firewall, so a power event at the GPU box takes out the management plane with it -- which is precisely why this incident needs a drive rather than a console session. Separate the two and a repeat leaves a live firewall, a live BMC, and remote eyes on a dark chassis.
- A 1500 VA unit was never going to hold this box. It has four cards, not the two every record claimed until 2026-09-12.
⚠ Do NOT use a UPS's surge-only outlets to get around its rating. Both outlet banks sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA unit that is a single NEMA 5-15P rated 12 A at maximum load, total across all outlets. The surge bank bypasses the inverter, not the current rating. Overloading the inverter trips or kills the unit; overloading the cord is a thermal problem in an unattended rack. Bypass the UPS entirely instead.
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
probably a 20 A circuit -- but size it from power.log, not from a spec sheet.
Afterwards
- ⭐ The OOB design gap this exposes. The cutover chose OPNsense-as-subnet-router specifically so the BMC stays reachable when the GPU box is down. That works for box-down/gateway-up. It does nothing for a site-wide power or gateway loss — exactly what happened — because the BMC's only path to the fleet is through that gateway. A genuine OOB path at FV needs something the FV circuit cannot take down: an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
- Get the actual circuit rating and the box's real peak draw, now that it is known to have four cards and not two. Until then, treat concurrent multi-card load at FV as unproven rather than safe.
services/flash-next-mtp-bench/power.logon the box holds the per-card draw right up to the cut. Recover it after boot — it is the only measurement of what the load actually drew, and it survives on/tank, not in the container.