Durable capture so tomorrow's session does not have to reconstruct either half. Built and verified before the power failed: Qwen3.8-Flash-Next serving on a single RTX PRO 6000 with its 51B n-gram table pinned in host RAM and read over CUDA UVA -- 74.36 GiB weights resident, 14.00 GiB KV for 560,654 tokens at the full 262,144 context, 67 GiB host RSS -- plus a gen-large gateway alias verified end to end. The five findings worth carrying: the offload is #54371 (UVA, merged) which supersedes the paused worker-based #53899 and designs out its entire bug family; text_config.ple_embedding_dtype is the load-or-fail discriminator for any community build; --kv-cache-memory makes vLLM SKIP memory profiling and ignore gpu-memory-utilization, which inverts the usual pin-bytes advice and let a 16 GiB pin nearly OOM with no visible failure; MTP is off pending measurement here rather than written off, because the recipe's number is cross-harness and tested k=3 only while the head is one layer run autoregressively; and a container once reported (healthy) with no published port at all, because the healthcheck runs inside the boundary it was trusted to validate. Then the outage. Records it as will-not-self-recover, so no session wastes effort polling a dead site, and carries the three things that change the visit: bypass the UPS rather than using its surge-only bank (both banks share one 12 A inlet -- the surge bank bypasses the inverter, not the current rating), recover power.log before anything else because it is the only load measurement that exists anywhere, and bring seats up one at a time because ten restart:unless-stopped containers loading at once is the largest transient the box can make into whatever just failed. Also records what is still half-done: the stale homepage labels on the 10 containers that died before they could be recreated, which the staged bring-up fixes as a side effect, and the eight drifted stacks plus three untracked host-only stacks that were deliberately left for a deliberate reconciliation.
12 KiB
FV site dark — 2026-09-13 ~06:56Z
Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley. Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one needs replacing, so there is no recovery to wait for and no point polling the site. Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is down so recovery does not have to be reconstructed from memory.
What is down
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
| target | result |
|---|---|
fv-ml1 10.251.50.54 |
100% loss |
fv-ml1 mesh 100.64.0.7 |
100% loss |
| FV gateway 10.251.50.1 | 100% loss |
| FV gateway mesh 100.64.0.8 | 100% loss |
| fv-ml1-bmc 10.251.250.50 | 100% loss |
fv.phasefinal.com (172.83.89.66) |
no answer |
My own path is healthy — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all answer with 0% loss, so this is FV-local, not an isolation of the observer.
Timeline
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
the two-card power risk. Production seat already live on GPU 2.
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
06:54:39 off_A arm healthy after 211 s
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
06:56:04 off_A rep 2 started <-- LAST LOG LINE
06:58:40 every FV address unreachable, including the BMC
So the site went dark inside a ~2.5 minute window, roughly five minutes into sustained bench load on GPU 3 with GPU 2 also resident/serving.
⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
Operator's read, and it fits the evidence better than the breaker-trip theory below, for a reason worth writing down:
A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective device to give. A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM, plus the firewall, plausibly clears the UPS rating while staying comfortably under the breaker. That explains the thing a breaker trip explains poorly: why it let go at TWO cards loaded and not four. It also explains the OPNsense box dying -- same UPS.
⚠ "Died" is the operationally important word. A tripped UPS resets; an overloaded one can kill its output stage or battery pack permanently. If it is dead rather than tripped, nothing on site can be reset back to life and the trip is wasted without the means to bypass it.
Bring / check list for the site visit
- The means to BYPASS the UPS entirely -- chassis and firewall straight to the PDU or wall. Do this first, get the site back, diagnose the UPS after.
- Read off the UPS before moving it: make/model, VA and W rating, fault LEDs, and any LCD or event log entry (many units record "overload" explicitly).
- Note which outlets were battery-backed vs surge-only. Mixed-bank units are common, and a GPU chassis on the battery-backed bank overloads soonest.
- ⭐ Recover
/tank/aimodels/flash-next-mtp-bench/power.logFIRST, before any bring-up. It sampled all four cards every 10 s right up to the cut and is the ONLY measurement of what the load actually drew. It lives on/tank, not in a container, so it survived. Without it, the replacement UPS gets sized by guesswork. - ⚠ Size the replacement for FOUR cards under load, not two. Anything sized to today's failure just relocates the trip point to the next person who loads all four.
⚠ No measured load figure exists yet. Idle draw was measured (GPU0 3.80 W, GPU1
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
and is recoverable only from power.log. Do not let anyone put a wattage in a
purchasing decision until that file has been read.
Three candidate causes, and what distinguishes them
Cannot be distinguished remotely, because every FV path — including the BMC — traverses the OPNsense gateway, and the gateway is also dark.
- Circuit tripped under two-card load. Fits the timing and is the predicted failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this same chassis, and the FV circuit was specced while every record still said the box had two GPUs rather than four. ⭐ Distinguishing evidence: the breaker is visibly tripped, and the OPNsense box is dark too (it draws ~20-30 W and would survive anything short of a circuit/utility loss).
- OPNsense gateway crashed or rebooted, taking all FV routing with it while the GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and the OPNsense box is the only dead thing.
- Upstream utility or colo-side power/network loss, unrelated to us. Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has power; neighbouring equipment is also dark.
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong circumstantial evidence, not a measurement — and the instrument that would have measured it (the power log on fv-ml1) died with the box.
Blast radius
19 of 30 LiteLLM aliases are dark — fv-ml1 backs most of the fleet's inference:
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
summarizer-large
⚠⚠ THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied there was. Probed 2026-09-13: every free local model on the gateway lives on fv-ml1. irv-ml1 runs no LLM chat seat at all -- it carries TTS (tts-gateway, breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and waterland-studio, and its two Ampere cards are partly occupied by them. The only non-fv chat backends on the gateway are paid: api.z.ai (9 aliases) and Moonshot/Kimi (2).
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
vendor credits. ⚠ If credits are spent, it must be under a NEW alias name that
callers opt into -- never by silently repointing summarizer/gen/classifier at
GLM. Silent model substitution behind a familiar name is a standing prohibition here
and has already been violated twice; an outage is not an exemption.
The gateway itself on ana-docker is healthy — it is the backends that are gone.
⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
Every vLLM seat on fv-ml1 carries restart: unless-stopped. On boot all of them
start loading simultaneously — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
embed, rerank, reward, coder, scriberr — which is the single largest power transient
the box can produce, fed straight into a circuit that may have just tripped. That is a
re-trip, and a re-trip during model load can leave a half-written page cache and a
much longer recovery.
Preferred sequence:
- Power the chassis on with Docker masked, so nothing auto-starts:
at the BMC/console, boot to the OS and before the network comes up run
systemctl mask docker containerd— or if the box is already up and loading,systemctl stop dockerimmediately. - Confirm
nvidia-smisees all four cards andzpool status tankis ONLINE. - Unmask, then bring seats up one at a time, waiting for each to report healthy:
genfirst (19 aliases depend on it), then embed/rerank/reward/coder, then mog-sec, then the rest.flash-nextLAST — it is the newest and least depended-on. - Do not restart the MTP campaign. It is the prime suspect.
- Watch power while seats come up:
nvidia-smi --query-gpu=index,power.draw --format=csv.
⭐ The staged bring-up also finishes the homepage-label fix, for free
The 2026-09-13 renumber (commit 3132a16) repaired 25 compose files on this box but
the 10 RUNNING containers were never recreated, so their labels still carried the dead
10.250.50.54. Those containers are gone with the power loss.
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. restart: unless-stopped restarts the existing
container with its existing labels; labels only attach at container CREATION. But the
staged docker compose up -d <svc> sequence above is a recreate, and the compose
files on disk are already corrected — so bringing seats up that way applies the new
labels as a side effect and the dashboard comes back correct. Bring them up with
compose up -d, not by letting Docker restore the old containers.
Afterwards, confirm with:
curl -s http://10.0.50.45:5100/api/services | \
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
Expect []. Before the outage that query returned 16 entries.
Do not reboot the OPNsense firewall (standing operator directive; its reboot API 403s anyway).
⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
This is a recommendation awaiting the operator's call, not settled intent. Written down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring the arrangement that just failed by default.
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
PDU / wall <- GPU chassis (no UPS in series)
Two reasons:
- It fixes the OOB gap this outage exposed. The BMC's only route to the fleet is through the firewall, so a power event at the GPU box takes out the management plane with it -- which is precisely why this incident needs a drive rather than a console session. Separate the two and a repeat leaves a live firewall, a live BMC, and remote eyes on a dark chassis.
- A 1500 VA unit was never going to hold this box. It has four cards, not the two every record claimed until 2026-09-12.
⚠ Do NOT use a UPS's surge-only outlets to get around its rating. Both outlet banks sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA unit that is a single NEMA 5-15P rated 12 A at maximum load, total across all outlets. The surge bank bypasses the inverter, not the current rating. Overloading the inverter trips or kills the unit; overloading the cord is a thermal problem in an unattended rack. Bypass the UPS entirely instead.
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
probably a 20 A circuit -- but size it from power.log, not from a spec sheet.
Afterwards
- ⭐ The OOB design gap this exposes. The cutover chose OPNsense-as-subnet-router specifically so the BMC stays reachable when the GPU box is down. That works for box-down/gateway-up. It does nothing for a site-wide power or gateway loss — exactly what happened — because the BMC's only path to the fleet is through that gateway. A genuine OOB path at FV needs something the FV circuit cannot take down: an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
- Get the actual circuit rating and the box's real peak draw, now that it is known to have four cards and not two. Until then, treat concurrent multi-card load at FV as unproven rather than safe.
services/flash-next-mtp-bench/power.logon the box holds the per-card draw right up to the cut. Recover it after boot — it is the only measurement of what the load actually drew, and it survives on/tank, not in the container.