Files
esh-pfi-infrastructure/docs/runbooks/fv-site-dark-20260913.md
T
vh 312725ddfb memory: snapshot — Flash-Next seat on one card, and the FV outage that followed
Durable capture so tomorrow's session does not have to reconstruct either half.

Built and verified before the power failed: Qwen3.8-Flash-Next serving on a single
RTX PRO 6000 with its 51B n-gram table pinned in host RAM and read over CUDA UVA --
74.36 GiB weights resident, 14.00 GiB KV for 560,654 tokens at the full 262,144
context, 67 GiB host RSS -- plus a gen-large gateway alias verified end to end.

The five findings worth carrying: the offload is #54371 (UVA, merged) which supersedes
the paused worker-based #53899 and designs out its entire bug family;
text_config.ple_embedding_dtype is the load-or-fail discriminator for any community
build; --kv-cache-memory makes vLLM SKIP memory profiling and ignore
gpu-memory-utilization, which inverts the usual pin-bytes advice and let a 16 GiB pin
nearly OOM with no visible failure; MTP is off pending measurement here rather than
written off, because the recipe's number is cross-harness and tested k=3 only while the
head is one layer run autoregressively; and a container once reported (healthy) with no
published port at all, because the healthcheck runs inside the boundary it was trusted
to validate.

Then the outage. Records it as will-not-self-recover, so no session wastes effort
polling a dead site, and carries the three things that change the visit: bypass the UPS
rather than using its surge-only bank (both banks share one 12 A inlet -- the surge
bank bypasses the inverter, not the current rating), recover power.log before anything
else because it is the only load measurement that exists anywhere, and bring seats up
one at a time because ten restart:unless-stopped containers loading at once is the
largest transient the box can make into whatever just failed.

Also records what is still half-done: the stale homepage labels on the 10 containers
that died before they could be recreated, which the staged bring-up fixes as a side
effect, and the eight drifted stacks plus three untracked host-only stacks that were
deliberately left for a deliberate reconciliation.
2026-09-13 00:14:46 -07:00

210 lines
12 KiB
Markdown

# FV site dark — 2026-09-13 ~06:56Z
**Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley.**
Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one
needs replacing, so there is no recovery to wait for and no point polling the site.
Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is
down so recovery does not have to be reconstructed from memory.
## What is down
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
| target | result |
|---|---|
| `fv-ml1` 10.251.50.54 | 100% loss |
| `fv-ml1` mesh 100.64.0.7 | 100% loss |
| FV gateway 10.251.50.1 | 100% loss |
| FV gateway mesh 100.64.0.8 | 100% loss |
| **fv-ml1-bmc 10.251.250.50** | **100% loss** |
| `fv.phasefinal.com` (172.83.89.66) | no answer |
**My own path is healthy** — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all
answer with 0% loss, so this is FV-local, not an isolation of the observer.
## Timeline
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
the two-card power risk. Production seat already live on GPU 2.
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
06:54:39 off_A arm healthy after 211 s
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
06:56:04 off_A rep 2 started <-- LAST LOG LINE
06:58:40 every FV address unreachable, including the BMC
So the site went dark inside a ~2.5 minute window, roughly five minutes into
sustained bench load on GPU 3 with GPU 2 also resident/serving.
## ⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
Operator's read, and it fits the evidence **better** than the breaker-trip theory
below, for a reason worth writing down:
**A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective
device to give.** A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit
carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM,
plus the firewall, plausibly clears the UPS rating while staying comfortably under the
breaker. That explains the thing a breaker trip explains poorly: **why it let go at
TWO cards loaded and not four.** It also explains the OPNsense box dying -- same UPS.
⚠ **"Died" is the operationally important word.** A tripped UPS resets; an overloaded
one can kill its output stage or battery pack permanently. If it is dead rather than
tripped, **nothing on site can be reset back to life** and the trip is wasted without
the means to bypass it.
### Bring / check list for the site visit
- **The means to BYPASS the UPS entirely** -- chassis and firewall straight to the
PDU or wall. Do this first, get the site back, diagnose the UPS after.
- **Read off the UPS before moving it:** make/model, VA and W rating, fault LEDs, and
any LCD or event log entry (many units record "overload" explicitly).
- **Note which outlets were battery-backed vs surge-only.** Mixed-bank units are
common, and a GPU chassis on the battery-backed bank overloads soonest.
- ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST, before any
bring-up.** It sampled all four cards every 10 s right up to the cut and is the
ONLY measurement of what the load actually drew. It lives on `/tank`, not in a
container, so it survived. Without it, the replacement UPS gets sized by guesswork.
- ⚠ **Size the replacement for FOUR cards under load, not two.** Anything sized to
today's failure just relocates the trip point to the next person who loads all four.
⚠ **No measured load figure exists yet.** Idle draw was measured (GPU0 3.80 W, GPU1
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
and is recoverable only from `power.log`. Do not let anyone put a wattage in a
purchasing decision until that file has been read.
## Three candidate causes, and what distinguishes them
Cannot be distinguished remotely, because every FV path — including the BMC —
traverses the OPNsense gateway, and the gateway is also dark.
1. **Circuit tripped under two-card load.** Fits the timing and is the predicted
failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this
same chassis, and the FV circuit was specced while every record still said the box
had **two** GPUs rather than four. ⭐ Distinguishing evidence: **the breaker is
visibly tripped**, and the OPNsense box is dark too (it draws ~20-30 W and would
survive anything short of a circuit/utility loss).
2. **OPNsense gateway crashed or rebooted**, taking all FV routing with it while the
GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and
the OPNsense box is the only dead thing.
3. **Upstream utility or colo-side power/network loss**, unrelated to us.
Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has
power; neighbouring equipment is also dark.
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong
circumstantial evidence, not a measurement — and the instrument that would have
measured it (the power log on fv-ml1) died with the box.
## Blast radius
**19 of 30 LiteLLM aliases are dark** — fv-ml1 backs most of the fleet's inference:
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
summarizer-large
⚠⚠ **THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied
there was.** Probed 2026-09-13: every free local model on the gateway lives on
fv-ml1. irv-ml1 runs **no LLM chat seat at all** -- it carries TTS (tts-gateway,
breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and
waterland-studio, and its two Ampere cards are partly occupied by them. The only
non-fv chat backends on the gateway are **paid**: api.z.ai (9 aliases) and
Moonshot/Kimi (2).
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
vendor credits. ⚠ **If credits are spent, it must be under a NEW alias name that
callers opt into** -- never by silently repointing `summarizer`/`gen`/`classifier` at
GLM. Silent model substitution behind a familiar name is a standing prohibition here
and has already been violated twice; an outage is not an exemption.
The gateway itself on ana-docker is healthy — it is the backends that are gone.
## ⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
Every vLLM seat on fv-ml1 carries `restart: unless-stopped`. On boot **all of them
start loading simultaneously** — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
embed, rerank, reward, coder, scriberr — which is the single largest power transient
the box can produce, fed straight into a circuit that may have just tripped. That is a
re-trip, and a re-trip during model load can leave a half-written page cache and a
much longer recovery.
**Preferred sequence:**
1. Power the chassis on with **Docker masked**, so nothing auto-starts:
at the BMC/console, boot to the OS and before the network comes up run
`systemctl mask docker containerd` — or if the box is already up and loading,
`systemctl stop docker` immediately.
2. Confirm `nvidia-smi` sees all four cards and `zpool status tank` is ONLINE.
3. Unmask, then bring seats up **one at a time**, waiting for each to report healthy:
`gen` first (19 aliases depend on it), then embed/rerank/reward/coder, then
mog-sec, then the rest. `flash-next` LAST — it is the newest and least depended-on.
4. **Do not restart the MTP campaign.** It is the prime suspect.
5. Watch power while seats come up: `nvidia-smi --query-gpu=index,power.draw --format=csv`.
### ⭐ The staged bring-up also finishes the homepage-label fix, for free
The 2026-09-13 renumber (commit `3132a16`) repaired 25 compose files on this box but
the 10 RUNNING containers were never recreated, so their labels still carried the dead
10.250.50.54. Those containers are gone with the power loss.
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. `restart: unless-stopped` restarts the existing
container with its existing labels; labels only attach at container CREATION. But the
staged `docker compose up -d <svc>` sequence above **is** a recreate, and the compose
files on disk are already corrected — so bringing seats up that way applies the new
labels as a side effect and the dashboard comes back correct. Bring them up with
`compose up -d`, not by letting Docker restore the old containers.
Afterwards, confirm with:
curl -s http://10.0.50.45:5100/api/services | \
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
Expect `[]`. Before the outage that query returned 16 entries.
**Do not reboot the OPNsense firewall** (standing operator directive; its reboot API
403s anyway).
## ⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
**This is a recommendation awaiting the operator's call, not settled intent.** Written
down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring
the arrangement that just failed by default.
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
PDU / wall <- GPU chassis (no UPS in series)
Two reasons:
1. **It fixes the OOB gap this outage exposed.** The BMC's only route to the fleet is
through the firewall, so a power event at the GPU box takes out the management plane
with it -- which is precisely why this incident needs a drive rather than a console
session. Separate the two and a repeat leaves a live firewall, a live BMC, and
remote eyes on a dark chassis.
2. **A 1500 VA unit was never going to hold this box.** It has four cards, not the two
every record claimed until 2026-09-12.
⚠ **Do NOT use a UPS's surge-only outlets to get around its rating.** Both outlet banks
sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA
unit that is a single NEMA 5-15P rated **12 A at maximum load**, total across all
outlets. The surge bank bypasses the inverter, not the current rating. Overloading the
inverter trips or kills the unit; overloading the cord is a thermal problem in an
unattended rack. Bypass the UPS entirely instead.
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
## Afterwards
- ⭐ **The OOB design gap this exposes.** The cutover chose OPNsense-as-subnet-router
specifically so the BMC stays reachable when the GPU box is down. That works for
box-down/gateway-up. It does **nothing** for a site-wide power or gateway loss —
exactly what happened — because the BMC's only path to the fleet is through that
gateway. A genuine OOB path at FV needs something the FV circuit cannot take down:
an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
- Get the actual circuit rating and the box's real peak draw, now that it is known to
have four cards and not two. Until then, treat concurrent multi-card load at FV as
unproven rather than safe.
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
up to the cut. Recover it after boot — it is the only measurement of what the load
actually drew, and it survives on `/tank`, not in the container.