Files
esh-pfi-infrastructure/docs/runbooks/fv-site-dark-20260913.md
T
vh 2da0c76d99 correct the hardware: fv-ml1 is 4x Blackwell Max-Q 300W, ana-ml3 is 2x Ada RTX 6000 — and four cards is a breaker problem
Operator clarification, and it separates two boxes I had been conflating. fv-ml1 is
4x Blackwell RTX PRO 6000 Max-Q at 300 W each (Max-Q being the reduced-TGP SKU; the
Workstation Edition is the 600 W part), 391 GB VRAM, deployed and currently dark.
ana-ml3 is 2x Ada Generation RTX 6000 at 300 W, 96 GB VRAM, not yet deployed. The 200 W
cap directive is ana-ml3's.

With the TGP known, the outage stops being a vague 'undersized' and acquires a
mechanism: two Max-Q cards at 300 W is ~600 W of card, plus a host carrying 566 GB of
RAM, drives, fans and PSU conversion loss at perhaps 200-350 W, against an Eaton 1500
VA's real ~900-1200 W. That lands at or just over the rating, which is precisely what
explains a full day of service on one card and failure minutes into the second. The host
term is the only one being guessed; idle-at-the-plug measures it directly.

It also surfaces something that is not a UPS question at all. Four cards at 300 W plus
~300 W of host is ~1500 W against a 15 A circuit's 1440 W continuous derating, so four
cards uncapped is marginal on the breaker with no UPS in the path. Capping therefore
belongs at fv-ml1 as well as ana-ml3, or fv-ml1 needs a 20 A feed -- and worth noting
today's incident only ever had two of the four cards working.

ana-ml3's placement constraints sharpen too: sm_89 has native FP8 but no NVFP4, so the
in-house NVFP4 quants stay at FV, and at 96 GB total it cannot host the Flash-Next seat
at all -- that needs 74 GiB resident on a single card, and the offload moves the n-gram
table rather than the experts.
2026-09-13 00:26:29 -07:00

441 lines
24 KiB
Markdown

# FV site dark — 2026-09-13 ~06:56Z
**Status: UNRESOLVED, WILL NOT SELF-RECOVER, needs hands at Fountain Valley.**
Operator 2026-09-13: an overload-tripped UPS does not clear itself and a dead one
needs replacing, so there is no recovery to wait for and no point polling the site.
Visit planned for 2026-09-14. Do not leave watchers running against FV addresses. Written while the site is
down so recovery does not have to be reconstructed from memory.
## What is down
Everything at Fountain Valley, measured 06:58:40Z from nh3-dev:
| target | result |
|---|---|
| `fv-ml1` 10.251.50.54 | 100% loss |
| `fv-ml1` mesh 100.64.0.7 | 100% loss |
| FV gateway 10.251.50.1 | 100% loss |
| FV gateway mesh 100.64.0.8 | 100% loss |
| **fv-ml1-bmc 10.251.250.50** | **100% loss** |
| `fv.phasefinal.com` (172.83.89.66) | no answer |
**My own path is healthy** — ana-docker, nh3-docker, esh-docker-vm and 1.1.1.1 all
answer with 0% loss, so this is FV-local, not an isolation of the observer.
## Timeline
06:51:08 MTP campaign started on GPU 3 (:8023), operator-authorised, knowing
the two-card power risk. Production seat already live on GPU 2.
06:51:17 power sampler baseline: GPU0 3.80 W, GPU1 3.88 W, GPU2 14.27 W, GPU3 6.75 W
06:54:39 off_A arm healthy after 211 s
06:54:39 off_A rep 1 ran clean: 75.49 / 212.33 / 387.83 tok/s at conc 1/4/8
06:56:04 off_A rep 2 started <-- LAST LOG LINE
06:58:40 every FV address unreachable, including the BMC
So the site went dark inside a ~2.5 minute window, roughly five minutes into
sustained bench load on GPU 3 with GPU 2 also resident/serving.
## ⭐ LEADING HYPOTHESIS (operator, 2026-09-13): the UPS overloaded and died
Operator's read, and it fits the evidence **better** than the breaker-trip theory
below, for a reason worth writing down:
**A UPS's output rating is far below the circuit's, so the UPS is the FIRST protective
device to give.** A common 1500 VA rack unit delivers ~900-1000 W; a 15 A circuit
carries ~1800 W. Two of these cards under load, plus a host carrying 566 GB of RAM,
plus the firewall, plausibly clears the UPS rating while staying comfortably under the
breaker. That explains the thing a breaker trip explains poorly: **why it let go at
TWO cards loaded and not four.** It also explains the OPNsense box dying -- same UPS.
⚠ **"Died" is the operationally important word.** A tripped UPS resets; an overloaded
one can kill its output stage or battery pack permanently. If it is dead rather than
tripped, **nothing on site can be reset back to life** and the trip is wasted without
the means to bypass it.
### Bring / check list for the site visit
- **The means to BYPASS the UPS entirely** -- chassis and firewall straight to the
PDU or wall. Do this first, get the site back, diagnose the UPS after.
- **Read off the UPS before moving it:** make/model, VA and W rating, fault LEDs, and
any LCD or event log entry (many units record "overload" explicitly).
- **Note which outlets were battery-backed vs surge-only.** Mixed-bank units are
common, and a GPU chassis on the battery-backed bank overloads soonest.
- ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST, before any
bring-up.** It sampled all four cards every 10 s right up to the cut and is the
ONLY measurement of what the load actually drew. It lives on `/tank`, not in a
container, so it survived. Without it, the replacement UPS gets sized by guesswork.
- ⚠ **Size the replacement for FOUR cards under load, not two.** Anything sized to
today's failure just relocates the trip point to the next person who loads all four.
⚠ **No measured load figure exists yet.** Idle draw was measured (GPU0 3.80 W, GPU1
3.88 W, GPU2 14.27 W, GPU3 6.75 W at 06:51:17Z) but the load figure died with the box
and is recoverable only from `power.log`. Do not let anyone put a wattage in a
purchasing decision until that file has been read.
## ⭐ OPERATOR RULING 2026-09-13: undersized UPS. NAT hypothesis DEMOTED.
> "highly doubt the nat hypothesis, it went effective (connectivity confirmed for
> previously dark path) -- then 20 minutes later, during load, site went dark. that's
> pretty unlikely to be the cause. for sure I think the ups was undersized."
Accepted, and the reasoning is better than the hypothesis-space argument below: **a
config change that went effective, was verified bidirectional, and then ran correctly
for twenty minutes does not spontaneously fail when an unrelated physical variable --
someone else's GPU load -- is introduced.** The load correlation is tight; the NAT
correlation is merely adjacent in time. Undersized UPS is the only candidate that
explains the *trigger*.
The NAT material below is retained as record, not as a live competing theory, and the
`power.log` / `uptime` check is retained as **free confirmation** rather than as a
decision point.
## Measurement plan for the site visit (operator bringing a PDU + ammeter)
⚠ **`power.log` is GPU-ONLY.** It samples `nvidia-smi` per-card draw and does **not**
include the host: CPU, 566 GB of RAM, drives, fans, or PSU conversion losses. The
number that matters against a UPS rating is the whole chassis **at the plug**. The
ammeter is therefore the primary instrument and `power.log` is a cross-check on the
GPU share.
Capture four states -- this is the first real sizing data that has ever existed for
this box:
| state | why it matters |
|---|---|
| all seats down, idle | the floor (GPU idle measured 3.80 / 3.88 / 14.27 / 6.75 W) |
| one card loaded | the condition that ran fine for a day |
| **two cards loaded** | the condition that took the site down |
| **four cards loaded** | the only number that can size a replacement honestly |
⚠⚠ **CAPTURE PEAK, NOT AVERAGE.** GPU power has fast transients and UPS overload
protection responds to short-term overload, so a 1-second-average reading can
under-read peaks badly. Use max-hold/peak capture if the meter has it. A figure like
"1100 W average" that hides 1600 W spikes will mis-size the replacement the same way
the current unit got mis-sized. **If the meter is average-only, record the number as a
FLOOR, not as the draw.**
⭐ **The four-card figure goes into `servers/fv-ml1/README.md` permanently.** The
cutover notes flagged that the FV circuit was "likely specced against half the real
draw" while every record still said two GPUs; this closes that with a measurement
instead of an assumption.
## The NAT change, retained as record (DEMOTED — see the ruling above)
**Another session applied a Tailscale SNAT rule to the FV gateway at ~06:22Z — 34
minutes before the site went dark.** See `docs/runbooks/fv-to-ana-nat.md` (uncommitted
as of this writing; not my work, left alone). So "we overloaded the power" is no longer
the only live hypothesis, and the UPS should not be replaced on the strength of a theory
until the one below is run.
**On the evidence, that change is the WRONG SHAPE to have caused this**, and I want that
on the record so nobody wastes the visit chasing it:
- It is one **outbound** SNAT rule, source-scoped to `10.251.50.54/32`, destination-
scoped to `10.250.0.0/16`. An outbound NAT rule cannot stop the gateway, the BMC or
the public WAN address from answering **inbound**.
- The runbook states no routes, filter rules, WAN settings or subnet advertisements were
touched, and that `pfctl -sr` came back byte-identical.
- It was verified working in **both** directions afterwards: FV→hub HTTP 200, FV→ANA
TCP 5432, FV internet HTTPS 200, **ANA→FV SSH reachable**, gateway management intact,
Beszel **18/18 up**.
⚠ Note their BMC observation used **`10.251.50.50`**, which is not the BMC — the BMC is
**`10.251.250.50`**, a different subnet. They correctly declined to claim BMC health, but
the datapoint is *void*, not negative. Do not reason from it either way.
### ⭐⭐ THE DISCRIMINATOR — run this before forming any conclusion
`power.log` is written **locally on `/tank`, every 10 seconds, by a shell loop on the
box.** It does not depend on the network. So:
| `power.log` last entry | what it means |
|---|---|
| **past 06:56Z** | the box **never lost power**. This is a routing/gateway fault, and the UPS is innocent. |
| **stops at ~06:56Z** | the box lost power. UPS/circuit confirmed. |
Cross-check with `uptime` and `journalctl --list-boots` the moment there is a console:
**continuous uptime across 06:56Z kills the UPS theory outright.**
⚠ **So the FIRST action on site is to read, not to fix.** `uptime`,
`journalctl --list-boots`, then `tail power.log`. Establish whether the machine ever
went down before anyone buys hardware or flips anything — the two hypotheses lead to
completely different remediations and only one of them needs a new UPS.
## Three further candidate causes, and what distinguishes them
Cannot be distinguished remotely, because every FV path — including the BMC —
traverses the OPNsense gateway, and the gateway is also dark.
1. **Circuit tripped under two-card load.** Fits the timing and is the predicted
failure: the Anaheim rack breaker tripped twice (2026-08-26, 2026-09-11) on this
same chassis, and the FV circuit was specced while every record still said the box
had **two** GPUs rather than four. ⭐ Distinguishing evidence: **the breaker is
visibly tripped**, and the OPNsense box is dark too (it draws ~20-30 W and would
survive anything short of a circuit/utility loss).
2. **OPNsense gateway crashed or rebooted**, taking all FV routing with it while the
GPU box is fine. Distinguishing evidence: the GPU box's PSU fans/lights are on and
the OPNsense box is the only dead thing.
3. **Upstream utility or colo-side power/network loss**, unrelated to us.
Distinguishing evidence: the breaker is NOT tripped and nothing at the rack has
power; neighbouring equipment is also dark.
⚠ Do not record cause 1 as fact until someone has looked. The timing is strong
circumstantial evidence, not a measurement — and the instrument that would have
measured it (the power log on fv-ml1) died with the box.
## Blast radius
**19 of 30 LiteLLM aliases are dark** — fv-ml1 backs most of the fleet's inference:
char-rp, char-rp-fast, char-rp-reasoning, chat-judge, classifier, coder-fast,
erp-tune-v2, gemma4-26b-a4b-it-base, gen, gen-large, gen-reasoning, image-judge,
qwen-image-bench, qwen3-embedding, reranker, sec, sec-reasoning, summarizer,
summarizer-large
⚠⚠ **THERE IS NO LOCAL FALLBACK, and an earlier note in this session wrongly implied
there was.** Probed 2026-09-13: every free local model on the gateway lives on
fv-ml1. irv-ml1 runs **no LLM chat seat at all** -- it carries TTS (tts-gateway,
breeze-tts, voice-studio, omnivoice-ref), ComfyUI, arbo, yt-voice-clipper, bragi and
waterland-studio, and its two Ampere cards are partly occupied by them. The only
non-fv chat backends on the gateway are **paid**: api.z.ai (9 aliases) and
Moonshot/Kimi (2).
So the choice during the outage is: leave the 19 aliases failing loudly, or spend
vendor credits. ⚠ **If credits are spent, it must be under a NEW alias name that
callers opt into** -- never by silently repointing `summarizer`/`gen`/`classifier` at
GLM. Silent model substitution behind a familiar name is a standing prohibition here
and has already been violated twice; an outage is not an exemption.
The gateway itself on ana-docker is healthy — it is the backends that are gone.
## ⚠⚠ RECOVERY — do NOT just reset the breaker and walk away
Every vLLM seat on fv-ml1 carries `restart: unless-stopped`. On boot **all of them
start loading simultaneously** — gen, mog-sec, erp-seat, gemma4-charrp, flash-next,
embed, rerank, reward, coder, scriberr — which is the single largest power transient
the box can produce, fed straight into a circuit that may have just tripped. That is a
re-trip, and a re-trip during model load can leave a half-written page cache and a
much longer recovery.
**Preferred sequence:**
1. Power the chassis on with **Docker masked**, so nothing auto-starts:
at the BMC/console, boot to the OS and before the network comes up run
`systemctl mask docker containerd` — or if the box is already up and loading,
`systemctl stop docker` immediately.
2. Confirm `nvidia-smi` sees all four cards and `zpool status tank` is ONLINE.
3. Unmask, then bring seats up **one at a time**, waiting for each to report healthy:
`gen` first (19 aliases depend on it), then embed/rerank/reward/coder, then
mog-sec, then the rest. `flash-next` LAST — it is the newest and least depended-on.
4. **Do not restart the MTP campaign.** It is the prime suspect.
5. Watch power while seats come up: `nvidia-smi --query-gpu=index,power.draw --format=csv`.
### ⭐ The staged bring-up also finishes the homepage-label fix, for free
The 2026-09-13 renumber (commit `3132a16`) repaired 25 compose files on this box but
the 10 RUNNING containers were never recreated, so their labels still carried the dead
10.250.50.54. Those containers are gone with the power loss.
⚠ A PLAIN POWER-ON DOES NOT FIX THEM. `restart: unless-stopped` restarts the existing
container with its existing labels; labels only attach at container CREATION. But the
staged `docker compose up -d <svc>` sequence above **is** a recreate, and the compose
files on disk are already corrected — so bringing seats up that way applies the new
labels as a side effect and the dashboard comes back correct. Bring them up with
`compose up -d`, not by letting Docker restore the old containers.
Afterwards, confirm with:
curl -s http://10.0.50.45:5100/api/services | \
python3 -c 'import json,sys;d=json.load(sys.stdin);print([s["href"] for g in d for s in (g.get("services") or []) if "10.250.50.54" in (s.get("href") or "")])'
Expect `[]`. Before the outage that query returned 16 entries.
**Do not reboot the OPNsense firewall** (standing operator directive; its reboot API
403s anyway).
## ⚠ THE CIRCUIT CASE — what split power does and does not buy (operator, 2026-09-13)
> "unless of course the thing trips the circuit anyway."
**It still helps, but only halfway, and the halfway matters.**
- ✅ **A breaker trip is exactly what the split survives.** Firewall + BMC on the UPS is
~25-40 W of load on a 1500 VA unit — hours of battery, not minutes. On a trip the UPS
stops being a load-bearing supply and goes back to being what it is for.
- ❌ **A live firewall is useless if the path OUT of the site is dead.** Our UPS covers
our gear; it does not cover the **colo's handoff** — their switch, ONT or demarc. If
that sits on the circuit we just tripped, the result is a firewall running happily on
battery with nothing upstream to talk to, and the drive happens anyway.
⭐ **ASK THE FACILITY: is the network handoff on our circuit or theirs, and is theirs
on facility UPS?** This is the question that decides whether split power actually
delivers remote diagnosis or merely feels like it does.
### ⚠⚠ And the case where none of the above matters
**If four cards plus host exceeds the circuit, no UPS arrangement helps** — the box does
not fit its feed. Removing an undersized UPS does not remove the constraint, it promotes
the next one:
UPS ~900-1200 W (the one that just gave way)
circuit ~1800 W @ 15 A / ~2400 W @ 20 A
Which side of those the four-card figure lands on decides everything, which is why that
single ammeter reading is the load-bearing measurement of the visit.
### ⭐ The lever that may avoid an electrician: per-card power limits
`nvidia-smi -pl <watts>` caps TGP per card. The box can be made to fit whatever the feed
turns out to be, at a **throughput** cost rather than a **rewiring** cost — four capped
cards on a 15 A circuit is a dial we control today, where a 20 A drop is a ticket and a
site visit.
- Read `nvidia-smi -q -d POWER` first for the enforced min/max range per card; do not
assume how much room the dial has.
- ⚠ **If capping is the answer it MUST be persisted** (systemd unit, or an `if-up`
equivalent). A limit that evaporates on reboot is worse than no limit, because it will
hold right up until the next power event and then silently stop holding.
### Three questions for the site visit
1. What is the **breaker rating** on that circuit?
2. Is the circuit **dedicated** to us, or shared with other racks/tenants?
3. Is the **network handoff** on our circuit or the facility's, and is the facility's on
their UPS?
## ⚠ PROPOSED, NOT RATIFIED — split the power so the management plane survives
**This is a recommendation awaiting the operator's call, not settled intent.** Written
down so tomorrow's rebuild can adopt or reject it deliberately rather than restoring
the arrangement that just failed by default.
UPS <- OPNsense firewall + fv-ml1 BMC only (tens of watts, long runtime)
PDU / wall <- GPU chassis (no UPS in series)
Two reasons:
1. **It fixes the OOB gap this outage exposed.** The BMC's only route to the fleet is
through the firewall, so a power event at the GPU box takes out the management plane
with it -- which is precisely why this incident needs a drive rather than a console
session. Separate the two and a repeat leaves a live firewall, a live BMC, and
remote eyes on a dark chassis.
2. **A 1500 VA unit was never going to hold this box.** It has four cards, not the two
every record claimed until 2026-09-12.
⚠ **Do NOT use a UPS's surge-only outlets to get around its rating.** Both outlet banks
sit downstream of the same input cord, inlet and internal breaker; for a 120 V 1500 VA
unit that is a single NEMA 5-15P rated **12 A at maximum load**, total across all
outlets. The surge bank bypasses the inverter, not the current rating. Overloading the
inverter trips or kills the unit; overloading the cord is a thermal problem in an
unattended rack. Bypass the UPS entirely instead.
If the GPU box is ever to go on battery, it is a 3000 VA / 2700 W-class unit and
probably a 20 A circuit -- but size it from `power.log`, not from a spec sheet.
## Afterwards
- ⭐ **The OOB design gap this exposes.** The cutover chose OPNsense-as-subnet-router
specifically so the BMC stays reachable when the GPU box is down. That works for
box-down/gateway-up. It does **nothing** for a site-wide power or gateway loss —
exactly what happened — because the BMC's only path to the fleet is through that
gateway. A genuine OOB path at FV needs something the FV circuit cannot take down:
an LTE/cellular console, or the BMC on a separate circuit with its own uplink.
- Get the actual circuit rating and the box's real peak draw, now that it is known to
have four cards and not two. Until then, treat concurrent multi-card load at FV as
unproven rather than safe.
- `services/flash-next-mtp-bench/power.log` on the box holds the per-card draw right
up to the cut. Recover it after boot — it is the only measurement of what the load
actually drew, and it survives on `/tank`, not in the container.
## ⭐ POLICY, forward-looking (operator, 2026-09-13): cap new cards at build time
> "i believe our ada cards for the other colo are rated 600w each, we'll want them
> power limited to 200w"
>
> **CLARIFIED BY OPERATOR 2026-09-13 — the two boxes are different hardware:**
>
> | box | cards | TGP each | VRAM total | status |
> |---|---|---|---|---|
> | **fv-ml1** | 4x **Blackwell** RTX PRO 6000 **Max-Q** | **300 W** (Max-Q is the reduced-TGP SKU; the Workstation Edition is 600 W) | 4x96 = 391 GB | deployed, currently dark |
> | **ana-ml3** | 2x **Ada Generation** RTX 6000 | **300 W** | 2x48 = 96 GB | **NOT YET DEPLOYED** |
>
> The 200 W cap directive applies to **ana-ml3**. It very likely wants applying to
> fv-ml1 as well — see the four-card circuit arithmetic below.
The generalised lesson from this outage: **decide the power envelope first and size the
cards into it**, rather than installing cards and discovering the constraint by tripping
it. 4 x 200 W = 800 W of card, which fits a real circuit with a real UPS and headroom.
Three things to settle before that is a plan:
1. ⚠ **200 W may be below the card's ENFORCEABLE FLOOR.** `nvidia-smi -pl` is bounded by
the part's own `Min Power Limit`, which on a high-TGP card is often around half the
rating. If the floor is 300 W, a 200 W request is refused — and, depending on how it
is scripted, refused quietly. **First command on the new hardware:**
nvidia-smi -q -d POWER | grep -iE 'power limit|default'
If the floor lands above 200 W, the envelope has to come from fewer cards or a bigger
feed, not from the cap.
2. ✅ **RESOLVED — both card types are 300 W.** So 200 W is a cap to **67% of TGP**, the
favourable part of the concave curve, not the severe 33% cap a 600 W part would have
implied. The enforceable-floor concern largely goes away too: 200 W was borderline
against a 600 W card's minimum and is very unlikely to sit below a 300 W card's. Worth
the one command; expect it to take.
**ana-ml3: 2 x 200 W = 400 W of card.** Modest, and pointed — ana-ml3 lands in the
**Anaheim** rack whose breaker tripped on 2026-08-26 and 2026-09-11, one of those
caused by this very chassis before it relocated. The cap there is remediation of a
circuit with a track record, not precaution.
### ⭐⭐ The outage arithmetic, now that the TGP is known
2 x Blackwell Max-Q @ 300 W ~ 600 W of card under load
host (board, 566 GB RAM, drives,
fans, PSU conversion loss) ~ 200-350 W <-- UNMEASURED, the gap
-----------
~ 800-950 W
Eaton 1500 VA real watt rating ~ 900-1200 W depending on model
**At or just over the line** — and this is what a vague "undersized" could not explain:
why it ran a full day on one card (~500-650 W, comfortably inside) and died minutes into
the second (~800-950 W, at or past the rating). The host term is the only one being
guessed at, and idle-at-the-plug with all seats down measures it directly.
### ⚠⚠ FOUR cards is a BREAKER problem, not a UPS problem
4 x 300 W card + ~300 W host ~ 1500 W
15 A circuit, 80% continuous = 1440 W
**Four cards uncapped is marginal on a 15 A circuit with no UPS in the path at all.** So
capping belongs at **fv-ml1 too**, not only ana-ml3. If the four-card ammeter reading
confirms it, fv-ml1 needs a per-card cap or a 20 A feed before anyone loads all four
again — and note that today's incident only ever had TWO cards working.
3. ⭐ **Decode tolerates a cap far better than training does**, which is lucky given what
this fleet mostly does. Decode is memory-bandwidth-bound; the perf/watt curve is
strongly concave, so ~60-70% of TGP typically costs ~10-15% of throughput. A cap to
33% of TGP is deeper into the steep region — the cost is real and should be measured
on the first card rather than predicted, and it will hurt a prefill-heavy or training
workload considerably more than a serving seat.
### ⚠ Two placement consequences of Ada, independent of power
- **sm_89 has native FP8 but NOT NVFP4** (Blackwell-only, sm_100/sm_120). Most of our
in-house quants are NVFP4, so **they will not run accelerated on that colo's cards.**
Its seats want FP8 W8A8 builds, or the NVFP4 checkpoints stay on fv-ml1. Same class of
constraint as the Ampere finding for irv-ml1, one generation up.
- ⭐ **It unparks the triton-backend item.** That is a hard no on Ampere — crashes every
render on the A6000, `fp8e4nv` unsupported on sm_86 — and was explicitly deferred TO
Ada. sm_89 has the FP8 support it needs, so it becomes testable on this hardware.
- **VRAM:** 4 x 48 GB = 192 GB, against fv-ml1's 4 x 96 = 391 GB. Big-model placement
stays at FV. The Flash-Next seat needs 74 GiB resident on ONE card and would not fit a
48 GB Ada card even with the n-gram table offloaded — the offload moves the *table*,
not the experts.
⚠ **PERSIST THE CAP.** A hand-set limit holds until the next reboot and then silently
stops holding — the worst possible failure shape, because the thing that reboots the box
is likely to be the power event the cap existed to prevent. Systemd unit, persistence
mode, ordered before Docker starts.