memory: snapshot — run 3 gated DO-NOT-SERVE, run 3c held on a tripped breaker
Run 3 trained, gated and dispositioned do-not-serve on a measured 44pp self-harm guardrail regression that its own preregistered rule passed -- a pooled preserve-list test cannot see a single-axis collapse. Run 3c (lr 20x cut, single variable) launched, killed by an Anaheim power-breaker trip at step 80, relaunched, then stopped by the operator at step 22 pending a weekend power triage. Also captured: the corpus mix was specified in a unit the optimiser never sees (45.8% dialogue by context, 24.2% by loss); the dose-response says benefit and damage are one direction in weight space, so the merge-back measures the problem rather than fixing it; four guests including the storage SPOF had onboot unset and never came back from the outage, now fixed with dependency ordering; and a transport failure that enters a measurement as a value looks like whatever you hoped to find -- which found a live defect in another agent's instrument an hour after it was reported. Auto-archived 8 entries to archival-memory.md (Recent decisions: 8, Tried and abandoned: 0); 4 held back on open deferred-work pointers.
This commit is contained in:
@@ -0,0 +1,70 @@
|
||||
# `[2026-08-27]` Anaheim tripped a power breaker — and four guests including the NAS never came back
|
||||
|
||||
Site-wide outage, ~90 minutes. **Operator-confirmed cause: a tripped power breaker**, not a
|
||||
fault and not the tunnel. The discriminator that established scope: `ana-srv1`
|
||||
(38.120.12.44:443, Anaheim's PUBLIC address) was dark **from the internet**, so it was not the
|
||||
NH3↔ANA IPsec tunnel stranding NH3 — the site was not answering on any path. `ana-ml2` returned
|
||||
with `up 1 min`, confirming a hard power event.
|
||||
|
||||
## ⚠ THE DURABLE FINDING — `onboot` was unset on four guests
|
||||
|
||||
pfi-pve came back and auto-started everything **except**:
|
||||
|
||||
CT109 ana-nas the storage SPOF
|
||||
CT113 ana-wg the WireGuard remote-access path
|
||||
CT112 ana-filebot
|
||||
VM106 corviduo-dev
|
||||
|
||||
All four had `onboot` unset. **Recovery was manual and would have been manual every time** —
|
||||
including for the NAS that postgres/PBS/cross-site-restic depend on, and the WireGuard host
|
||||
that is the way in when the site misbehaves.
|
||||
|
||||
**FIXED, with dependency ordering** (operator-authorised):
|
||||
|
||||
CT109 ana-nas onboot=1 order=1,up=45 <- first; 45s for NFS to SERVE
|
||||
CT113 ana-wg onboot=1 order=2 <- remote access before anything can fail
|
||||
VM104/105 Mongo/Postgres order=3,up=60 (pre-existing)
|
||||
VM102 ANA-Docker order=4 (pre-existing)
|
||||
VM101 ANA-DC order=5,up=120 (pre-existing)
|
||||
CT112 ana-filebot onboot=1 order=10
|
||||
VM106 corviduo-dev onboot=1 order=10
|
||||
|
||||
Every guest on pfi-pve now auto-starts. ana-nas precedes the databases deliberately; the
|
||||
`up=45` is for NFS to be *serving*, not merely for the container to be *running* — the exact
|
||||
distinction that killed `rest-server` on ana-docker, which came up before the NAS existed,
|
||||
found nothing to serve, and exited 255.
|
||||
|
||||
## ⚠ `/tank` came back DEGRADED — a disk is genuinely gone
|
||||
|
||||
tank DEGRADED, raidz2-0, 7 devices ONLINE
|
||||
9477159196657038377 FAULTED was /dev/nvme4n1p1
|
||||
errors: No known data errors
|
||||
|
||||
**Only 7 physical NVMe present where the pool expects 8** — checked, so not renumbering. One
|
||||
drive did not re-enumerate. raidz2 carries two disks of parity; one is spent. Operator taking
|
||||
it; chassis is a Supermicro AS-4125GS-TNRT2 with PCIe hot-plug slots, so a swap should not
|
||||
need a power-down.
|
||||
|
||||
## ⚠ `/mnt/smithy` is manual by design — it will be missing after EVERY reboot
|
||||
|
||||
Not in fstab, and **deliberately so**: a cross-site NFS entry can hang boot on a GPU host, and
|
||||
it is `soft` rather than `hard` because ana-ml2 is cross-site from that NAS and a hard mount
|
||||
turns a link blip into unkillable D-state. Remount with the recorded spec, do NOT "fix" it into
|
||||
fstab:
|
||||
|
||||
sudo mount -t nfs4 -o ro,soft,timeo=30,retrans=3,proto=tcp,vers=4.1 \
|
||||
10.100.50.50:/volume1/smithy /mnt/smithy
|
||||
|
||||
Full rationale: [[2026-08-23-smithy-mount-ana-ml2]].
|
||||
|
||||
## Power capacity is now the open item
|
||||
|
||||
Operator: *"we'll triage this weekend, probably shut down some seats."* ana-ml2 alone was
|
||||
pulling ~600 W across both GPUs at their 300 W caps during training. `gen` stays up by
|
||||
instruction; everything else on that box is idle.
|
||||
|
||||
## Blast radius beyond us
|
||||
|
||||
heid lost **both gateway-routed arms of a four-arm panel** mid-dispatch and discovered the
|
||||
outage by losing half a panel. That report produced the single most valuable artifact of the
|
||||
incident — see [[2026-08-27-empty-response-as-a-datum]].
|
||||
Reference in New Issue
Block a user