Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-11-ana-ml2-fv-ml1-relocating-to.md
T
vh 838132cd6b memory: snapshot — FV cross-site routing fixed, fleet conventions pinned
Session captured: the FV outbound-NAT root cause and its diagnostic signature,
the fv-ml1 dead man's switch, fleet identity/group/path conventions and the
root:docker normalization, nh3-dev's ts-input reachability fix, ESPHome
modernisation and the kb KB-search tool, and the Hermes bearer rotation
release. Six new detail files.

Tried-and-abandoned gains three: probing OPNsense endpoints by POSTing at them
(which rebooted the FV firewall), advertising a /32 from nh3-dev, and the
nh3-scale remote-site masquerade rules that fired but were not the fix.

Housekeeping: 8 Recent-decisions entries archived to archival-memory.md, and 21
oversized inline entries split into detail files per the two-tier rule -- they
had been sitting fully inline in the index, which is what the split exists to
prevent. Two pointers to a detail file archived this run were repointed at
archival-memory.md.

The index is 389 lines, still over the ~300 soft cap. The archival guards stop
it there: only 4 further entries are old enough to move and every one carries an
open deferred-work pointer. An over-cap file that keeps live decisions beats a
scannable one that lost a deferred call.
2026-09-15 00:53:48 -07:00

2.1 KiB

[2026-09-11] ana-ml2 → fv-ml1: relocating to a NEW Fountain Valley colo TOMORROW (operator decision). Its power draw (dual

⭐ ana-ml2 → fv-ml1: relocating to a NEW Fountain Valley colo TOMORROW (operator decision). Its power draw (dual Blackwell PRO 6000, ~1.5 kW peak) is the ROOT CAUSE of the repeated Anaheim rack-breaker trips (2026-08-26, 2026-09-11) — moving it to its own circuit fixes the recurring whole-site outage. New site fv, same shape as Anaheim: server subnet 10.251.50.0/24 (fv-ml1 = 10.251.50.54, mirroring the old host octet), mgmt/BMC 10.251.250.0/24 (fv-ml1-bmc = 10.251.250.50). OPNsense firewall is the multi-homed gateway (.1 in every FV VLAN) AND the tailscale/headscale subnet-router advertising 10.251.0.0/16 — chosen over ana-ml2-as-endpoint specifically because the firewall stays up when the GPU box is down, giving out-of-band BMC access over the mesh — the exact thing the fleet LACKED during today's outage (no OOB path, BMC islanded). Rename to fv-ml1, full fv.internal DNS name. DNS approach: PIGGYBACK — dns-sync builds name.site.zone with no check that the site is in the sites: block, so fv-ml1/fv-ml1-bmc records with site: fv resolve fleet-wide from the existing ana/esh/nh3 resolvers immediately; add a real fv resolver only when FV needs LOCAL resolution (OPNsense can't host the AdGuard the sync targets — it's FreeBSD/Unbound). Clean cutover: the box is already down (BMC dark, no power since the outage), and /tank is LOCAL ZFS with NO NFS from ana-nas, so data travels with the chassis. ⚠ Load-bearing repoint = stacks/litellm/conf/config.yaml (~10 api_base: 10.250.50.54:{8015,8016,8018,8019} → 10.251.50.54; darkens every inference alias if missed) — gateway STAYS on ana-docker so fv-ml1 serves cross-site (FV↔Anaheim metro, fine). Everything staged, nothing deployed: runbook docs/runbooks/fv-ml1-cutover.md (commit ce04f9d; exact DNS + LiteLLM commands) + scripts/fv-ml1-rename-sweep.sh (8400f3a; scoped, dry-run default, history/provenance-safe, manual-review list for judgement calls).