5.7 KiB
Cutover: ana-ml2 → fv-ml1 (Anaheim → Fountain Valley colo)
Why: ana-ml2's power draw (dual Blackwell PRO 6000, ~1.5 kW peak) repeatedly trips the Anaheim rack breaker — the root cause of the 2026-08-26 and 2026-09-11 whole-site outages. Moving it to the Fountain Valley colo on its own circuit.
Clean-cutover property: the box is already down (BMC dark, no power) as of the
2026-09-11 outage, so this is a relocation, not a take-down. /tank is LOCAL ZFS
(no NFS from ana-nas — verified), so the data travels with the chassis; nothing
copies over the wire.
Address map (same shape as Anaheim)
| thing | Anaheim (old) | Fountain Valley (new) |
|---|---|---|
| server subnet | 10.250.0.0/16 | 10.251.0.0/16 |
| box | ana-ml2 10.250.50.54 |
fv-ml1 10.251.50.54 (confirmed) |
| BMC | ana-ml2-bmc 10.250.250.50 |
fv-ml1-bmc 10.251.250.50 |
| site gateway / OPNsense | ana-gw 10.250.0.1 | OPNsense = .1 in EVERY FV VLAN (server 10.251.50.1, mgmt 10.251.250.1, …) — multi-homed gateway |
| DNS name | ana-ml2.ana.internal | fv-ml1.fv.internal (full new fv site) |
| mesh | via ana-scale | OPNsense = tailscale subnet-router for 10.251.0.0/16 |
DNS approach (decided): PIGGYBACK now, promote later. dns-sync builds each
FQDN as name.site.zone with NO check that the site is in the sites: block, and
publishes every record to every resolver. So fv-ml1 + fv-ml1-bmc records with
site: fv resolve fleet-wide from the existing ana/esh/nh3 resolvers immediately —
you get the full fv-ml1.fv.internal name with zero new infra. Add a real fv
resolver (AdGuard container on a future FV utility box, or extend dns-sync for
OPNsense Unbound) only when FV needs LOCAL resolution. Do NOT add fv to sites:
until that resolver exists — the sync would fail trying to reach it.
Mesh: OPNsense as subnet-router (NOT ana-ml2)
os-tailscale plugin on the OPNsense, advertise 10.251.0.0/16, approve the route
in headscale. Rationale: mirrors ana-scale/nh3-scale/esh-scale; stays up when
fv-ml1 is powered off, so the BMC stays reachable over the mesh — the out-of-band
path the fleet lacked on 2026-09-11. Do NOT tie the mesh node to fv-ml1 itself.
Blast radius — every reference to repoint (from grep 10.250.50.54)
Load-bearing (traffic breaks if missed):
stacks/litellm/conf/config.yaml— THE critical one (every inference alias: gen/summarizer/classifier/image-judge/mog-sec/char-rp/meromero). At cutover, once fv-ml1 serves at FV:Verify:sed -i 's/10\.250\.50\.54/10.251.50.54/g' stacks/litellm/conf/config.yaml scripts/deploy-stack.sh infra-ops@10.250.50.70 litellm --conf # push to ana-docker ssh infra-ops@10.250.50.70 'cd /opt/docker/compose/litellm && sudo docker compose up -d litellm'gen/summarizer/mog-secroute to fv-ml1. ⚠ Gateway STAYS on ana-docker → cross-site to FV (metro link, fine).dns/internal.yaml— add the two records below (piggyback; NOT asites:entry), comment out the oldana-ml2/ana-ml2-bmclines. Thenscripts/dns-sync.py --dry-run→ apply.# ---- fv: Fountain Valley colo (piggyback DNS until FV has its own resolver) ---- - {name: fv-ml1, site: fv, v4: 10.251.50.54, note: GPU inference, dual RTX PRO 6000 (ex ana-ml2, relocated 2026-09)} - {name: fv-ml1-bmc, site: fv, v4: 10.251.250.50, note: BMC for fv-ml1}- mesh route to 10.251.0.0/16 approved in headscale (OPNsense subnet-router).
stacks/ana-ml2-proxy/— the Homepage Docker-API relay (10.250.50.54:2375). Rename/repoint to the FV IP or retire if Homepage reaches FV another way.
Cosmetic (no traffic break — rename sweep):
- homepage
href/pinglabels: stacks/{llama-swap,gemma4-charrp,meromero-charrp,erp-seat,sglang,heretic2-charrp-reasoning}/compose.yaml, stacks/homepage/conf/services.yaml - servers/ana-ml2/ → servers/fv-ml1/ (dir rename), ssh-target 10.250.50.54→10.251.50.54, ssh_config, remote-ssh allowlist (10.0.0.0/8 already covers 10.251 — no change)
- playbooks/ana-ml2, elway invocations infra-ops@10.250.50.54 → @10.251.50.54
- stacks/beszel deploy target; docs/*, README.md, CLAUDE.md server table, disaster-recovery.md
- model
.PROVENANCE.txtfiles reference ana-ml2 in prose (accuracy, not load-bearing)
Order of operations (cutover day)
- Physically install fv-ml1 at FV, on its dedicated circuit. Confirm the circuit handles ~1.5 kW peak (the whole point).
- OPNsense up: LAN 10.251.50.1 / mgmt 10.251.250.1, WAN,
os-tailscalesubnet-router for 10.251.0.0/16, approve route in headscale. - fv-ml1 up on 10.251.50.54, BMC on 10.251.250.50. Verify BMC reachable over mesh (out-of-band check).
- Bring up the vLLM seats (same /tank, same compose) — verify :8015/:8019 serve locally.
- Apply the LiteLLM api_base repoint on ana-docker →
up -d litellm→ verifygen/summarizer/mog-secroute to FV. - DNS: internal.yaml
fvsite + records →dns-sync.py. - Rename sweep (cosmetic) — a scripted
git mv+ sed pass, one commit. - Beszel/Homepage repoint; verify dashboards.
Rollback
The physical move is not reversible same-day, but the config is: keep the old
ana-ml2 blocks in git history; if FV networking isn't ready, the seats simply
stay down (they already are) — no worse than the current state. Do NOT delete the
ana-ml2 DNS/LiteLLM entries until FV serves; comment them.
Cross-site note
LiteLLM gateway stays on ana-docker (Anaheim); fv-ml1 serves cross-site. FV↔Anaheim is metro (both Orange County, ~10 mi) so latency is a non-issue, but it's a NEW dependency: if the FV↔Anaheim link drops, inference dark. Acceptable; revisit moving the gateway to FV if FV becomes the inference hub.