8a742f59b8
The link died when ESH lost its public IP during the fiber cutover. Two independent causes, and the second would have defeated the obvious fix: - phase1 ana-to-eshudm was type static, pinned to 70.181.90.232, an address that no longer exists. - nattraversal was disable, so ESP could not have crossed NAT even with the peer IP corrected. pfi-ana-nh3 shares that setting and survives only because NH3 is publicly addressed, which is why the two tunnels diverged. FortiOS refuses `set type dynamic` on an existing tunnel -- "Cannot change tunnel type once configured" -- and rolled back cleanly, so the fix could not be an edit. Rather than delete and recreate, which cascades into the phase2, two static routes and ten policies, the replacement was built alongside: new phase1+phase2 ana-eshudm-dyn (type dynamic, ikev2, aes256-sha1, dh14, NAT-T on, PSK read from the ESH UDM API so neither side needed a new key), static route id 10 at distance 20, and two consolidated multi-zone policies 73/74. The old tunnel is left in place, dead and harmless, as rollback. Verified up: ana-eshudm-dyn_0 97.170.236.56:4500 selectors 1/1 -- the _0 suffix is a dialup child, :4500 is NAT-T, and the address is the carrier's, which is precisely what could never have been pinned. ESH reaches all four colo hosts at 40-56ms, the colo reaches all three ESH hosts, and traceroute drops from eight hops leaking into the carrier network to three hops fully encapsulated. Config was backed up before any write (1.17MB, 36903 lines, off-box). Residual fragility recorded: the UDM's ipsec_local_ip demands a literal address -- empty is rejected as api.err.InvalidPayload -- so it still needs updating when the fiber changes ESH's WAN address. The gateway end is now address-agnostic; the UniFi end is not.
docs/
Navigation map for the documentation tree. New session? Read
orientation.md first — it's the narrative overview
of the fleet, backup architecture, governing principles, and gotchas,
and it points at everything else.
Tree
docs/
├── orientation.md # start here — fleet overview + where-to-look guide
├── runbooks/ # ops runbooks (recovery, deployment phases)
│ ├── disaster-recovery.md
│ ├── nh3-prune-ritual.md
│ └── pbs-deployment.md
└── pfi/ # PFI-specific reference (services, models, VMs)
├── docker-stack.md
├── model-list.md
├── proxmox-vms.md
├── recommended-model-settings.md
├── vm-102-matrix-appservice.md
└── vm-102-matrix-synapse.md
What goes where
runbooks/— step-by-step ops procedures. Anything you'd reach for during an incident or while standing up new infrastructure. Examples: disaster recovery (blast-radius tiers + restoration steps), PBS deployment (9-phase rollout). New runbook → new file here.pfi/— PFI-specific reference material that's too narrow for the top-level CLAUDE.md but doesn't change incident response. AI model inventory, recommended inference settings, Matrix bridge config, Proxmox VM map. New stable reference → new file here.- Top-level (
docs/orientation.md,docs/README.md) — narrative guides about the workspace itself, not about specific infra.
Cross-references
- Fleet topology + servers table: top-level
CLAUDE.md. - Open work + recent milestones: top-level
STATUS.md. - Durable cross-session facts:
~/.claude/projects/-home-lkraven-development-eshpfi-management/memory/.
Conventions
- Markdown, GitHub-flavored. CommonMark renders fine in most viewers.
- File names are lowercase-kebab-case, descriptive. No dates in filenames — git history covers that.
- One topic per file. If a file grows past ~500 lines, look for a natural split before adding more.
- No checked-in binaries or checksums. Build/release artifacts belong
in a build pipeline or
tools/, notdocs/.