memory: snapshot — FV cross-site routing fixed, fleet conventions pinned

Session captured: the FV outbound-NAT root cause and its diagnostic signature,
the fv-ml1 dead man's switch, fleet identity/group/path conventions and the
root:docker normalization, nh3-dev's ts-input reachability fix, ESPHome
modernisation and the kb KB-search tool, and the Hermes bearer rotation
release. Six new detail files.

Tried-and-abandoned gains three: probing OPNsense endpoints by POSTing at them
(which rebooted the FV firewall), advertising a /32 from nh3-dev, and the
nh3-scale remote-site masquerade rules that fired but were not the fix.

Housekeeping: 8 Recent-decisions entries archived to archival-memory.md, and 21
oversized inline entries split into detail files per the two-tier rule -- they
had been sitting fully inline in the index, which is what the split exists to
prevent. Two pointers to a detail file archived this run were repointed at
archival-memory.md.

The index is 389 lines, still over the ~300 soft cap. The archival guards stop
it there: only 4 further entries are old enough to move and every one carries an
open deferred-work pointer. An over-cap file that keeps live decisions beats a
scannable one that lost a deferred call.
This commit is contained in:
vh
2026-09-15 00:53:48 -07:00
parent 0ab9da5b89
commit 838132cd6b
35 changed files with 1717 additions and 1289 deletions
@@ -0,0 +1,42 @@
# `[2026-09-15]` Dead man's switch on fv-ml1 — mesh watchdog
Every path into Fountain Valley runs through equipment at FV. When fv-ml1 loses
its way back to the fleet there is no console, no local hands, and the BMC sits
behind the same gateway. This is the net under the next routing change.
`/usr/local/sbin/fv-mesh-watchdog.sh` + `fv-mesh-watchdog.{service,timer}`,
every 60 s. Canonical copies in `servers/fv-ml1/`. Commit `8c8559b`.
## Design choices that matter
- **Two anchors that cannot share a failure mode** — a plain-internet one
(`1.1.1.1`) and a mesh-only one (`100.64.0.1`). If only the mesh anchor fails,
the mesh is the problem and it acts. ⭐ **If BOTH fail it deliberately does
nothing** — the site uplink is down, Tailscale cannot fix that, and thrashing
tailscaled during an ISP outage turns a wait into an incident.
- **Threshold 5 consecutive failures**, counter reset on recovery.
- **Narrow remit**: only `tailscale set --accept-routes=false` + re-`up` with a
stored key + `systemctl restart tailscaled`. It touches no routes, no
firewall, no services — a watchdog with a wide remit is a second way to lose
the box.
- **Disable file** `/etc/fv-watchdog.disable` for planned work.
## Proven, not assumed
Positive control against a black-holed anchor (`MESH_ANCHOR=192.0.2.1` via the
conf file, real WAN anchor left in place so the uplink guard did not
short-circuit):
run1..run4 counted 1/5 .. 4/5, no action
run5 fired — tailscale up ran, tailscaled restarted, "restore attempt complete"
after counter reset to 0 once the real anchor returned
fv-ml1 stayed reachable throughout.
## Why it exists
Earlier the same session, `tailscale up --accept-routes` on fv-ml1 black-holed
it from its own LAN: it accepted `10.251.0.0/16` from the gateway — **its own
subnet** — and routed the local network through the tunnel. Recovery only worked
because its mesh address happened to still answer. Same family as the
2026-09-06 nh3-dev incident; see [[2026-09-15-fv-cross-site-snat]].