memory: snapshot — FV cross-site routing fixed, fleet conventions pinned
Session captured: the FV outbound-NAT root cause and its diagnostic signature, the fv-ml1 dead man's switch, fleet identity/group/path conventions and the root:docker normalization, nh3-dev's ts-input reachability fix, ESPHome modernisation and the kb KB-search tool, and the Hermes bearer rotation release. Six new detail files. Tried-and-abandoned gains three: probing OPNsense endpoints by POSTing at them (which rebooted the FV firewall), advertising a /32 from nh3-dev, and the nh3-scale remote-site masquerade rules that fired but were not the fix. Housekeeping: 8 Recent-decisions entries archived to archival-memory.md, and 21 oversized inline entries split into detail files per the two-tier rule -- they had been sitting fully inline in the index, which is what the split exists to prevent. Two pointers to a detail file archived this run were repointed at archival-memory.md. The index is 389 lines, still over the ~300 soft cap. The archival guards stop it there: only 4 further entries are old enough to move and every one carries an open deferred-work pointer. An over-cap file that keeps live decisions beats a scannable one that lost a deferred call.
This commit is contained in:
@@ -0,0 +1,42 @@
|
||||
# `[2026-09-15]` Dead man's switch on fv-ml1 — mesh watchdog
|
||||
|
||||
Every path into Fountain Valley runs through equipment at FV. When fv-ml1 loses
|
||||
its way back to the fleet there is no console, no local hands, and the BMC sits
|
||||
behind the same gateway. This is the net under the next routing change.
|
||||
|
||||
`/usr/local/sbin/fv-mesh-watchdog.sh` + `fv-mesh-watchdog.{service,timer}`,
|
||||
every 60 s. Canonical copies in `servers/fv-ml1/`. Commit `8c8559b`.
|
||||
|
||||
## Design choices that matter
|
||||
|
||||
- **Two anchors that cannot share a failure mode** — a plain-internet one
|
||||
(`1.1.1.1`) and a mesh-only one (`100.64.0.1`). If only the mesh anchor fails,
|
||||
the mesh is the problem and it acts. ⭐ **If BOTH fail it deliberately does
|
||||
nothing** — the site uplink is down, Tailscale cannot fix that, and thrashing
|
||||
tailscaled during an ISP outage turns a wait into an incident.
|
||||
- **Threshold 5 consecutive failures**, counter reset on recovery.
|
||||
- **Narrow remit**: only `tailscale set --accept-routes=false` + re-`up` with a
|
||||
stored key + `systemctl restart tailscaled`. It touches no routes, no
|
||||
firewall, no services — a watchdog with a wide remit is a second way to lose
|
||||
the box.
|
||||
- **Disable file** `/etc/fv-watchdog.disable` for planned work.
|
||||
|
||||
## Proven, not assumed
|
||||
|
||||
Positive control against a black-holed anchor (`MESH_ANCHOR=192.0.2.1` via the
|
||||
conf file, real WAN anchor left in place so the uplink guard did not
|
||||
short-circuit):
|
||||
|
||||
run1..run4 counted 1/5 .. 4/5, no action
|
||||
run5 fired — tailscale up ran, tailscaled restarted, "restore attempt complete"
|
||||
after counter reset to 0 once the real anchor returned
|
||||
|
||||
fv-ml1 stayed reachable throughout.
|
||||
|
||||
## Why it exists
|
||||
|
||||
Earlier the same session, `tailscale up --accept-routes` on fv-ml1 black-holed
|
||||
it from its own LAN: it accepted `10.251.0.0/16` from the gateway — **its own
|
||||
subnet** — and routed the local network through the tunnel. Recovery only worked
|
||||
because its mesh address happened to still answer. Same family as the
|
||||
2026-09-06 nh3-dev incident; see [[2026-09-15-fv-cross-site-snat]].
|
||||
Reference in New Issue
Block a user