Files
esh-pfi-infrastructure/servers/nh3-pve/mesh-exit-masq.sh
T
vh 8c8559b8ec feat(fv): mesh dead-man's switch on fv-ml1; partial progress on FV cross-site routing
WATCHDOG (done, proven). fv-mesh-watchdog probes two independent anchors every
minute and, after 5 consecutive failures, puts Tailscale back to known-good:
accept-routes off, re-up against headscale with a stored key. It touches
nothing else — a watchdog with a wide remit is a second way to lose the box.

Two anchors that cannot share a failure mode: a plain-internet one and a
mesh-only one. If BOTH fail the site uplink is down, Tailscale cannot fix that,
and it deliberately does nothing — thrashing tailscaled during an ISP outage
turns a wait into an incident. Disable file at /etc/fv-watchdog.disable for
planned work.

Proven by positive control, not assumed: counter incremented 1..4 without
acting, fired the restore at 5 (tailscale up ran, tailscaled restarted), and
reset to 0 once the real anchor returned. fv-ml1 stayed reachable throughout.

This exists because a  on fv-ml1 black-holed it
from its own LAN earlier the same day: it accepted 10.251.0.0/16 from the
gateway — its OWN subnet — and routed the local network through the tunnel.

FV CROSS-SITE ROUTING (partial). Two changes landed, the path is still broken:

  1. acceptSubnetRoutes 0 -> 1 on the FV gateway's tailscale plugin, via
     settings/set + service/reconfigure (the documented apply, not a reboot).
     The GATEWAY now has 10.0/16, 10.100/16 and 10.250/16 in its routing table
     and reaches NH3 and ESH itself. It could not before.

  2. Remote-site MASQUERADE rules on nh3-scale. The existing jump matched only
     -s 100.64.0.0/10, so traffic from another site's LAN never entered
     MESH-EXIT and kept its original source; an NH3 host then replied via its
     own LAN router instead of back through nh3-scale, making the path
     asymmetric. The rule is confirmed firing (counter increments on FV
     traffic) but does not complete the path.

Still failing: fv-ml1 -> NH3/ESH LAN addresses. Mesh addresses work perfectly
from fv-ml1 (100.64.0.1, 100.64.0.4), Anaheim works over the metro link, and
the FV firewall log shows the outbound passing on tailscale0 with
src=10.251.50.54 and no reply ever returning. The remaining gap is forwarded
FV-LAN traffic specifically, not the gateway's own.

Full regression sweep clean: nh3-dev, ana-docker and esh-docker-vm all reach
all four sites plus the internet.
2026-09-15 00:25:55 -07:00

54 lines
3.3 KiB
Bash

#!/bin/sh
# Masquerade mesh clients' INTERNET-bound (exit-node) traffic only; preserve site-to-site source.
iptables -t nat -F MESH-EXIT 2>/dev/null || iptables -t nat -N MESH-EXIT
# ── Mesh-member hosts on this routed subnet — MUST come before the RFC1918 RETURNs ──
# A host that runs Tailscale itself installs an anti-spoof rule:
# -A ts-input -s 100.64.0.0/10 ! -i tailscale0 -j DROP
# With source preservation, a mesh client's packet reaches that host's ETHERNET
# interface still carrying its 100.64.x source, and is dropped there — silently,
# before anything can answer. Hosts that do NOT run Tailscale are unaffected,
# which is why every other NH3 address worked and only this one did not.
# Masquerading just these destinations makes them behave like every other host
# while leaving source preservation absolute for the rest of the subnet.
# (2026-09-14. Tailscale's own default is --snat-subnet-routes=true, i.e. SNAT
# everything; this file is the deliberate departure from that, so the exception
# belongs here rather than as a reason to abandon the design.)
#
# ⚠ Do NOT "fix" this instead by advertising the host's /32 from the host
# itself. That was tried on nh3-dev 2026-09-14 and black-holed it from ESH,
# Anaheim, FV and Irvine — `ip rule` there puts `lookup 52` at priority 5270,
# ahead of main at 32766, and becoming a subnet router let table 52 capture
# cross-site traffic the node had no accepted route for. Its own LAN and the
# internet kept working, so a narrow check looks clean. Fix it at the router.
iptables -t nat -A MESH-EXIT -d 10.100.10.50/32 -j MASQUERADE # nh3-dev
iptables -t nat -A MESH-EXIT -d 10.0.0.0/8 -j RETURN
iptables -t nat -A MESH-EXIT -d 172.16.0.0/12 -j RETURN
iptables -t nat -A MESH-EXIT -d 192.168.0.0/16 -j RETURN
iptables -t nat -A MESH-EXIT -d 100.64.0.0/10 -j RETURN
iptables -t nat -A MESH-EXIT -j MASQUERADE
iptables -t nat -C POSTROUTING -s 100.64.0.0/10 -o eth0 -j MESH-EXIT 2>/dev/null \
|| iptables -t nat -A POSTROUTING -s 100.64.0.0/10 -o eth0 -j MESH-EXIT
# ── Remote-SITE sources, not just mesh clients ────────────────────────────────
# The jump above matches only 100.64.0.0/10, so traffic from another site's LAN
# arriving over the mesh never enters MESH-EXIT and keeps its original source.
# An NH3 host then replies via its own LAN router instead of back through this
# node, the path is asymmetric, and the reply is lost. Measured 2026-09-15:
# fv-ml1 reached 100.64.0.1 and 100.64.0.4 fine while 10.100.50.40 failed
# outright, and the FV firewall log showed the outbound passing with
# src=10.251.50.54 and nothing ever coming back.
#
# Masquerading remote-site sources onto this node's LAN address makes the reply
# return here, where the conntrack state lives. It costs source visibility for
# cross-site traffic on the NH3 LAN — the same trade as the nh3-dev rule above,
# and the alternative is no connectivity at all.
#
# ⚠ Sites are listed explicitly rather than using 10.0.0.0/8: a blanket rule
# would also masquerade NH3-local traffic that has no business being rewritten.
for site_net in 10.251.0.0/16 10.0.0.0/16 10.250.0.0/16 10.6.110.0/24; do
iptables -t nat -C POSTROUTING -s "$site_net" -o eth0 -j MASQUERADE 2>/dev/null \
|| iptables -t nat -A POSTROUTING -s "$site_net" -o eth0 -j MASQUERADE
done