FV could not reach any site but Anaheim. The cause was a single outbound-NAT rule on the FV gateway, added 2026-09-13 and scoped to Anaheim only -- docs/runbooks/fv-to-ana-nat.md says so in as many words: "Other remote sites remain outside this fix's scope." Three mirrors added, same interface and source, only the destination differing: 10.100.0.0/16, 10.0.0.0/16 and 10.6.110.0/24. After: fv-ml1 reaches NH3, ESH, Anaheim, Irvine, the mesh and the internet. Regression sweep clean across nh3-dev, nh3-docker and esh-docker-vm. The runbook now records what the failure looks like, because it presents as a routing or Tailscale fault and is neither. fv-ml1 reached mesh addresses perfectly and LAN addresses not at all; the FV firewall log showed the outbound passing with src=10.251.50.54 and no reply returning; temporary counting rules proved nh3-scale received 5 packets and sent 4 replies; both peers' AllowedIPs were correct. The discriminator that settles it is that every other site pair works -- nh3-docker to esh/ana/FV and esh-docker-vm to FV all succeed -- so a general subnet-to-subnet limitation is ruled out and only outbound SNAT is left. Also reverts the remote-site MASQUERADE rules added to nh3-scale earlier on the asymmetric-return theory. They fired but were not the fix, so they are removed rather than left to accumulate as NAT that achieves nothing. Applied via source_nat/add_rule + apply with a pre-change config backup taken first. Source scope is still fv-ml1's /32, so a second FV host will hit this again -- flagged in the runbook.
34 lines
2.0 KiB
Bash
34 lines
2.0 KiB
Bash
#!/bin/sh
|
|
# Masquerade mesh clients' INTERNET-bound (exit-node) traffic only; preserve site-to-site source.
|
|
iptables -t nat -F MESH-EXIT 2>/dev/null || iptables -t nat -N MESH-EXIT
|
|
|
|
# ── Mesh-member hosts on this routed subnet — MUST come before the RFC1918 RETURNs ──
|
|
# A host that runs Tailscale itself installs an anti-spoof rule:
|
|
# -A ts-input -s 100.64.0.0/10 ! -i tailscale0 -j DROP
|
|
# With source preservation, a mesh client's packet reaches that host's ETHERNET
|
|
# interface still carrying its 100.64.x source, and is dropped there — silently,
|
|
# before anything can answer. Hosts that do NOT run Tailscale are unaffected,
|
|
# which is why every other NH3 address worked and only this one did not.
|
|
# Masquerading just these destinations makes them behave like every other host
|
|
# while leaving source preservation absolute for the rest of the subnet.
|
|
# (2026-09-14. Tailscale's own default is --snat-subnet-routes=true, i.e. SNAT
|
|
# everything; this file is the deliberate departure from that, so the exception
|
|
# belongs here rather than as a reason to abandon the design.)
|
|
#
|
|
# ⚠ Do NOT "fix" this instead by advertising the host's /32 from the host
|
|
# itself. That was tried on nh3-dev 2026-09-14 and black-holed it from ESH,
|
|
# Anaheim, FV and Irvine — `ip rule` there puts `lookup 52` at priority 5270,
|
|
# ahead of main at 32766, and becoming a subnet router let table 52 capture
|
|
# cross-site traffic the node had no accepted route for. Its own LAN and the
|
|
# internet kept working, so a narrow check looks clean. Fix it at the router.
|
|
iptables -t nat -A MESH-EXIT -d 10.100.10.50/32 -j MASQUERADE # nh3-dev
|
|
|
|
iptables -t nat -A MESH-EXIT -d 10.0.0.0/8 -j RETURN
|
|
iptables -t nat -A MESH-EXIT -d 172.16.0.0/12 -j RETURN
|
|
iptables -t nat -A MESH-EXIT -d 192.168.0.0/16 -j RETURN
|
|
iptables -t nat -A MESH-EXIT -d 100.64.0.0/10 -j RETURN
|
|
iptables -t nat -A MESH-EXIT -j MASQUERADE
|
|
iptables -t nat -C POSTROUTING -s 100.64.0.0/10 -o eth0 -j MESH-EXIT 2>/dev/null \
|
|
|| iptables -t nat -A POSTROUTING -s 100.64.0.0/10 -o eth0 -j MESH-EXIT
|
|
|