405 lines
28 KiB
Markdown
405 lines
28 KiB
Markdown
# Headscale overlay mesh — plan (2026-09-06)
|
||
|
||
**Status:** planning. Operator direction 2026-09-06: replace Site Magic (NH3↔ESH) and the
|
||
FortiGate IPsec tunnels (colo↔NH3, colo↔ESH) with a self-hosted Tailscale-protocol overlay
|
||
(Headscale), one ultra-light LXC per PVE host, keeping the old tunnels as an emergency
|
||
backup. This document is the plan that a `.contract.md` will be cut from; nothing here is
|
||
provisioned yet.
|
||
|
||
## Why Headscale (decided 2026-09-06)
|
||
|
||
Tailscale clients are the best in class on the devices actually used for remote access
|
||
(iPad, Mac, Linux); the future OPNsense colo edge can be a Tailscale node natively
|
||
(`os-tailscale`); Headscale is one binary + one SQLite file to self-host, versus NetBird's
|
||
six-service stack with a mandatory IdP. NetBird's built-in UI and HA routing groups are
|
||
real but not decisive for three sites. UDP-blocked networks (in-flight Wi-Fi) work because
|
||
the client falls back to a relay over TCP 443; plain WireGuard (ana-wg) cannot.
|
||
|
||
## Site facts that drive placement
|
||
|
||
| site | edge | v4 | v6 | change risk |
|
||
|---|---|---|---|---|
|
||
| ANA colo | FortiGate 80F → **OPNsense on R420 within the month** | public static | pending | HIGH — edge in flux; incident history 2026 (breaker, PSU1, WAN admin closed) |
|
||
| NH3 | UDM SE, stable | dynamic-but-stable (DDNS) | single /64, **reserved for meshing** | low |
|
||
| ESH | UDM Pro Max | CGNAT now, **static soon**; 2G symmetric soon | /56 | low, improving |
|
||
|
||
**Control plane at NH3 now** (only stable edge; static v6 + DDNS v4 → one DNS name).
|
||
**Reassess ESH** once its static v4 lands — it becomes the relay site regardless (2G
|
||
symmetric). **Never the colo:** the mesh's brain must not live in the building it exists
|
||
to reach in an emergency. Migration later = copy one SQLite file + move the DNS name;
|
||
clients follow the URL and never re-enrol.
|
||
|
||
## Topology
|
||
|
||
```
|
||
headscale.phasefinal.com (A via DDNS, AAAA static)
|
||
│ HTTPS 443 (+ DERP later)
|
||
┌─────────────────────────────┼───────────────────────────────┐
|
||
│ NH3 nh3-pve │ │
|
||
│ CT nh3-headscale ──control plane (1 core / 512M / 8G) │
|
||
│ CT nh3-mesh-rtr ──subnet router, advertises 10.100.0.0/16│
|
||
├──────────────────────────────────────────────────────────────┤
|
||
│ ESH esh-pve │
|
||
│ CT esh-mesh-rtr ──subnet router, advertises 10.0.0.0/16 │ (relay/DERP here later)
|
||
├──────────────────────────────────────────────────────────────┤
|
||
│ ANA pfi-pve │
|
||
│ CT ana-mesh-rtr ──subnet router, advertises 10.250.0.0/16│ (OPNsense takes this over later)
|
||
└──────────────────────────────────────────────────────────────┘
|
||
Per-device clients: nh3-dev, laptops/iPad, corviduo-dev, ana-ml2 … (MagicDNS names)
|
||
```
|
||
|
||
Four unprivileged Debian 13 LXCs, each 1 vCPU / 512 MB / 8 GB, `onboot=1`, backed up by
|
||
the existing `pbs-ana` job. Routers need `/dev/net/tun` passed in and `ip_forward` (both
|
||
routine for unprivileged CTs; net sysctls are namespaced).
|
||
|
||
Proposed ids/names (DHCP with reservation like every existing CT; recorded in
|
||
`dns/internal.yaml` as `<name>.<site>.internal`):
|
||
|
||
| PVE | next id | CT | note |
|
||
|---|---|---|---|
|
||
| nh3-pve (`root@10.100.250.60`, PVE 8.4.1) | 106 | `nh3-headscale` | control plane |
|
||
| nh3-pve | 107 | `nh3-mesh-rtr` | subnet router |
|
||
| esh-pve (`root@10.0.250.35`, PVE 8.4.20) | 108 | `esh-mesh-rtr` | subnet router |
|
||
| pfi-pve (`root@10.250.250.31`, PVE 8.3.5) | 114 | `ana-mesh-rtr` | subnet router; ana-wg (113) stays as independent WG fallback |
|
||
|
||
`debian-13-standard_13.6-1_amd64.tar.zst` is in `pveam available` on all three (not yet
|
||
downloaded on any).
|
||
|
||
## Access & credentials — all in hand (verified 2026-09-06)
|
||
|
||
| need | have |
|
||
|---|---|
|
||
| provision LXCs on all three PVEs | `infra-ops@` with NOPASSWD sudo on all four PVE hosts since 2026-09-06 (elway `bootstrap-infra-ops-user.yaml`; PVE needed `sudo` installed first); `root@` also works |
|
||
| public DNS name + DDNS | Cloudflare all-zones DNS-edit token, vault `nh3-dev/.config/cloudflare/infra-ops-dns-token` |
|
||
| NH3 UDM port-forward 443 → nh3-headscale; static routes on both UDMs | UDM API keys, vault `unifi/pfi-udmse-api-key`, `unifi/esh-udmpm-api-key` (classic `/rest/*` read+write) |
|
||
| colo static route toward ana-mesh-rtr | FortiGate infra-ops SSH pw vaulted; reachable at 10.250.0.1 via the tunnel (`execute backup config` first). Moot once OPNsense lands |
|
||
| TLS for headscale | Let's Encrypt via Cloudflare DNS-01 (same token) — no inbound 80 needed |
|
||
| secrets (pre-auth keys, API key, DB) | vault under `nh3-headscale/…` |
|
||
|
||
Nothing outstanding on credentials. Inputs still needed from the operator: confirm names/ids
|
||
above; 443 direct on the UDM vs behind the existing Caddy on nh3-dev; ACL posture (flat
|
||
"everything can reach everything" first, tighten later, is the recommendation).
|
||
|
||
## Phases
|
||
|
||
0. **Pre-flight (no changes):** confirm NH3 v6 /64 address for the AAAA; confirm DHCP
|
||
reservation ranges on the three sites; `pveam download` Debian 13 on all three.
|
||
1. **Control plane:** CT `nh3-headscale`; headscale in a container or the .deb (prefer .deb —
|
||
fewer layers in a 512M CT); Caddy/own TLS via DNS-01; UDM forward 443; DDNS updater
|
||
(UDM → Cloudflare, or a `ddclient`/cron in the CT); DERP = Tailscale public map
|
||
initially; headplane UI optional. Backup: `/var/lib/headscale/db.sqlite` is the only
|
||
state — PBS covers the CT; add a nightly `sqlite3 .backup` to `backupStore` too.
|
||
2. **First nodes:** nh3-dev + the operator's laptop/iPad enrol → prove MagicDNS, prove
|
||
TCP-443 relay path from a UDP-blocked network.
|
||
3. **Subnet routers:** three router CTs, `--advertise-routes` per site, approve routes in
|
||
headscale, SNAT off (`--snat-subnet-routes=false`) so source IPs survive, static routes
|
||
on each site gateway pointing the other two /16s at the local router. Test
|
||
site-to-site from clientless hosts.
|
||
4. **Cut over:** move day-to-day traffic onto the mesh; **disable (not delete)** Site Magic
|
||
and the two IPsec tunnels. Running all three at once makes route precedence murky on
|
||
the UDM and produces asymmetric paths. ana-wg stays as the out-of-band WG fallback.
|
||
5. **Later:** relay (DERP) at ESH when the 2G circuit is in; OPNsense as the colo router
|
||
node; HA router pairs (Tailscale HA subnet routers) per site; tighten ACLs.
|
||
|
||
## Risks / open questions
|
||
|
||
- Control-plane outage does not drop existing tunnels but stalls new logins and key
|
||
rotation — hence the SQLite backup and the "not at the colo" rule.
|
||
- Tailscale public DERPs carry relayed (encrypted) traffic through third-party infra until
|
||
a self-hosted DERP exists. Acceptable at first; fix in phase 5.
|
||
- CGNAT at ESH: outbound-only is fine; nothing at ESH needs to be dialled inbound.
|
||
- Headscale lags Tailscale feature-wise (tailnet lock, some ACL syntax). Not relevant to
|
||
phases 0–4.
|
||
- VM 106 on pfi-pve was once named "PFI-Tailscale" (pre-2026 inventory). No config
|
||
survives; nothing to reuse.
|
||
|
||
## Status — 2026-09-06 05:35Z: phases 0–1 done, phase 3 half-done, NOT cut over
|
||
|
||
Operator inputs (05:00Z): 443 direct on the UDM; names `nh3-headscale`, `nh3-scale`,
|
||
`esh-scale`, `ana-scale`; ids my call → sequential. ACLs: flat (database policy mode,
|
||
empty policy) until told otherwise.
|
||
|
||
**Live:**
|
||
|
||
| what | where | detail |
|
||
|---|---|---|
|
||
| control plane | CT 106 `nh3-headscale` 10.100.50.45 (nh3-pve, Debian 12, 1c/512M/8G) | headscale **v0.29.3** (.deb), `https://headscale.phasefinal.com` on **443 direct**, Let's Encrypt via TLS-ALPN-01 (no Caddy, no DNS-01), SQLite at `/var/lib/headscale/db.sqlite`, DERP = Tailscale public map, MagicDNS base `mesh.phasefinal.com`, split DNS `internal` → the three AdGuards, `override_local_dns: false`. Public health check 200 with a valid cert. |
|
||
| public name | Cloudflare A `headscale.phasefinal.com` → NH3 WAN (70.230.226.88 at creation), unproxied, TTL 120 | **DDNS:** `~/.local/bin/headscale-ddns.sh` on nh3-dev, user timer `headscale-ddns.timer` every 10 min, token pulled from the vault at run time (nothing stored in the CT). No AAAA yet — NH3 LAN v6 is deliberately off. |
|
||
| UDM forward | NH3 UDM SE `headscale-443`: WAN tcp/443 → 10.100.50.45:443 | added via the classic REST API. Hairpin from inside NH3 works. |
|
||
| routers | CT 107 `nh3-scale` 10.100.50.46 (100.64.0.1) · CT 108 `esh-scale` 10.0.50.65 (100.64.0.2, Debian 13) · CT 114 `ana-scale` 10.250.50.45 (100.64.0.3) | tailscale 1.102.3, `/dev/net/tun` passed, forwarding on, each advertising its site /16, **routes approved and serving**, `--snat-subnet-routes=false`, `--accept-routes=false` (see lesson 1). |
|
||
| first client | nh3-dev (100.64.0.4) | `--accept-routes=false`, `--accept-dns=false`. Direct paths to all three routers (7–10 ms). |
|
||
| DNS | `*.internal` names for all four CTs synced to the three AdGuards | |
|
||
| backups | pfi-pve + nh3-pve jobs are `all`; esh-pve job vmid list extended with 108 | |
|
||
| vault | `headscale/preauth-router-48h-20260906`, `headscale/preauth-client-48h-20260906` (reusable, expire 2026-09-08) | mint fresh ones on the CT: `headscale preauthkeys create --user 1 --reusable --expiration 48h` |
|
||
|
||
**Lessons (do not repeat):**
|
||
|
||
1. ⚠ **`--accept-routes` on a client BEFORE the return path exists breaks that client's
|
||
LAN.** nh3-dev accepted all three /16s including its OWN site's; Linux table-52 routes
|
||
beat the main table, so every 10.100.x packet went into the tunnel to nh3-scale and
|
||
nothing came back (no SNAT → the far host has no route to 100.64/10). Took ssh to
|
||
nh3-pve, esh-pve and pfi-pve down for ~1 min until `tailscale set --accept-routes=false`.
|
||
Phase order is therefore: **return path first** (either SNAT on the routers, or static
|
||
routes for `100.64.0.0/10` → local router on every site gateway), *then* accept-routes on
|
||
clients, and exclude a node's own site route.
|
||
2. The routers currently find each other "direct" via LAN addresses (`10.0.50.65:41641`
|
||
etc.) — i.e. **the mesh is riding inside the old Site Magic / IPsec tunnels**. Public-
|
||
endpoint / relay path discovery is untested until those tunnels are disabled (phase 4).
|
||
Test that with one tunnel down before declaring cut-over ready.
|
||
3. pfi-pve (PVE 8.3.5) and nh3-pve (8.4.1) refuse the Debian 13 template
|
||
(`unsupported debian version '13.6'`); esh-pve (8.4.20) accepts it. Three CTs are
|
||
Debian 12, esh-scale is Debian 13. A failed `pct create` followed by an `>>` into the
|
||
conf leaves a stub conf that makes `pct` say "missing arch" — delete the stub. Follow-up:
|
||
bring pfi-pve/nh3-pve up to current 8.4.
|
||
4. headscale 0.29 removed `randomize_client_port` and renamed
|
||
`ephemeral_node_inactivity_timeout` → `node.ephemeral.inactivity_timeout`; with a bad
|
||
key `headscale serve` silently falls back to defaults (127.0.0.1:8080). Always
|
||
`headscale configtest` first.
|
||
|
||
**Next (in order):** (a) decide the return path — recommendation: static routes for
|
||
`100.64.0.0/10` on each site gateway → local router (keeps source IPs), SNAT as the fallback;
|
||
(b) enrol the operator's devices with the client key; (c) test a UDP-blocked path (relay over
|
||
443); (d) disable ONE old tunnel and re-test router paths; (e) cut over: static routes for the
|
||
other two /16s on each gateway, disable Site Magic + IPsec; (f) v6 AAAA once NH3 LAN v6 is
|
||
on; (g) self-hosted DERP at ESH; (h) PVE upgrades on pfi-pve/nh3-pve.
|
||
|
||
### 06:05Z — return path built and verified (operator approved 05:50Z)
|
||
|
||
Static route `100.64.0.0/10` → local router on all three gateways: NH3 UDM
|
||
(`mesh-100.64-via-nh3-scale` → 10.100.50.46, classic REST), ESH UDM
|
||
(`mesh-100.64-via-esh-scale` → 10.0.50.65; ⚠ that vault item is a 19-line note — extract with
|
||
`secret get unifi/esh-udmpm-api-key | grep '^key:' | awk '{print $2}'`, a raw pipe gives nginx
|
||
400 and the NH3 key gives 401), FortiGate (`config router static`, dst 100.64.0.0/255.192.0.0,
|
||
gw 10.250.50.45, device `servers`; config backed up to flash as `pre-mesh-route-20260906`
|
||
first). Routers re-set to `--accept-routes=true` — a router MUST accept the other sites' routes
|
||
or its replies to remote-LAN sources leave via its default gateway with a 100.64 source and die;
|
||
a node never accepts its own advertised route, so the nh3-dev failure mode does not apply to
|
||
routers.
|
||
|
||
Verified (LAN host → gateway → local router → mesh → far router, and back): nh3-docker→ana-scale
|
||
9 ms, esh-docker-vm→ana-scale and →nh3-scale 0% loss, ana-docker→nh3-scale and →esh-scale 0%
|
||
loss. All three sites can reach mesh addresses from plain LAN hosts with no client installed.
|
||
|
||
**LAN-resident clients (nh3-dev) keep `--accept-routes=false`:** they reach other sites via
|
||
their gateway like any LAN host. `--accept-routes` is for roaming devices (laptop/iPad off-site)
|
||
only. Consequence: a LAN-resident client is reachable by its mesh IP only from the mesh, not
|
||
from other sites' LANs (ana-docker→100.64.0.4 fails by design; →100.64.0.1 works).
|
||
|
||
**Enrolling GUI devices (Mac/iPad) without OIDC:** the apps do not take a pre-auth key. Point
|
||
the app at `https://headscale.phasefinal.com`, it opens a registration page showing a
|
||
`headscale nodes register --user pfi --key mkey:…` command; run that on nh3-headscale
|
||
(`pct exec 106 -- headscale nodes register --user 1 --key <mkey>`). CLI Linux/macOS builds can
|
||
use the vaulted client key instead.
|
||
|
||
### 06:20Z — first roaming client verified
|
||
|
||
Operator's MacBook enrolled via the GUI (Option-click → Debug → Custom Login Server; headscale
|
||
0.29 registration command is `headscale auth register --auth-id hskey-authreq-… --user pfi`),
|
||
node 5 `vhlk-mba26` (100.64.0.5), "Use Tailscale subnets" on. From ESH: `ping 100.64.0.3`
|
||
and `ssh infra-ops@ana-docker.ana.internal` both work → colo subnet route + split DNS for
|
||
`*.internal` proven from a client. (Path was still via the ESH LAN / old tunnels.)
|
||
|
||
### Note — the mesh is also the durable fix for CGNAT CrowdSec false-bans
|
||
|
||
2026-09-06: CrowdSec on ana-docker banned ESH's shared CGNAT egress (23.164.40.160,
|
||
`custom/gitea-aggressive-crawl`) and the fortigate-mirror bouncer pushed it to the colo
|
||
edge, blackholing Matrix/gitea/chat for the whole ESH site (see
|
||
`~/.claude/.../memory/incident_crowdsec_cgnat_false_ban.md`). Once ESH consumes colo
|
||
services over the mesh (100.64/10 via ana-scale) rather than the public FortiGate VIP,
|
||
that traffic never reaches CrowdSec — a concrete win beyond replacing the tunnels, worth
|
||
weighing when prioritising the cut-over.
|
||
|
||
### 06:35Z — ESH egress whitelisted in CrowdSec (operator-directed, temporary)
|
||
|
||
`/opt/docker/conf/crowdsec/postoverflows/s01-whitelist/pfi-esh-egress.yaml` (bind mount,
|
||
persists) whitelists 23.164.40.160; SIGHUP reload, `crowdsec -t` clean, parser loaded.
|
||
TEMPORARY — remove when ESH gets its static IP. crowdsec stack is NOT in stacks/ canonical
|
||
(only stacks-mirror), so this lives on the host; re-mirror with sync-stacks.sh.
|
||
|
||
## CUTOVER EXECUTED 2026-09-06 — IPsec dormant, mesh primary; Site Magic pending operator
|
||
|
||
Operator goal: replace Site Magic + IPsec with headscale, tunnels dormant as backup.
|
||
|
||
**Safety net first:** FortiGate WAN SSH enabled — `wan1 set allowaccess ping ssh`, admin
|
||
`infra-ops` trusthost2/3 = 70.230.226.88/32 (NH3 egress) + 23.164.40.160/32 (ESH egress);
|
||
trusthost1 10.0.0.0/8 kept. Config backed up to flash `pre-wan-ssh-cutover-20260906`.
|
||
Verified: `ssh infra-ops@38.120.12.42` works from nh3-dev, tunnel-independent. ⚠ TEMPORARY —
|
||
scoped to two source IPs, not 0.0.0.0; remove after the edge is retired. Drive the FortiGate
|
||
via WAN SSH (`scratchpad/fgw.py`) during/after cutover, not over the tunnel.
|
||
|
||
**Method (learned the hard way):** the tunnel and a mesh static route for the same /16 on the
|
||
same gateway FIGHT — forward/return split across transports, asymmetric traffic dropped
|
||
(broke colo from NH3 at 06:xx while both were up; rolled back). Correct order per pair:
|
||
disable the tunnel end(s) FIRST (route withdraws), THEN add the mesh /16 route (now
|
||
unambiguous). Never do it with the tunnel up.
|
||
|
||
**colo↔NH3 (IPsec):** NH3 UDM `pfi-nh3-ana` disabled (enabled=false, API); mesh routes
|
||
NH3 `10.250.0.0/16→10.100.50.46 (nh3-scale)`, FortiGate `10.100.0.0/16→10.250.50.45
|
||
(ana-scale)`. Verified: nh3-dev→ana-docker + pfi-pve OPEN; FortiGate reachable via mesh AND
|
||
WAN SSH; ESH→rest-server-ana (restic) OPEN. FortiGate phase1 left enabled (no peer → idle);
|
||
re-enable NH3 side to restore.
|
||
|
||
**colo↔ESH (IPsec):** ESH UDM `esh-ana` disabled; mesh routes ESH
|
||
`10.250.0.0/16→10.0.50.65 (esh-scale)`, FortiGate `10.0.0.0/16→10.250.50.45`. Verified
|
||
bidirectional: esh-docker-vm↔ana-docker/pfi-pve OPEN both ways. Proves ESH (CGNAT) mesh over
|
||
public/DERP.
|
||
|
||
**NH3↔ESH (Site Magic) — NOT done by API.** `sdwan-mesh-tunnel` objects (`ESH-UDMPM` on NH3,
|
||
`PFI-UDMSE` on ESH) return `api.err.NoEdit` — Site Magic is cloud-orchestrated, not togglable
|
||
via the classic REST API. Mesh routes PRE-STAGED (NH3 `10.0.0.0/16→nh3-scale`, ESH
|
||
`10.100.0.0/16→esh-scale`), currently SHADOWED by Site Magic (traffic still on the SD-WAN
|
||
tunnel, traceroute via 192.168.1.x transit). esh-scale↔nh3-scale hold a DERP(lax) 9ms path
|
||
independent of Site Magic → when the operator disables Site Magic in the UniFi UI / Site
|
||
Manager, the staged routes take over immediately. **Operator to toggle; mesh verified ready.**
|
||
|
||
**Dormancy = disabled-but-retained**, not auto-failover (overlapping /16s across two
|
||
transports reintroduce the asymmetry). Re-enable a pair = flip the UDM object back to
|
||
enabled=true (+ re-enable Site Magic in UI). No automatic mesh→tunnel failover is wired.
|
||
|
||
**Bonus (operator noted):** exit nodes at each site give a free multi-location egress proxy
|
||
(`tailscale up --advertise-exit-node` + approve); replaces the nh3-dev SOCKS5 residential
|
||
egress with a native feature. Not yet configured — parked follow-up.
|
||
|
||
**Follow-ups:** operator disables Site Magic; (optional) disable FortiGate phase1 ends for
|
||
tidiness; advertise exit nodes; self-hosted DERP at ESH on the 2G circuit; remove the
|
||
FortiGate WAN-SSH trusthosts when the edge is retired; PVE 8.4 upgrade on pfi-pve/nh3-pve.
|
||
|
||
## ✅ CUTOVER COMPLETE 2026-09-06 — all three site-pairs on the mesh
|
||
|
||
Operator disabled Site Magic in the UniFi UI. NH3↔ESH re-homed to the mesh immediately —
|
||
traceroute now esh-scale(10.0.50.65)→nh3-scale(100.64.0.1)→dest, no 192.168.1.x SD-WAN
|
||
transit; esh-scale↔nh3-scale hole-punched a DIRECT public path (70.230.226.88:41641, 8ms),
|
||
not even DERP. Full matrix verified, all six directions OPEN:
|
||
NH3↔colo, NH3↔ESH, colo↔ESH. FortiGate mgmt via WAN SSH OPEN; ana-wg WG fallback untouched.
|
||
|
||
State: **Site Magic disabled (UI); both IPsec tunnels dormant (UDM ends disabled, config
|
||
retained); headscale is the sole active site-to-site transport.** Tunnels re-enablable for
|
||
backup (Site Magic in UI; IPsec by flipping the UDM objects enabled=true). Goal met.
|
||
|
||
## 2026-09-06 — exit nodes advertised (multi-location egress proxy)
|
||
|
||
All three routers advertise + serve exit nodes; approved on headscale (each shows
|
||
`0.0.0.0/0` + `::/0` Serving). Clients see three: `nh3-scale` (100.64.0.1, NH3 residential
|
||
egress 70.230.226.88), `esh-scale` (100.64.0.2, ESH 23.164.40.160), `ana-scale`
|
||
(100.64.0.3, colo 38.120.12.42).
|
||
|
||
⚠ **Design reversal:** exit nodes REQUIRE SNAT, so `--snat-subnet-routes` was flipped back to
|
||
**true** (NoSNAT=false) on all three routers — this undoes the earlier source-preserving
|
||
choice. Consequence: cross-site subnet-route traffic now appears to come from the local
|
||
router's IP, not the original host. Site-to-site re-verified OK after the flip. If per-host
|
||
source visibility is later needed, the alternative is NoSNAT + a manual selective masquerade
|
||
rule (`-s 100.64.0.0/10 ! -d 10.0.0.0/8 -o eth0 -j MASQUERADE`) on each router, made
|
||
persistent — not done.
|
||
|
||
**Use one from a client:** `tailscale set --exit-node=nh3-scale` (or `esh-scale` /
|
||
`ana-scale`), `--exit-node=` to clear. A node that itself advertises an exit node cannot
|
||
consume one (so the routers can't be used to test egress; use the laptop/iPad).
|
||
Replaces the nh3-dev SOCKS5 residential-egress proxy — retire that once confirmed on real
|
||
clients.
|
||
|
||
## 2026-09-06 — selective masquerade applied (supersedes the SNAT-reversal note above)
|
||
|
||
Restored `--snat-subnet-routes=false` (NoSNAT=true) on all three routers AND added a
|
||
persistent selective-masquerade rule, so BOTH properties now hold: exit nodes work (internet
|
||
egress masqueraded) and site-to-site keeps the original host source.
|
||
|
||
Per router: `/usr/local/sbin/mesh-exit-masq.sh` + `mesh-exit-masq.service` (oneshot, After
|
||
tailscaled, enabled). Rule = a `MESH-EXIT` nat chain jumped from POSTROUTING for
|
||
`-s 100.64.0.0/10 -o eth0`: RETURN for 10/8, 172.16/12, 192.168/16, 100.64/10 (source
|
||
preserved), MASQUERADE the rest (real internet). ⚠ nf_tables iptables forbids multiple `-d`
|
||
in one rule — hence the chain-with-RETURNs shape, not a single `! -d ... ! -d ...` rule.
|
||
|
||
Verified: nh3-docker→ana-docker seen at colo as **10.100.50.40** (real host, preserved);
|
||
nh3-dev via ana-scale exit → egress **38.120.12.42** (colo), via nh3-scale → 70.230.226.88
|
||
(NH3). Persistent across CT reboot via the systemd unit. `tailscale set --operator` NOT set;
|
||
drive nh3-dev's tailscale via `ssh infra-ops@10.100.10.50 sudo tailscale ...`.
|
||
|
||
## 2026-09-06 — throughput baseline (mesh vs the old IPsec)
|
||
|
||
NH3↔colo over the direct WireGuard path (5ms), iperf3 between the 1-vCPU router CTs, medians
|
||
of 3: **download colo→NH3 761 Mb/s (688–761), upload NH3→colo 464 Mb/s (391–507)**, parallel-8
|
||
up ~530. Old IPsec ceiling was ~250 both ways (poor AES offload) — so ~3x down, ~1.8x up. CPU
|
||
during upload sampled ~32% (single reading) → not clearly the cap; upload likely limited by
|
||
userspace-WireGuard single-flow, the CT vNIC, or the WAN. Operator: the ANA AT&T line is
|
||
SHARED fiber, so the upload figure may be the line, not the routers — left as-is (no CT core
|
||
bump). Baseline for future comparison, not a hardware limit.
|
||
|
||
## 2026-09-06 — irv-ml1 folded in ADDITIVELY (operator inaccessible 5 days — lifeline preserved)
|
||
|
||
Constraint: operator has NO physical access to the Irvine building for 5 days; irv-ml1 must
|
||
not go dark unrecoverably. Its ONLY remote path is the bespoke WireGuard tunnel `wg0`
|
||
(10.100.79.3 → nh3-dev, AllowedIPs 10.0/16+10.100/16+10.250/16); its LAN 10.6.110.50 is not
|
||
reachable from the fleet.
|
||
|
||
Enrolled irv-ml1 as headscale node **100.64.0.6** (`infra-ops@10.100.79.3` has NOPASSWD root;
|
||
`tailscaled` was already running but logged out with STALE prefs → **public** Tailscale
|
||
`controlplane.tailscale.com` + `RouteAll:true`). Used `tailscale up --reset --login-server
|
||
https://headscale.phasefinal.com --authkey … --accept-routes=false --accept-dns=false
|
||
--hostname irv-ml1`. ⚠ **accept-routes MUST stay OFF**: wg0's AllowedIPs are the exact /16s the
|
||
mesh routers advertise, so accepting them would fight the lifeline over the prefix my SSH
|
||
returns on = 5-day lockout. Armed an auto-revert dead-man switch (`tailscale down` in 10m
|
||
unless disarmed) before the change; disarmed via sentinel after verifying.
|
||
|
||
Result: **DUAL PATH** — wg0 (10.100.79.3, untouched, fleet routes still via it) AND mesh
|
||
(100.64.0.6). Both verified from nh3-dev. Reboot-safe: `wg-quick@wg0` + `tailscaled` both
|
||
enabled, prefs now headscale + accept-routes=false + WantRunning=true. Mesh peering currently
|
||
rides over wg0 as transport (11ms).
|
||
|
||
⚠ **DEFERRED to the on-site window (do NOT do remotely):** retire wg0; enable accept-routes /
|
||
subnet-routing over the mesh; migrate yt-voice-clipper egress off the dante SOCKS proxy. Each
|
||
touches the exact routing that is the lifeline. When on-site (physical console as fallback):
|
||
flip accept-routes on, confirm fleet reachability over the mesh, THEN down wg0.
|
||
|
||
⚠ **Lesson (repeat):** `pkill -f` / `pgrep -f <pattern>` over SSH self-matches when the pattern
|
||
is in your own ssh argv — killed my session twice disarming the switch. Kill by explicit PID,
|
||
or make the switch sentinel-based (which saved it here).
|
||
|
||
## 2026-09-06 — irv-ml1 CUTOVER COMPLETE + dante retired (done remotely, 5-day no-access)
|
||
|
||
Operator wanted the cutover finished tonight despite no physical Irvine access for 5 days, and
|
||
directed: dante proxy goes away (exit nodes instead), yt-voice-clipper may stay dead until
|
||
scoped exit-node egress is wired.
|
||
|
||
**Three independent recovery paths built BEFORE cutting wg0** (the discipline that made it safe):
|
||
1. Mesh over DERP-LA (~12ms) — irv-ml1's own WAN (38.120.94.3, easy NAT), independent of wg0.
|
||
2. **Reverse SSH tunnel** `revtun-nh3.service` on irv-ml1 → nh3-dev via NH3 public WAN
|
||
(UDM forward `irv-revtun-ssh` tcp/47822 → 10.100.10.50:22, **src-restricted to 38.120.94.3**;
|
||
key `/root/.ssh/revtun_ed25519`, authorized on nh3-dev lkraven with `restrict,port-forwarding`).
|
||
Reach irv-ml1: `ssh -i ~/.ssh/infra-ops_ed25519 -p 2201 infra-ops@127.0.0.1` on nh3-dev.
|
||
Independent of BOTH wg0 and the mesh/control-plane.
|
||
3. **`wg0-watchdog.service`** on irv-ml1 — `wg-quick up wg0` if the fleet is unreachable over the
|
||
mesh for ~5min (auto-failover). Plus wg-quick@wg0 stays enabled → reboot restores wg0.
|
||
|
||
**Cut:** accept-routes=true on irv-ml1 (fleet /16s → tailscale0 table 52), then `wg-quick down
|
||
wg0` over the reverse tunnel under a 12-min auto-revert dead-man (disarmed after verify).
|
||
Verified: irv-ml1 reaches nh3-docker/colo/ESH/nh3-pve over pure mesh; unreachable ONLY 10.100.10.50
|
||
(nh3-dev's own LAN IP — a mesh-node-reaching-a-mesh-node quirk; every other host fine; moot now
|
||
dante is gone). wg0 dormant (down, enabled). Recovery via reverse tunnel + mesh IP both OK.
|
||
|
||
**Dante retired** on nh3-dev: `danted` disabled+stopped, :1080 closed, `/etc/danted.conf` →
|
||
`.retired-20260906`. Only live consumer was YTVC (operator-exempted). searxng ref was already
|
||
commented out.
|
||
|
||
**OPEN follow-ups (on-site or later):** wire yt-voice-clipper egress via a SCOPED exit-node path
|
||
(do NOT set irv-ml1 `--exit-node` globally — it would route the reverse tunnel through the mesh and
|
||
kill the independent lifeline; use tailscale `--socks5-server`+exit-node or a per-container netns);
|
||
then bring YTVC back. On-site: decide whether to keep the reverse tunnel + UDM forward + revtun key
|
||
or remove them (small WAN attack surface, src-restricted). Fully retire wg0 config if desired.
|
||
|
||
## 2026-09-06 — wg0 made non-reboot-restorable + Irvine subnet discovery (operator corrections)
|
||
|
||
**wg0 no longer reboot-restorable** (supersedes the earlier "wg-quick@wg0 stays enabled" note):
|
||
`systemctl disable wg-quick@wg0` on irv-ml1 — the cutover is real, wg0 does not silently return on
|
||
reboot. Config `/etc/wireguard/wg0.conf` retained so the failover paths can still bring it up. The
|
||
FAILOVER is now: `wg0-watchdog.service` (`wg-quick up wg0` if the fleet is unreachable over the mesh
|
||
~5min — works regardless of the unit's enable state) + the independent reverse tunnel. Rationale
|
||
(operator): reboot-restore was redundant with the watchdog and made the cutover half-real; accept-routes
|
||
table-52 would keep the mesh primary anyway, but a clean cutover shouldn't resurrect the retired tunnel.
|
||
|
||
**Irvine subnet discovery LIVE (bidirectional):** irv-ml1 now `--advertise-routes=10.6.110.0/24`
|
||
(approved on headscale, serving; ip_forward=1; default SNAT-on so mesh-unaware Irvine hosts reply to
|
||
irv-ml1). Fleet-wide reach added via gateway static routes `10.6.110.0/24 → local mesh router`
|
||
(NH3 UDM→nh3-scale, ESH UDM→esh-scale, FortiGate route 22→ana-scale). Verified: ana-docker→10.6.110.1
|
||
and esh-docker-vm→10.6.110.50 both OK over the mesh. So irv-ml1 both REACHES the fleet /16s
|
||
(accept-routes) and EXPOSES its Irvine LAN to the fleet (advertise-routes) — a full mesh subnet router.
|