feat(cira-tunnel-watchdog): auto-drop stale AMT CIRA tunnels on MeshCentral's MPS

This commit is contained in:
vh
2026-10-03 12:34:35 -07:00
parent fe1d472aae
commit 00a6fbd849
9 changed files with 161 additions and 0 deletions
+36
View File
@@ -0,0 +1,36 @@
# cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)
**Fixes:** MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at **"Setup…"** forever.
**Why it happens.** A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new
CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay
(`webserver.js` `handleRelayWebSocket`) can pick the dead one. A TLS failure there is only debug-logged and never
closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because
MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites:
esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up:
`servers/pfi-tacticalrmm/README.md`.
**What it does.** A systemd timer on pfi-tacticalrmm runs `/usr/local/sbin/cira-tunnel-watchdog` every minute.
Any ESTABLISHED socket on :4433 that has received nothing for **180 s** (`CIRA_MAX_SILENT_MS`) is destroyed with
`ss -K` and logged (`journalctl -t cira-tunnel-watchdog`). MeshCentral then drops it, and the live tunnel serves.
- **Margin:** healthy tunnels showed `lastrcv` of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s.
- **A wrong drop costs seconds:** the AMT policy reconnects every 10 s, and a reconnect of a known device is the
safe code path (only a device's FIRST connect hits the MeshCentral crash race).
**Install / update:** `scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml`
(installed 2026-10-03 1232, Prime's go).
**Tests (controls, run before install):**
```bash
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
# → nothing null control
```
On the live host, `CIRA_MAX_SILENT_MS=1 … --dry-run` must list every tunnel (proves the parser sees them), and the
kill filter `( sport = :4433 ) and dst <peer>` must select exactly one socket. Both were checked at install.
**What it does not cover:** an AMT with no link at all (cable out, or Linux has downed a shared igc port). There
is then no tunnel to fix, and the device shows offline in MeshCentral.
Uninstall: `systemctl disable --now cira-tunnel-watchdog.timer`, then remove the two units and the script.
+56
View File
@@ -0,0 +1,56 @@
#!/bin/bash
# cira-tunnel-watchdog: drop half-open Intel AMT phone-home (CIRA) tunnels on MeshCentral's MPS.
#
# Why (servers/pfi-tacticalrmm/README.md): when a host whose AMT shares the network chip power-cycles, AMT opens
# a new CIRA tunnel without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay (KVM "HW Connect",
# SOL, IDE-R, MeshCommander) can pick the dead one and hang at "Setup..." forever. Its own 90 s idle timeout
# does not fire, because its writes to the dead socket keep resetting it. Seen 3 of 3 power events, at two sites
# (2026-10-02/03).
#
# What: every run, any ESTABLISHED socket on the MPS port that has received nothing for MAX_SILENT_MS is
# destroyed with `ss -K`. MeshCentral then drops it, and the live tunnel serves. Healthy tunnels hear from their
# AMT every ~6-30 s (measured), so the 180 s default leaves a wide margin. A tunnel dropped by mistake comes
# back on its own: the AMT policy reconnects every 10 s.
#
# cira-tunnel-watchdog act (root: ss -K needs CAP_NET_ADMIN)
# cira-tunnel-watchdog --dry-run report what would be dropped
# cira-tunnel-watchdog --dry-run --from-file F decide from a saved `ss -tnio` capture (tests)
set -euo pipefail
PORT="${CIRA_PORT:-4433}"
MAX_SILENT_MS="${CIRA_MAX_SILENT_MS:-180000}"
DRY=0
SRC=""
while [ $# -gt 0 ]; do
case "$1" in
--dry-run) DRY=1 ;;
--from-file) SRC="${2:?--from-file needs a path}"; shift ;;
*) echo "usage: $0 [--dry-run] [--from-file <ss capture>]" >&2; exit 2 ;;
esac
shift
done
if [ -n "$SRC" ]; then
[ "$DRY" = 1 ] || { echo "--from-file is a test mode and needs --dry-run" >&2; exit 2; }
out="$(cat "$SRC")"
else
out="$(ss -tnio state established "( sport = :$PORT )")"
fi
# `ss -tnio state established` prints a socket line (Recv-Q Send-Q Local Peer) followed by an indented info line.
stale="$(printf '%s\n' "$out" | awk -v max="$MAX_SILENT_MS" '
/^[0-9]/ { peer = $4; next }
/lastrcv:/ {
lr = -1
for (i = 1; i <= NF; i++) if ($i ~ /^lastrcv:/) { split($i, a, ":"); lr = a[2] + 0 }
if (peer != "" && lr > max) print peer, lr
peer = ""
}')"
[ -z "$stale" ] && exit 0
while read -r peer lr; do
[[ "$peer" =~ ^\[[0-9a-fA-F:.]+\]:[0-9]+$ ]] || { echo "skip unparsable peer [$peer]" >&2; continue; }
if [ "$DRY" = 1 ]; then
echo "would drop $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
else
ss -K -tn "( sport = :$PORT ) and dst $peer" >/dev/null
logger -t cira-tunnel-watchdog "dropped stale CIRA tunnel $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
echo "dropped $peer (silent ${lr} ms)"
fi
done <<<"$stale"
@@ -0,0 +1,8 @@
# Drops half-open Intel AMT CIRA tunnels on MeshCentral's MPS (:4433). See README.md beside this file.
[Unit]
Description=Drop stale Intel AMT CIRA tunnels on MeshCentral's MPS
After=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/cira-tunnel-watchdog
@@ -0,0 +1,10 @@
[Unit]
Description=Every minute: drop stale Intel AMT CIRA tunnels
[Timer]
OnBootSec=2min
OnUnitActiveSec=1min
AccuracySec=10s
[Install]
WantedBy=timers.target
@@ -0,0 +1,5 @@
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 cwnd:10 lastsnd:22588 lastrcv:19332 lastack:22584
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:16994
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 cwnd:10 lastsnd:16360 lastrcv:30136 lastack:16356
@@ -0,0 +1,7 @@
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 pmtu:1500 rcvmss:838 advmss:1460 cwnd:10 bytes_sent:71724 bytes_received:372619 lastsnd:22588 lastrcv:22588 lastack:22584 app_limited busy:3684ms
0 0 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16992
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 pmtu:1500 cwnd:10 bytes_sent:49346 bytes_received:77389 lastsnd:16360 lastrcv:16360 lastack:16356 app_limited busy:632ms
0 2989 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16995 timer:(on,1min48sec,14)
cubic wscale:0,7 rto:120000 backoff:14 rtt:4.056/0.136 ato:40 mss:1460 pmtu:1500 cwnd:1 bytes_sent:53105 bytes_retrans:7350 lastsnd:11780 lastrcv:782484 lastack:782480 unacked:1 retrans:1/14 lost:1 notsent:2464