2.7 KiB
cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)
Fixes: MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at "Setup…" forever.
Why it happens. A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new
CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay
(webserver.js handleRelayWebSocket) can pick the dead one. A TLS failure there is only debug-logged and never
closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because
MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites:
esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up:
servers/pfi-tacticalrmm/README.md.
What it does. A systemd timer on pfi-tacticalrmm runs /usr/local/sbin/cira-tunnel-watchdog every minute.
Any ESTABLISHED socket on :4433 that has received nothing for 180 s (CIRA_MAX_SILENT_MS) is destroyed with
ss -K and logged (journalctl -t cira-tunnel-watchdog). MeshCentral then drops it, and the live tunnel serves.
- Margin: healthy tunnels showed
lastrcvof 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s. - A wrong drop costs seconds: the AMT policy reconnects every 10 s, and a reconnect of a known device is the safe code path (only a device's FIRST connect hits the MeshCentral crash race).
Install / update: scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml
(installed 2026-10-03 1232, Prime's go).
Tests (controls, run before install):
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
# → nothing null control
On the live host, CIRA_MAX_SILENT_MS=1 … --dry-run must list every tunnel (proves the parser sees them), and the
kill filter ( sport = :4433 ) and dst <peer> must select exactly one socket. Both were checked at install.
What it does not cover: an AMT with no link at all (cable out, or Linux has downed a shared igc port). There is then no tunnel to fix, and the device shows offline in MeshCentral.
Uninstall: systemctl disable --now cira-tunnel-watchdog.timer, then remove the two units and the script.