Files
esh-pfi-infrastructure/services/cira-tunnel-watchdog/README.md
T

2.7 KiB
Raw Blame History

cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)

Fixes: MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at "Setup…" forever.

Why it happens. A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay (webserver.js handleRelayWebSocket) can pick the dead one. A TLS failure there is only debug-logged and never closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites: esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up: servers/pfi-tacticalrmm/README.md.

What it does. A systemd timer on pfi-tacticalrmm runs /usr/local/sbin/cira-tunnel-watchdog every minute. Any ESTABLISHED socket on :4433 that has received nothing for 180 s (CIRA_MAX_SILENT_MS) is destroyed with ss -K and logged (journalctl -t cira-tunnel-watchdog). MeshCentral then drops it, and the live tunnel serves.

  • Margin: healthy tunnels showed lastrcv of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s.
  • A wrong drop costs seconds: the AMT policy reconnects every 10 s, and a reconnect of a known device is the safe code path (only a device's FIRST connect hits the MeshCentral crash race).

Install / update: scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml (installed 2026-10-03 1232, Prime's go).

Tests (controls, run before install):

services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
#   → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel)      positive control
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
#   → nothing                                                                            null control

On the live host, CIRA_MAX_SILENT_MS=1 … --dry-run must list every tunnel (proves the parser sees them), and the kill filter ( sport = :4433 ) and dst <peer> must select exactly one socket. Both were checked at install.

What it does not cover: an AMT with no link at all (cable out, or Linux has downed a shared igc port). There is then no tunnel to fix, and the device shows offline in MeshCentral.

Uninstall: systemctl disable --now cira-tunnel-watchdog.timer, then remove the two units and the script.