Files
esh-pfi-infrastructure/services/cira-tunnel-watchdog/README.md
T

37 lines
2.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)
**Fixes:** MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at **"Setup…"** forever.
**Why it happens.** A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new
CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay
(`webserver.js` `handleRelayWebSocket`) can pick the dead one. A TLS failure there is only debug-logged and never
closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because
MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites:
esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up:
`servers/pfi-tacticalrmm/README.md`.
**What it does.** A systemd timer on pfi-tacticalrmm runs `/usr/local/sbin/cira-tunnel-watchdog` every minute.
Any ESTABLISHED socket on :4433 that has received nothing for **180 s** (`CIRA_MAX_SILENT_MS`) is destroyed with
`ss -K` and logged (`journalctl -t cira-tunnel-watchdog`). MeshCentral then drops it, and the live tunnel serves.
- **Margin:** healthy tunnels showed `lastrcv` of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s.
- **A wrong drop costs seconds:** the AMT policy reconnects every 10 s, and a reconnect of a known device is the
safe code path (only a device's FIRST connect hits the MeshCentral crash race).
**Install / update:** `scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml`
(installed 2026-10-03 1232, Prime's go).
**Tests (controls, run before install):**
```bash
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
# → nothing null control
```
On the live host, `CIRA_MAX_SILENT_MS=1 … --dry-run` must list every tunnel (proves the parser sees them), and the
kill filter `( sport = :4433 ) and dst <peer>` must select exactly one socket. Both were checked at install.
**What it does not cover:** an AMT with no link at all (cable out, or Linux has downed a shared igc port). There
is then no tunnel to fix, and the device shows offline in MeshCentral.
Uninstall: `systemctl disable --now cira-tunnel-watchdog.timer`, then remove the two units and the script.