ops(meshcentral): diagnose and fix HW Connect stuck at Setup (stale CIRA tunnel); add relay probe

This commit is contained in:
vh
2026-10-02 22:49:27 -07:00
parent 55004090d1
commit ad0322bf83
3 changed files with 64 additions and 0 deletions
+12
View File
@@ -54,6 +54,18 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc.
has none and its two first connects did not crash. That is inferred, not measured. Reconnects of
an existing device take the other branch and are safe. **Onboard new AMTs at a quiet hour**; expect one ~7 s
MeshCentral restart.
- ⚠ **"HW Connect" stuck at "Setup…" = a stale CIRA tunnel** (2026-10-02 2247, esh-pve-2-amt; Prime power-cycled
the host several times). The AMT opened a new tunnel without closing the old one. MeshCentral then held two:
the live one (power actions and the AMT manager worked through it) and a dead one, `backoff:14`, nothing
received for 782 s. `webrelay.ashx` (KVM, SOL, IDE-R, MeshCommander) picked the dead one. A relay TLS failure
is only logged at debug level and never closes the websocket, so the browser waits forever. MPS's 90 s idle
timeout did not fire, because MeshCentral's own writes to the dead socket kept resetting it. TCP gives up on
its own after ~15–30 min of retransmits.
**Fix:** `sudo ss -tnio state established "( sport = :4433 )"`. Find the device's public IP with a large
`lastrcv` (ms) or a `backoff`. Healthy tunnels show `lastrcv` of seconds. Then
`sudo ss -K -tn "dst [::ffff:<ip>]:<its port>"`. Verify with `scripts/meshcentral-amt-relay-probe.js`, using
another AMT as a positive control. Likely whenever a host whose AMT shares its NIC (esh-pve-2) power-cycles.
Seen once so far.
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's