feat(cira-tunnel-watchdog): auto-drop stale AMT CIRA tunnels on MeshCentral's MPS

This commit is contained in:
vh
2026-10-03 12:34:35 -07:00
parent fe1d472aae
commit 00a6fbd849
9 changed files with 161 additions and 0 deletions
+1
View File
@@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
+36
View File
@@ -0,0 +1,36 @@
# Install the CIRA tunnel watchdog on pfi-tacticalrmm (MeshCentral's host). Idempotent.
# scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml
# Source + rationale: services/cira-tunnel-watchdog/.
steps:
- name: script
sudo: true
upload:
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog
dest: /usr/local/sbin/cira-tunnel-watchdog
mode: "0755"
- name: service unit
sudo: true
upload:
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.service
dest: /etc/systemd/system/cira-tunnel-watchdog.service
mode: "0644"
- name: timer unit
sudo: true
upload:
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer
dest: /etc/systemd/system/cira-tunnel-watchdog.timer
mode: "0644"
- name: reload and enable the timer
sudo: true
shell: systemctl daemon-reload && systemctl enable --now cira-tunnel-watchdog.timer
changed_when: "false"
verify:
- name: timer is active and scheduled
sudo: true
shell: systemctl is-active --quiet cira-tunnel-watchdog.timer && systemctl list-timers cira-tunnel-watchdog.timer --no-pager | grep -q cira-tunnel-watchdog
changed_when: "false"
- name: the live MPS capture parses (dry run exits 0)
sudo: true
shell: /usr/local/sbin/cira-tunnel-watchdog --dry-run
changed_when: "false"
+2
View File
@@ -68,6 +68,8 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc.
tunnel dead within seconds, and it was gone by 2252. The first one lingered 13+ min, probably because repeated
connect attempts kept writing to it (inferred). **Third time 2026-10-03 1218, on nh3-pve-2** (NH3, Prime power-cycling it
on site), so it is not ESH- or NAT-specific: 3 of 3 power-event episodes, at two sites.
**Automated since 2026-10-03 1232 (Prime's go): `cira-tunnel-watchdog` timer on this host drops any :4433 tunnel
silent >180 s, every minute** (`services/cira-tunnel-watchdog/`, `journalctl -t cira-tunnel-watchdog`).
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's
+36
View File
@@ -0,0 +1,36 @@
# cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)
**Fixes:** MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at **"Setup…"** forever.
**Why it happens.** A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new
CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay
(`webserver.js` `handleRelayWebSocket`) can pick the dead one. A TLS failure there is only debug-logged and never
closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because
MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites:
esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up:
`servers/pfi-tacticalrmm/README.md`.
**What it does.** A systemd timer on pfi-tacticalrmm runs `/usr/local/sbin/cira-tunnel-watchdog` every minute.
Any ESTABLISHED socket on :4433 that has received nothing for **180 s** (`CIRA_MAX_SILENT_MS`) is destroyed with
`ss -K` and logged (`journalctl -t cira-tunnel-watchdog`). MeshCentral then drops it, and the live tunnel serves.
- **Margin:** healthy tunnels showed `lastrcv` of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s.
- **A wrong drop costs seconds:** the AMT policy reconnects every 10 s, and a reconnect of a known device is the
safe code path (only a device's FIRST connect hits the MeshCentral crash race).
**Install / update:** `scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml`
(installed 2026-10-03 1232, Prime's go).
**Tests (controls, run before install):**
```bash
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
# → nothing null control
```
On the live host, `CIRA_MAX_SILENT_MS=1 … --dry-run` must list every tunnel (proves the parser sees them), and the
kill filter `( sport = :4433 ) and dst <peer>` must select exactly one socket. Both were checked at install.
**What it does not cover:** an AMT with no link at all (cable out, or Linux has downed a shared igc port). There
is then no tunnel to fix, and the device shows offline in MeshCentral.
Uninstall: `systemctl disable --now cira-tunnel-watchdog.timer`, then remove the two units and the script.
+56
View File
@@ -0,0 +1,56 @@
#!/bin/bash
# cira-tunnel-watchdog: drop half-open Intel AMT phone-home (CIRA) tunnels on MeshCentral's MPS.
#
# Why (servers/pfi-tacticalrmm/README.md): when a host whose AMT shares the network chip power-cycles, AMT opens
# a new CIRA tunnel without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay (KVM "HW Connect",
# SOL, IDE-R, MeshCommander) can pick the dead one and hang at "Setup..." forever. Its own 90 s idle timeout
# does not fire, because its writes to the dead socket keep resetting it. Seen 3 of 3 power events, at two sites
# (2026-10-02/03).
#
# What: every run, any ESTABLISHED socket on the MPS port that has received nothing for MAX_SILENT_MS is
# destroyed with `ss -K`. MeshCentral then drops it, and the live tunnel serves. Healthy tunnels hear from their
# AMT every ~6-30 s (measured), so the 180 s default leaves a wide margin. A tunnel dropped by mistake comes
# back on its own: the AMT policy reconnects every 10 s.
#
# cira-tunnel-watchdog act (root: ss -K needs CAP_NET_ADMIN)
# cira-tunnel-watchdog --dry-run report what would be dropped
# cira-tunnel-watchdog --dry-run --from-file F decide from a saved `ss -tnio` capture (tests)
set -euo pipefail
PORT="${CIRA_PORT:-4433}"
MAX_SILENT_MS="${CIRA_MAX_SILENT_MS:-180000}"
DRY=0
SRC=""
while [ $# -gt 0 ]; do
case "$1" in
--dry-run) DRY=1 ;;
--from-file) SRC="${2:?--from-file needs a path}"; shift ;;
*) echo "usage: $0 [--dry-run] [--from-file <ss capture>]" >&2; exit 2 ;;
esac
shift
done
if [ -n "$SRC" ]; then
[ "$DRY" = 1 ] || { echo "--from-file is a test mode and needs --dry-run" >&2; exit 2; }
out="$(cat "$SRC")"
else
out="$(ss -tnio state established "( sport = :$PORT )")"
fi
# `ss -tnio state established` prints a socket line (Recv-Q Send-Q Local Peer) followed by an indented info line.
stale="$(printf '%s\n' "$out" | awk -v max="$MAX_SILENT_MS" '
/^[0-9]/ { peer = $4; next }
/lastrcv:/ {
lr = -1
for (i = 1; i <= NF; i++) if ($i ~ /^lastrcv:/) { split($i, a, ":"); lr = a[2] + 0 }
if (peer != "" && lr > max) print peer, lr
peer = ""
}')"
[ -z "$stale" ] && exit 0
while read -r peer lr; do
[[ "$peer" =~ ^\[[0-9a-fA-F:.]+\]:[0-9]+$ ]] || { echo "skip unparsable peer [$peer]" >&2; continue; }
if [ "$DRY" = 1 ]; then
echo "would drop $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
else
ss -K -tn "( sport = :$PORT ) and dst $peer" >/dev/null
logger -t cira-tunnel-watchdog "dropped stale CIRA tunnel $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
echo "dropped $peer (silent ${lr} ms)"
fi
done <<<"$stale"
@@ -0,0 +1,8 @@
# Drops half-open Intel AMT CIRA tunnels on MeshCentral's MPS (:4433). See README.md beside this file.
[Unit]
Description=Drop stale Intel AMT CIRA tunnels on MeshCentral's MPS
After=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/cira-tunnel-watchdog
@@ -0,0 +1,10 @@
[Unit]
Description=Every minute: drop stale Intel AMT CIRA tunnels
[Timer]
OnBootSec=2min
OnUnitActiveSec=1min
AccuracySec=10s
[Install]
WantedBy=timers.target
@@ -0,0 +1,5 @@
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 cwnd:10 lastsnd:22588 lastrcv:19332 lastack:22584
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:16994
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 cwnd:10 lastsnd:16360 lastrcv:30136 lastack:16356
@@ -0,0 +1,7 @@
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 pmtu:1500 rcvmss:838 advmss:1460 cwnd:10 bytes_sent:71724 bytes_received:372619 lastsnd:22588 lastrcv:22588 lastack:22584 app_limited busy:3684ms
0 0 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16992
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 pmtu:1500 cwnd:10 bytes_sent:49346 bytes_received:77389 lastsnd:16360 lastrcv:16360 lastack:16356 app_limited busy:632ms
0 2989 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16995 timer:(on,1min48sec,14)
cubic wscale:0,7 rto:120000 backoff:14 rtt:4.056/0.136 ato:40 mss:1460 pmtu:1500 cwnd:1 bytes_sent:53105 bytes_retrans:7350 lastsnd:11780 lastrcv:782484 lastack:782480 unacked:1 retrans:1/14 lost:1 notsent:2464