feat(cira-tunnel-watchdog): auto-drop stale AMT CIRA tunnels on MeshCentral's MPS
This commit is contained in:
@@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._
|
|||||||
|
|
||||||
## Recent decisions
|
## Recent decisions
|
||||||
|
|
||||||
|
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
|
||||||
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
|
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
|
||||||
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
|
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
|
||||||
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
|
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
|
||||||
|
|||||||
@@ -0,0 +1,36 @@
|
|||||||
|
# Install the CIRA tunnel watchdog on pfi-tacticalrmm (MeshCentral's host). Idempotent.
|
||||||
|
# scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml
|
||||||
|
# Source + rationale: services/cira-tunnel-watchdog/.
|
||||||
|
steps:
|
||||||
|
- name: script
|
||||||
|
sudo: true
|
||||||
|
upload:
|
||||||
|
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog
|
||||||
|
dest: /usr/local/sbin/cira-tunnel-watchdog
|
||||||
|
mode: "0755"
|
||||||
|
- name: service unit
|
||||||
|
sudo: true
|
||||||
|
upload:
|
||||||
|
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.service
|
||||||
|
dest: /etc/systemd/system/cira-tunnel-watchdog.service
|
||||||
|
mode: "0644"
|
||||||
|
- name: timer unit
|
||||||
|
sudo: true
|
||||||
|
upload:
|
||||||
|
src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer
|
||||||
|
dest: /etc/systemd/system/cira-tunnel-watchdog.timer
|
||||||
|
mode: "0644"
|
||||||
|
- name: reload and enable the timer
|
||||||
|
sudo: true
|
||||||
|
shell: systemctl daemon-reload && systemctl enable --now cira-tunnel-watchdog.timer
|
||||||
|
changed_when: "false"
|
||||||
|
|
||||||
|
verify:
|
||||||
|
- name: timer is active and scheduled
|
||||||
|
sudo: true
|
||||||
|
shell: systemctl is-active --quiet cira-tunnel-watchdog.timer && systemctl list-timers cira-tunnel-watchdog.timer --no-pager | grep -q cira-tunnel-watchdog
|
||||||
|
changed_when: "false"
|
||||||
|
- name: the live MPS capture parses (dry run exits 0)
|
||||||
|
sudo: true
|
||||||
|
shell: /usr/local/sbin/cira-tunnel-watchdog --dry-run
|
||||||
|
changed_when: "false"
|
||||||
@@ -68,6 +68,8 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc.
|
|||||||
tunnel dead within seconds, and it was gone by 2252. The first one lingered 13+ min, probably because repeated
|
tunnel dead within seconds, and it was gone by 2252. The first one lingered 13+ min, probably because repeated
|
||||||
connect attempts kept writing to it (inferred). **Third time 2026-10-03 1218, on nh3-pve-2** (NH3, Prime power-cycling it
|
connect attempts kept writing to it (inferred). **Third time 2026-10-03 1218, on nh3-pve-2** (NH3, Prime power-cycling it
|
||||||
on site), so it is not ESH- or NAT-specific: 3 of 3 power-event episodes, at two sites.
|
on site), so it is not ESH- or NAT-specific: 3 of 3 power-event episodes, at two sites.
|
||||||
|
**Automated since 2026-10-03 1232 (Prime's go): `cira-tunnel-watchdog` timer on this host drops any :4433 tunnel
|
||||||
|
silent >180 s, every minute** (`services/cira-tunnel-watchdog/`, `journalctl -t cira-tunnel-watchdog`).
|
||||||
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
|
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
|
||||||
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
|
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
|
||||||
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's
|
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's
|
||||||
|
|||||||
@@ -0,0 +1,36 @@
|
|||||||
|
# cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm)
|
||||||
|
|
||||||
|
**Fixes:** MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at **"Setup…"** forever.
|
||||||
|
|
||||||
|
**Why it happens.** A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new
|
||||||
|
CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay
|
||||||
|
(`webserver.js` `handleRelayWebSocket`) can pick the dead one. A TLS failure there is only debug-logged and never
|
||||||
|
closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because
|
||||||
|
MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites:
|
||||||
|
esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up:
|
||||||
|
`servers/pfi-tacticalrmm/README.md`.
|
||||||
|
|
||||||
|
**What it does.** A systemd timer on pfi-tacticalrmm runs `/usr/local/sbin/cira-tunnel-watchdog` every minute.
|
||||||
|
Any ESTABLISHED socket on :4433 that has received nothing for **180 s** (`CIRA_MAX_SILENT_MS`) is destroyed with
|
||||||
|
`ss -K` and logged (`journalctl -t cira-tunnel-watchdog`). MeshCentral then drops it, and the live tunnel serves.
|
||||||
|
- **Margin:** healthy tunnels showed `lastrcv` of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s.
|
||||||
|
- **A wrong drop costs seconds:** the AMT policy reconnects every 10 s, and a reconnect of a known device is the
|
||||||
|
safe code path (only a device's FIRST connect hits the MeshCentral crash race).
|
||||||
|
|
||||||
|
**Install / update:** `scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml`
|
||||||
|
(installed 2026-10-03 1232, Prime's go).
|
||||||
|
|
||||||
|
**Tests (controls, run before install):**
|
||||||
|
```bash
|
||||||
|
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss
|
||||||
|
# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control
|
||||||
|
services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss
|
||||||
|
# → nothing null control
|
||||||
|
```
|
||||||
|
On the live host, `CIRA_MAX_SILENT_MS=1 … --dry-run` must list every tunnel (proves the parser sees them), and the
|
||||||
|
kill filter `( sport = :4433 ) and dst <peer>` must select exactly one socket. Both were checked at install.
|
||||||
|
|
||||||
|
**What it does not cover:** an AMT with no link at all (cable out, or Linux has downed a shared igc port). There
|
||||||
|
is then no tunnel to fix, and the device shows offline in MeshCentral.
|
||||||
|
|
||||||
|
Uninstall: `systemctl disable --now cira-tunnel-watchdog.timer`, then remove the two units and the script.
|
||||||
+56
@@ -0,0 +1,56 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# cira-tunnel-watchdog: drop half-open Intel AMT phone-home (CIRA) tunnels on MeshCentral's MPS.
|
||||||
|
#
|
||||||
|
# Why (servers/pfi-tacticalrmm/README.md): when a host whose AMT shares the network chip power-cycles, AMT opens
|
||||||
|
# a new CIRA tunnel without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay (KVM "HW Connect",
|
||||||
|
# SOL, IDE-R, MeshCommander) can pick the dead one and hang at "Setup..." forever. Its own 90 s idle timeout
|
||||||
|
# does not fire, because its writes to the dead socket keep resetting it. Seen 3 of 3 power events, at two sites
|
||||||
|
# (2026-10-02/03).
|
||||||
|
#
|
||||||
|
# What: every run, any ESTABLISHED socket on the MPS port that has received nothing for MAX_SILENT_MS is
|
||||||
|
# destroyed with `ss -K`. MeshCentral then drops it, and the live tunnel serves. Healthy tunnels hear from their
|
||||||
|
# AMT every ~6-30 s (measured), so the 180 s default leaves a wide margin. A tunnel dropped by mistake comes
|
||||||
|
# back on its own: the AMT policy reconnects every 10 s.
|
||||||
|
#
|
||||||
|
# cira-tunnel-watchdog act (root: ss -K needs CAP_NET_ADMIN)
|
||||||
|
# cira-tunnel-watchdog --dry-run report what would be dropped
|
||||||
|
# cira-tunnel-watchdog --dry-run --from-file F decide from a saved `ss -tnio` capture (tests)
|
||||||
|
set -euo pipefail
|
||||||
|
PORT="${CIRA_PORT:-4433}"
|
||||||
|
MAX_SILENT_MS="${CIRA_MAX_SILENT_MS:-180000}"
|
||||||
|
DRY=0
|
||||||
|
SRC=""
|
||||||
|
while [ $# -gt 0 ]; do
|
||||||
|
case "$1" in
|
||||||
|
--dry-run) DRY=1 ;;
|
||||||
|
--from-file) SRC="${2:?--from-file needs a path}"; shift ;;
|
||||||
|
*) echo "usage: $0 [--dry-run] [--from-file <ss capture>]" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
shift
|
||||||
|
done
|
||||||
|
if [ -n "$SRC" ]; then
|
||||||
|
[ "$DRY" = 1 ] || { echo "--from-file is a test mode and needs --dry-run" >&2; exit 2; }
|
||||||
|
out="$(cat "$SRC")"
|
||||||
|
else
|
||||||
|
out="$(ss -tnio state established "( sport = :$PORT )")"
|
||||||
|
fi
|
||||||
|
# `ss -tnio state established` prints a socket line (Recv-Q Send-Q Local Peer) followed by an indented info line.
|
||||||
|
stale="$(printf '%s\n' "$out" | awk -v max="$MAX_SILENT_MS" '
|
||||||
|
/^[0-9]/ { peer = $4; next }
|
||||||
|
/lastrcv:/ {
|
||||||
|
lr = -1
|
||||||
|
for (i = 1; i <= NF; i++) if ($i ~ /^lastrcv:/) { split($i, a, ":"); lr = a[2] + 0 }
|
||||||
|
if (peer != "" && lr > max) print peer, lr
|
||||||
|
peer = ""
|
||||||
|
}')"
|
||||||
|
[ -z "$stale" ] && exit 0
|
||||||
|
while read -r peer lr; do
|
||||||
|
[[ "$peer" =~ ^\[[0-9a-fA-F:.]+\]:[0-9]+$ ]] || { echo "skip unparsable peer [$peer]" >&2; continue; }
|
||||||
|
if [ "$DRY" = 1 ]; then
|
||||||
|
echo "would drop $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
|
||||||
|
else
|
||||||
|
ss -K -tn "( sport = :$PORT ) and dst $peer" >/dev/null
|
||||||
|
logger -t cira-tunnel-watchdog "dropped stale CIRA tunnel $peer (silent ${lr} ms > ${MAX_SILENT_MS})"
|
||||||
|
echo "dropped $peer (silent ${lr} ms)"
|
||||||
|
fi
|
||||||
|
done <<<"$stale"
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
# Drops half-open Intel AMT CIRA tunnels on MeshCentral's MPS (:4433). See README.md beside this file.
|
||||||
|
[Unit]
|
||||||
|
Description=Drop stale Intel AMT CIRA tunnels on MeshCentral's MPS
|
||||||
|
After=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=/usr/local/sbin/cira-tunnel-watchdog
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Every minute: drop stale Intel AMT CIRA tunnels
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnBootSec=2min
|
||||||
|
OnUnitActiveSec=1min
|
||||||
|
AccuracySec=10s
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
|
||||||
|
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
|
||||||
|
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 cwnd:10 lastsnd:22588 lastrcv:19332 lastack:22584
|
||||||
|
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:16994
|
||||||
|
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 cwnd:10 lastsnd:16360 lastrcv:30136 lastack:16356
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
|
||||||
|
0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664
|
||||||
|
cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 pmtu:1500 rcvmss:838 advmss:1460 cwnd:10 bytes_sent:71724 bytes_received:372619 lastsnd:22588 lastrcv:22588 lastack:22584 app_limited busy:3684ms
|
||||||
|
0 0 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16992
|
||||||
|
cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 pmtu:1500 cwnd:10 bytes_sent:49346 bytes_received:77389 lastsnd:16360 lastrcv:16360 lastack:16356 app_limited busy:632ms
|
||||||
|
0 2989 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16995 timer:(on,1min48sec,14)
|
||||||
|
cubic wscale:0,7 rto:120000 backoff:14 rtt:4.056/0.136 ato:40 mss:1460 pmtu:1500 cwnd:1 bytes_sent:53105 bytes_retrans:7350 lastsnd:11780 lastrcv:782484 lastack:782480 unacked:1 retrans:1/14 lost:1 notsent:2464
|
||||||
Reference in New Issue
Block a user