diff --git a/persistent-memory.md b/persistent-memory.md index 7cb27a4..4e6a78c 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._ ## Recent decisions +- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235. - `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day. - `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md` - `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md` diff --git a/playbooks/cira-tunnel-watchdog.yaml b/playbooks/cira-tunnel-watchdog.yaml new file mode 100644 index 0000000..2f352bd --- /dev/null +++ b/playbooks/cira-tunnel-watchdog.yaml @@ -0,0 +1,36 @@ +# Install the CIRA tunnel watchdog on pfi-tacticalrmm (MeshCentral's host). Idempotent. +# scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml +# Source + rationale: services/cira-tunnel-watchdog/. +steps: + - name: script + sudo: true + upload: + src: services/cira-tunnel-watchdog/cira-tunnel-watchdog + dest: /usr/local/sbin/cira-tunnel-watchdog + mode: "0755" + - name: service unit + sudo: true + upload: + src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.service + dest: /etc/systemd/system/cira-tunnel-watchdog.service + mode: "0644" + - name: timer unit + sudo: true + upload: + src: services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer + dest: /etc/systemd/system/cira-tunnel-watchdog.timer + mode: "0644" + - name: reload and enable the timer + sudo: true + shell: systemctl daemon-reload && systemctl enable --now cira-tunnel-watchdog.timer + changed_when: "false" + +verify: + - name: timer is active and scheduled + sudo: true + shell: systemctl is-active --quiet cira-tunnel-watchdog.timer && systemctl list-timers cira-tunnel-watchdog.timer --no-pager | grep -q cira-tunnel-watchdog + changed_when: "false" + - name: the live MPS capture parses (dry run exits 0) + sudo: true + shell: /usr/local/sbin/cira-tunnel-watchdog --dry-run + changed_when: "false" diff --git a/servers/pfi-tacticalrmm/README.md b/servers/pfi-tacticalrmm/README.md index 4e812bc..720b864 100644 --- a/servers/pfi-tacticalrmm/README.md +++ b/servers/pfi-tacticalrmm/README.md @@ -68,6 +68,8 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc. tunnel dead within seconds, and it was gone by 2252. The first one lingered 13+ min, probably because repeated connect attempts kept writing to it (inferred). **Third time 2026-10-03 1218, on nh3-pve-2** (NH3, Prime power-cycling it on site), so it is not ESH- or NAT-specific: 3 of 3 power-event episodes, at two sites. + **Automated since 2026-10-03 1232 (Prime's go): `cira-tunnel-watchdog` timer on this host drops any :4433 tunnel + silent >180 s, every minute** (`services/cira-tunnel-watchdog/`, `journalctl -t cira-tunnel-watchdog`). - Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS "Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's diff --git a/services/cira-tunnel-watchdog/README.md b/services/cira-tunnel-watchdog/README.md new file mode 100644 index 0000000..8be118c --- /dev/null +++ b/services/cira-tunnel-watchdog/README.md @@ -0,0 +1,36 @@ +# cira-tunnel-watchdog — drop stale Intel AMT phone-home tunnels (pfi-tacticalrmm) + +**Fixes:** MeshCentral "HW Connect" (and SOL, IDE-R, the embedded MeshCommander) stuck at **"Setup…"** forever. + +**Why it happens.** A host whose AMT shares the network chip power-cycles, or its link drops. AMT then opens a new +CIRA tunnel to MeshCentral's MPS (:4433) without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay +(`webserver.js` `handleRelayWebSocket`) can pick the dead one. A TLS failure there is only debug-logged and never +closes the browser's websocket, so the browser waits forever. The MPS's 90 s idle timeout does not fire, because +MeshCentral's own writes into the dead socket keep resetting it. Seen 3 of 3 power-event episodes, at two sites: +esh-pve-2 twice on 2026-10-02 (2247, 2252), nh3-pve-2 on 2026-10-03 (1218). Full write-up: +`servers/pfi-tacticalrmm/README.md`. + +**What it does.** A systemd timer on pfi-tacticalrmm runs `/usr/local/sbin/cira-tunnel-watchdog` every minute. +Any ESTABLISHED socket on :4433 that has received nothing for **180 s** (`CIRA_MAX_SILENT_MS`) is destroyed with +`ss -K` and logged (`journalctl -t cira-tunnel-watchdog`). MeshCentral then drops it, and the live tunnel serves. +- **Margin:** healthy tunnels showed `lastrcv` of 6–30 s (measured 2026-10-02/03), and the dead ones 101–782 s. +- **A wrong drop costs seconds:** the AMT policy reconnects every 10 s, and a reconnect of a known device is the + safe code path (only a device's FIRST connect hits the MeshCentral crash race). + +**Install / update:** `scripts/elway infra-ops@10.250.50.57 --playbook playbooks/cira-tunnel-watchdog.yaml` +(installed 2026-10-03 1232, Prime's go). + +**Tests (controls, run before install):** +```bash +services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/dead-tunnel.ss +# → would drop [::ffff:128.177.138.182]:16995 (the real 2026-10-02 dead tunnel) positive control +services/cira-tunnel-watchdog/cira-tunnel-watchdog --dry-run --from-file services/cira-tunnel-watchdog/test/all-live.ss +# → nothing null control +``` +On the live host, `CIRA_MAX_SILENT_MS=1 … --dry-run` must list every tunnel (proves the parser sees them), and the +kill filter `( sport = :4433 ) and dst ` must select exactly one socket. Both were checked at install. + +**What it does not cover:** an AMT with no link at all (cable out, or Linux has downed a shared igc port). There +is then no tunnel to fix, and the device shows offline in MeshCentral. + +Uninstall: `systemctl disable --now cira-tunnel-watchdog.timer`, then remove the two units and the script. diff --git a/services/cira-tunnel-watchdog/cira-tunnel-watchdog b/services/cira-tunnel-watchdog/cira-tunnel-watchdog new file mode 100755 index 0000000..ca7cafc --- /dev/null +++ b/services/cira-tunnel-watchdog/cira-tunnel-watchdog @@ -0,0 +1,56 @@ +#!/bin/bash +# cira-tunnel-watchdog: drop half-open Intel AMT phone-home (CIRA) tunnels on MeshCentral's MPS. +# +# Why (servers/pfi-tacticalrmm/README.md): when a host whose AMT shares the network chip power-cycles, AMT opens +# a new CIRA tunnel without closing the old one. MeshCentral 1.2.0 keeps both. Its web relay (KVM "HW Connect", +# SOL, IDE-R, MeshCommander) can pick the dead one and hang at "Setup..." forever. Its own 90 s idle timeout +# does not fire, because its writes to the dead socket keep resetting it. Seen 3 of 3 power events, at two sites +# (2026-10-02/03). +# +# What: every run, any ESTABLISHED socket on the MPS port that has received nothing for MAX_SILENT_MS is +# destroyed with `ss -K`. MeshCentral then drops it, and the live tunnel serves. Healthy tunnels hear from their +# AMT every ~6-30 s (measured), so the 180 s default leaves a wide margin. A tunnel dropped by mistake comes +# back on its own: the AMT policy reconnects every 10 s. +# +# cira-tunnel-watchdog act (root: ss -K needs CAP_NET_ADMIN) +# cira-tunnel-watchdog --dry-run report what would be dropped +# cira-tunnel-watchdog --dry-run --from-file F decide from a saved `ss -tnio` capture (tests) +set -euo pipefail +PORT="${CIRA_PORT:-4433}" +MAX_SILENT_MS="${CIRA_MAX_SILENT_MS:-180000}" +DRY=0 +SRC="" +while [ $# -gt 0 ]; do + case "$1" in + --dry-run) DRY=1 ;; + --from-file) SRC="${2:?--from-file needs a path}"; shift ;; + *) echo "usage: $0 [--dry-run] [--from-file ]" >&2; exit 2 ;; + esac + shift +done +if [ -n "$SRC" ]; then + [ "$DRY" = 1 ] || { echo "--from-file is a test mode and needs --dry-run" >&2; exit 2; } + out="$(cat "$SRC")" +else + out="$(ss -tnio state established "( sport = :$PORT )")" +fi +# `ss -tnio state established` prints a socket line (Recv-Q Send-Q Local Peer) followed by an indented info line. +stale="$(printf '%s\n' "$out" | awk -v max="$MAX_SILENT_MS" ' + /^[0-9]/ { peer = $4; next } + /lastrcv:/ { + lr = -1 + for (i = 1; i <= NF; i++) if ($i ~ /^lastrcv:/) { split($i, a, ":"); lr = a[2] + 0 } + if (peer != "" && lr > max) print peer, lr + peer = "" + }')" +[ -z "$stale" ] && exit 0 +while read -r peer lr; do + [[ "$peer" =~ ^\[[0-9a-fA-F:.]+\]:[0-9]+$ ]] || { echo "skip unparsable peer [$peer]" >&2; continue; } + if [ "$DRY" = 1 ]; then + echo "would drop $peer (silent ${lr} ms > ${MAX_SILENT_MS})" + else + ss -K -tn "( sport = :$PORT ) and dst $peer" >/dev/null + logger -t cira-tunnel-watchdog "dropped stale CIRA tunnel $peer (silent ${lr} ms > ${MAX_SILENT_MS})" + echo "dropped $peer (silent ${lr} ms)" + fi +done <<<"$stale" diff --git a/services/cira-tunnel-watchdog/cira-tunnel-watchdog.service b/services/cira-tunnel-watchdog/cira-tunnel-watchdog.service new file mode 100644 index 0000000..ac84081 --- /dev/null +++ b/services/cira-tunnel-watchdog/cira-tunnel-watchdog.service @@ -0,0 +1,8 @@ +# Drops half-open Intel AMT CIRA tunnels on MeshCentral's MPS (:4433). See README.md beside this file. +[Unit] +Description=Drop stale Intel AMT CIRA tunnels on MeshCentral's MPS +After=network-online.target + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/cira-tunnel-watchdog diff --git a/services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer b/services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer new file mode 100644 index 0000000..6e32892 --- /dev/null +++ b/services/cira-tunnel-watchdog/cira-tunnel-watchdog.timer @@ -0,0 +1,10 @@ +[Unit] +Description=Every minute: drop stale Intel AMT CIRA tunnels + +[Timer] +OnBootSec=2min +OnUnitActiveSec=1min +AccuracySec=10s + +[Install] +WantedBy=timers.target diff --git a/services/cira-tunnel-watchdog/test/all-live.ss b/services/cira-tunnel-watchdog/test/all-live.ss new file mode 100644 index 0000000..281dd85 --- /dev/null +++ b/services/cira-tunnel-watchdog/test/all-live.ss @@ -0,0 +1,5 @@ +Recv-Q Send-Q Local Address:Port Peer Address:Port Process +0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664 + cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 cwnd:10 lastsnd:22588 lastrcv:19332 lastack:22584 +0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:16994 + cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 cwnd:10 lastsnd:16360 lastrcv:30136 lastack:16356 diff --git a/services/cira-tunnel-watchdog/test/dead-tunnel.ss b/services/cira-tunnel-watchdog/test/dead-tunnel.ss new file mode 100644 index 0000000..a8f8210 --- /dev/null +++ b/services/cira-tunnel-watchdog/test/dead-tunnel.ss @@ -0,0 +1,7 @@ +Recv-Q Send-Q Local Address:Port Peer Address:Port Process +0 0 [::ffff:10.250.50.57]:4433 [::ffff:70.230.226.88]:664 + cubic wscale:0,7 rto:208 rtt:6.552/0.088 ato:40 mss:1460 pmtu:1500 rcvmss:838 advmss:1460 cwnd:10 bytes_sent:71724 bytes_received:372619 lastsnd:22588 lastrcv:22588 lastack:22584 app_limited busy:3684ms +0 0 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16992 + cubic wscale:0,7 rto:208 rtt:4.153/0.189 ato:40 mss:1460 pmtu:1500 cwnd:10 bytes_sent:49346 bytes_received:77389 lastsnd:16360 lastrcv:16360 lastack:16356 app_limited busy:632ms +0 2989 [::ffff:10.250.50.57]:4433 [::ffff:128.177.138.182]:16995 timer:(on,1min48sec,14) + cubic wscale:0,7 rto:120000 backoff:14 rtt:4.056/0.136 ato:40 mss:1460 pmtu:1500 cwnd:1 bytes_sent:53105 bytes_retrans:7350 lastsnd:11780 lastrcv:782484 lastack:782480 unacked:1 retrans:1/14 lost:1 notsent:2464