ops(esh-pve-2): register host, AMT phoning home to MeshCentral; note MeshCentral first-CIRA crash race

This commit is contained in:
vh
2026-10-02 22:25:37 -07:00
parent e707d87713
commit 55004090d1
4 changed files with 75 additions and 0 deletions
+1
View File
@@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions ## Recent decisions
- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md`
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423. - `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md` - `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today. - `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
+60
View File
@@ -0,0 +1,60 @@
# esh-pve-2 — Minisforum MS-03, Proxmox VE 9 (ESH)
The second of the two MS-03s (the other is `nh3-pve-2`). Planned home of **`esh-dev`**, which will take over most of
nh3-dev's sessions (Prime, 2026-10-02; the move itself is not yet planned). Standalone node, **not** in the ESH
cluster (`pve` + `esh-nas-pve`, still on PVE 8.4).
- **Address: `10.0.10.70`, DHCP, TEMPORARY.** Prime: the permanent address "can wait". When it is chosen, reserve it
on MAC `38:05:25:3b:9c:12` (see AMT below: the host and AMT share that MAC and that address).
- **Access:** `ssh infra-ops@10.0.10.70` (fleet key, NOPASSWD sudo, `Defaults:infra-ops log_output`), set up by Prime
2026-10-02. Web UI `https://10.0.10.70:8006`.
- PVE 9.2.21, kernel 7.0.14-20-pve (2026-10-02 2210 boot).
## Hardware
| | |
|---|---|
| Board | MS-03 (`PTWSA`), Intel Panther Lake |
| NICs | `nic1` **I226-LM 2.5G (vPro/AMT)**, the only cabled port and vmbr0's port; `nic0` RTL8127 10G; `nic2`/`nic3` X710 SFP+ |
| Boot disk | Toshiba KXG60ZNV256G 256 GB, serial `199A3537K01N`: ESP, LVM `pve` (root 70 G, swap 4 G, `local-lvm` 14 G) |
| VM disk | Crucial CT1000E100SSD8 1 TB, serial `2611EAD0291A`: whole-disk LVM-thin **`vmstore`** (913 GiB) |
⚠ **`nvme0`/`nvme1` swap between boots** (seen 2026-10-02: the 1 TB was `nvme0` before the reboot and `nvme1`
after). Address disks by `/dev/disk/by-id/…` (serial), never `nvmeXn1`.
## Disks and boot (2026-10-02, Prime)
`playbooks/esh-pve-2-disk-prep.yaml`, re-runnable:
- **Firmware boot entries:** `Boot0000 proxmox` (256 GB, shim) first; `Boot0006 UEFI OS` (same ESP's fallback
loader, kept); built-in EFI shell (inactive). Deleted `Boot0005`, the fallback loader of an older Proxmox install
on the 1 TB drive. GRUB itself never listed that install: `os-prober` is not installed and PVE disables it.
- **1 TB drive:** held that old install (VG renamed `pve-OLD-2E68F512` by the installer, thin pool 0.00% used).
VG and PV removed, signatures wiped, GPT zapped, whole drive discarded, then `pvesh create …/disks/lvmthin`
(the GUI path) made `vmstore`, content `images,rootdir`, node-restricted. Verified with a 1 GiB alloc/free,
again after the reboot.
- **Format choice: LVM-thin**, the PVE default and the same as the boot drive and esh-pve. ZFS was the alternative
(checksums, compression, replication) but is heavy on a single DRAM-less consumer drive. Cheap to change only
while `vmstore` is empty.
## Intel AMT — phones home to MeshCentral (2026-10-02 2218)
- **AMT and Proxmox share the I226-LM port, its MAC and its IP** (AMT DHCP, `SharedMAC`/`SharedDynamicIP` true):
AMT answers on `10.0.10.70` for its own ports only. AMT 21.0.6, **Admin Control Mode**, MEBx password = the one
on nh3-pve / nh3-pve-2, vaulted `esh-pve-2/amt-admin`.
- Set over WS-Man before phone-home: KVM enabled, redirection listener on (IDER/SOL/KVM), **OptInRequired 0**
(no consent code). Then `scripts/amt-cira-setup.py --apply` (MPS `rmm-mesh.phasefinal.com:4433`, user
`CtDDEpGX0VLlJ1X9`, periodic 10 s policy, random environment-detection domain).
- **MeshCentral device `esh-pve-2-amt`** (group `PFI-AMT`; credentials + `tls` 1 set on the device). Tunnel arrives
from ESH's public IP `128.177.138.182`. Read back 2221: CIRA connected, power on, AMT 21.0.6.
- ⚠ **Manage it through MeshCentral.** In phone-home ("outside") mode AMT refuses LAN management, so
`scripts/amt-wsman.py` against `10.0.10.70` no longer works. Last resort: MEBx (Ctrl+P at boot), on site.
- ⚠ **Keep `nic1` up.** On these boards igc powers the PHY off when Linux downs the port, and AMT loses its link
(auto-memory `reference_ms01_amt_port_must_stay_up`). Today it is vmbr0's bridge port, so it stays up. If vmbr0
ever moves to another NIC, give `nic1` its own `auto nic1` / `iface nic1 inet manual`.
- ⚠ **AMT must stay on DHCP** for phone-home (Intel: CIRA does not work on a static IP). When the permanent address is
chosen, make it a DHCP reservation on `38:05:25:3b:9c:12`. If the host goes static, use that same address so the
two keep sharing one IP.
- Power schemes offered: "Mobile: ON in S0" and "ON in S0, ME Wake in S3, S4-5 (AC only)". Not checked which is
active. Its twin nh3-pve-2 kept its tunnel while powered off (2026-10-02 2221), so this one probably does too.
- **KVM needs an active iGPU output.** A 4K dummy plug blacked the AMT console on nh3-pve-2 once Linux took the
display; use a 1080p plug.
+1
View File
@@ -0,0 +1 @@
10.0.10.70
+13
View File
@@ -41,6 +41,19 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc.
MeshCentral restart re-ran its AMT manager, which logged in at once (16.1.25, power on). ⚠ MeshCentral MeshCentral restart re-ran its AMT manager, which logged in at once (16.1.25, power on). ⚠ MeshCentral
1.2.0 `amtmanager.js` picks TLS-vs-not over CIRA with `boundPorts.indexOf('16992')` used as a boolean 1.2.0 `amtmanager.js` picks TLS-vs-not over CIRA with `boundPorts.indexOf('16992')` used as a boolean
(−1 is truthy), so a TLS-only AMT can be tried without TLS. Set `tls` explicitly as above. (−1 is truthy), so a TLS-only AMT can be tried without TLS. Set `tls` explicitly as above.
- **Third CIRA device: `esh-pve-2-amt`** (MS-03 at ESH, AMT 21.0.6, tunnel from ESH's public IP 128.177.138.182),
2026-10-02 2218. Same recipe. Instead of restarting MeshCentral after setting credentials, only that device's
tunnel was dropped (`sudo ss -K -tn "dst [::ffff:<public-ip>]:664"`). AMT reconnected within seconds, and the AMT
manager logged in with the new credentials. The other tunnels and the agents were untouched.
- ⚠ **A device's FIRST CIRA connection can crash MeshCentral 1.2.0** (it did once, 2026-10-02 22:18:07; the launcher
restarted it at 22:18:14 and all 14 agents came back). Read at source: in `mpsserver.js`, password-auth path for
a device not yet in the DB, `socket.tag.meshid` is set only inside the reverse-DNS callback, while
`addCiraConnection()` runs at once and its 300 ms timer reads `meshid`. `SetConnectivityState` →
`NotifyUserOfDeviceStateChange` then calls `meshid.split` on undefined, which is uncaught, so the whole server
restarts. The likely trigger is a slow PTR lookup on the source IP: ESH's resolves at zayo.com, while NH3's
has none and its two first connects did not crash. That is inferred, not measured. Reconnects of
an existing device take the other branch and are safe. **Onboard new AMTs at a quiet hour**; expect one ~7 s
MeshCentral restart.
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS - Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no "Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's