diff --git a/persistent-memory.md b/persistent-memory.md index 0641543..7c91c22 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._ ## Recent decisions +- `[2026-10-02]` **esh-pve-2 (MS-03 at ESH, future `esh-dev` host) onboarded: infra-ops (Prime), 1 TB → LVM-thin `vmstore`, AMT phoning home as `esh-pve-2-amt`.** The stale firmware boot entry for an old install on the 1 TB was deleted; that drive was wiped and pvesh-created as `vmstore` (913 GiB; `playbooks/esh-pve-2-disk-prep.yaml`); the reboot onto kernel 7.0.14-20 is verified. AMT 21 (ACM, same MEBx pw, vaulted `esh-pve-2/amt-admin`) shares nic1 (I226-LM), its MAC and its DHCP IP with the host: KVM on, OptIn 0, listener on, then `amt-cira-setup.py --apply`. ⚠ Its FIRST CIRA connect crashed MeshCentral once (~7 s, agents back; MeshCentral 1.2.0 race in mpsserver.js, read at source); credentials were then picked up by dropping only its tunnel (`ss -K`), with no restart. Address 10.0.10.70 is TEMPORARY (Prime: can wait) and must become a reservation on MAC `38:05:25:3b:9c:12`. → `servers/esh-pve-2/README.md`, `servers/pfi-tacticalrmm/README.md` - `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423. - `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md` - `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today. diff --git a/servers/esh-pve-2/README.md b/servers/esh-pve-2/README.md new file mode 100644 index 0000000..43644d4 --- /dev/null +++ b/servers/esh-pve-2/README.md @@ -0,0 +1,60 @@ +# esh-pve-2 — Minisforum MS-03, Proxmox VE 9 (ESH) + +The second of the two MS-03s (the other is `nh3-pve-2`). Planned home of **`esh-dev`**, which will take over most of +nh3-dev's sessions (Prime, 2026-10-02; the move itself is not yet planned). Standalone node, **not** in the ESH +cluster (`pve` + `esh-nas-pve`, still on PVE 8.4). + +- **Address: `10.0.10.70`, DHCP, TEMPORARY.** Prime: the permanent address "can wait". When it is chosen, reserve it + on MAC `38:05:25:3b:9c:12` (see AMT below: the host and AMT share that MAC and that address). +- **Access:** `ssh infra-ops@10.0.10.70` (fleet key, NOPASSWD sudo, `Defaults:infra-ops log_output`), set up by Prime + 2026-10-02. Web UI `https://10.0.10.70:8006`. +- PVE 9.2.21, kernel 7.0.14-20-pve (2026-10-02 2210 boot). + +## Hardware + +| | | +|---|---| +| Board | MS-03 (`PTWSA`), Intel Panther Lake | +| NICs | `nic1` **I226-LM 2.5G (vPro/AMT)**, the only cabled port and vmbr0's port; `nic0` RTL8127 10G; `nic2`/`nic3` X710 SFP+ | +| Boot disk | Toshiba KXG60ZNV256G 256 GB, serial `199A3537K01N`: ESP, LVM `pve` (root 70 G, swap 4 G, `local-lvm` 14 G) | +| VM disk | Crucial CT1000E100SSD8 1 TB, serial `2611EAD0291A`: whole-disk LVM-thin **`vmstore`** (913 GiB) | + +⚠ **`nvme0`/`nvme1` swap between boots** (seen 2026-10-02: the 1 TB was `nvme0` before the reboot and `nvme1` +after). Address disks by `/dev/disk/by-id/…` (serial), never `nvmeXn1`. + +## Disks and boot (2026-10-02, Prime) + +`playbooks/esh-pve-2-disk-prep.yaml`, re-runnable: +- **Firmware boot entries:** `Boot0000 proxmox` (256 GB, shim) first; `Boot0006 UEFI OS` (same ESP's fallback + loader, kept); built-in EFI shell (inactive). Deleted `Boot0005`, the fallback loader of an older Proxmox install + on the 1 TB drive. GRUB itself never listed that install: `os-prober` is not installed and PVE disables it. +- **1 TB drive:** held that old install (VG renamed `pve-OLD-2E68F512` by the installer, thin pool 0.00% used). + VG and PV removed, signatures wiped, GPT zapped, whole drive discarded, then `pvesh create …/disks/lvmthin` + (the GUI path) made `vmstore`, content `images,rootdir`, node-restricted. Verified with a 1 GiB alloc/free, + again after the reboot. +- **Format choice: LVM-thin**, the PVE default and the same as the boot drive and esh-pve. ZFS was the alternative + (checksums, compression, replication) but is heavy on a single DRAM-less consumer drive. Cheap to change only + while `vmstore` is empty. + +## Intel AMT — phones home to MeshCentral (2026-10-02 2218) + +- **AMT and Proxmox share the I226-LM port, its MAC and its IP** (AMT DHCP, `SharedMAC`/`SharedDynamicIP` true): + AMT answers on `10.0.10.70` for its own ports only. AMT 21.0.6, **Admin Control Mode**, MEBx password = the one + on nh3-pve / nh3-pve-2, vaulted `esh-pve-2/amt-admin`. +- Set over WS-Man before phone-home: KVM enabled, redirection listener on (IDER/SOL/KVM), **OptInRequired 0** + (no consent code). Then `scripts/amt-cira-setup.py --apply` (MPS `rmm-mesh.phasefinal.com:4433`, user + `CtDDEpGX0VLlJ1X9`, periodic 10 s policy, random environment-detection domain). +- **MeshCentral device `esh-pve-2-amt`** (group `PFI-AMT`; credentials + `tls` 1 set on the device). Tunnel arrives + from ESH's public IP `128.177.138.182`. Read back 2221: CIRA connected, power on, AMT 21.0.6. +- ⚠ **Manage it through MeshCentral.** In phone-home ("outside") mode AMT refuses LAN management, so + `scripts/amt-wsman.py` against `10.0.10.70` no longer works. Last resort: MEBx (Ctrl+P at boot), on site. +- ⚠ **Keep `nic1` up.** On these boards igc powers the PHY off when Linux downs the port, and AMT loses its link + (auto-memory `reference_ms01_amt_port_must_stay_up`). Today it is vmbr0's bridge port, so it stays up. If vmbr0 + ever moves to another NIC, give `nic1` its own `auto nic1` / `iface nic1 inet manual`. +- ⚠ **AMT must stay on DHCP** for phone-home (Intel: CIRA does not work on a static IP). When the permanent address is + chosen, make it a DHCP reservation on `38:05:25:3b:9c:12`. If the host goes static, use that same address so the + two keep sharing one IP. +- Power schemes offered: "Mobile: ON in S0" and "ON in S0, ME Wake in S3, S4-5 (AC only)". Not checked which is + active. Its twin nh3-pve-2 kept its tunnel while powered off (2026-10-02 2221), so this one probably does too. +- **KVM needs an active iGPU output.** A 4K dummy plug blacked the AMT console on nh3-pve-2 once Linux took the + display; use a 1080p plug. diff --git a/servers/esh-pve-2/ssh-target b/servers/esh-pve-2/ssh-target new file mode 100644 index 0000000..02f89a6 --- /dev/null +++ b/servers/esh-pve-2/ssh-target @@ -0,0 +1 @@ +10.0.10.70 diff --git a/servers/pfi-tacticalrmm/README.md b/servers/pfi-tacticalrmm/README.md index c8fc157..1c91e95 100644 --- a/servers/pfi-tacticalrmm/README.md +++ b/servers/pfi-tacticalrmm/README.md @@ -41,6 +41,19 @@ Monitors and manages endpoints, pushes patches, runs scripts, etc. MeshCentral restart re-ran its AMT manager, which logged in at once (16.1.25, power on). ⚠ MeshCentral 1.2.0 `amtmanager.js` picks TLS-vs-not over CIRA with `boundPorts.indexOf('16992')` used as a boolean (−1 is truthy), so a TLS-only AMT can be tried without TLS. Set `tls` explicitly as above. +- **Third CIRA device: `esh-pve-2-amt`** (MS-03 at ESH, AMT 21.0.6, tunnel from ESH's public IP 128.177.138.182), + 2026-10-02 2218. Same recipe. Instead of restarting MeshCentral after setting credentials, only that device's + tunnel was dropped (`sudo ss -K -tn "dst [::ffff:]:664"`). AMT reconnected within seconds, and the AMT + manager logged in with the new credentials. The other tunnels and the agents were untouched. +- ⚠ **A device's FIRST CIRA connection can crash MeshCentral 1.2.0** (it did once, 2026-10-02 22:18:07; the launcher + restarted it at 22:18:14 and all 14 agents came back). Read at source: in `mpsserver.js`, password-auth path for + a device not yet in the DB, `socket.tag.meshid` is set only inside the reverse-DNS callback, while + `addCiraConnection()` runs at once and its 300 ms timer reads `meshid`. `SetConnectivityState` → + `NotifyUserOfDeviceStateChange` then calls `meshid.split` on undefined, which is uncaught, so the whole server + restarts. The likely trigger is a slow PTR lookup on the source IP: ESH's resolves at zayo.com, while NH3's + has none and its two first connects did not crash. That is inferred, not measured. Reconnects of + an existing device take the other branch and are safe. **Onboard new AMTs at a quiet hour**; expect one ~7 s + MeshCentral restart. - Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS "Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's