ops(nh3-pve-2): register host, network/AMT/disk layout and the re-IP lesson

This commit is contained in:
vh
2026-10-03 13:17:46 -07:00
parent 9b1cfa7344
commit 22224b5335
3 changed files with 32 additions and 0 deletions
+1
View File
@@ -287,6 +287,7 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions ## Recent decisions
- `[2026-10-03]` **nh3-pve-2 finished (Prime on site): host 10.100.250.62 on nh3-mgmt via X710 `nic3`, AMT on `nic0` with `auto` kept up, 1 TB wiped into `vmstore` (913 GiB), stale boot entry gone.** My remote re-IP at 1259 failed: a live ifreload read the new vmbr0 as "not a bridge" and downed it. My rollback timer then restored the old /etc/hosts. Prime fixed the network at the console at 1312; I fixed `auto nic0` and /etc/hosts at 1316 (no reload). Lesson: auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`. → `servers/nh3-pve-2/README.md`
- `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235. - `[2026-10-03]` **CIRA tunnel watchdog LIVE on pfi-tacticalrmm (Prime go, 1232)** fixes MeshCentral "HW Connect stuck at Setup": after a power event AMT leaves its old tunnel half-open, and MeshCentral's relay picks it (3 of 3 episodes: esh-pve-2 twice, nh3-pve-2 once). A timer runs every minute and `ss -K`s any :4433 socket silent >180 s (healthy tunnels 6–30 s). Controls passed. → `services/cira-tunnel-watchdog/`. **Same day: nh3-pve-2's host moved to its X710 10G port** (`38:05:25:3b:a0:a8`, USW Pro 24 port 25 now native nh3-mgmt + tagged auto, reservation and DNS `nh3-pve-2.nh3.internal` = 10.100.250.62). ⚠ At 1233 its AMT (`:a6`, UDM port 5) had NO link, probably because the I226 port has no `auto` line now (MS-01 lesson, unconfirmed); flagged to Prime at 1235.
- `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day. - `[2026-10-03]` **Worldtree memory-gate duty REFRAMED: regression check, not a gate** (infra-ops msg 01M413W4). The 3-PASS streak completed 10-02 and U11b step 5 ran 2026-10-02 14:54:03Z on Prime's go; 20261003T143101Z was the first post-deletion PASS — recall held without legacy data. Keep the daily batch until worldtree-dev says retire. A FAIL now = recall regression on the LIVE plane (no streak to reset): send infra-ops AND worldtree-dev directly. Cron prompt and skill updated same-day.
- `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md` - `[2026-10-03]` **albok-service 0.1.2 LIVE on nh3-docker (albok-dev ask 6955, 0251), fixing a leak that had broken 0.1.1.** 0.1.1 leaked a Chroma client per health probe: after 9 h it held 1018 of 1024 fds and `/health`/`/search` answered 500. 0.1.2 (`sha256:6f91a33f…`, built from albok-dev's LOCAL `544ecfd` before the push; origin's tags verified = `544ecfd` at 1000) holds 59 fds, 0 deleted, over 30 health calls. Added restricted wing `personal/agent-feedback` (group `albok-feedback` 1512, default-deny). Nemi credentials: service token re-minted with agent-feedback grants (`albok/token-nemi`; old token still live until albok-dev revokes it) and LiteLLM key `albok-nemi` (gen-small/summarizer/qwen3-embedding; `albok/litellm-key-nemi`). **Nemi timer INSTALLED 0559** (user units `albok-nemi.{service,timer}` on nh3-dev, hourly; OnFailure → infra-ops inbox; `services/albok-nemi/`). Gate: albok-dev's steady-state walk, median 594 s over n=3; my test walk through the unit took 589 s with exit 0. ⚠ Reports grow ~55 MB/day (retention asked of albok-dev); manual walks go through `systemctl --user start albok-nemi.service` (no overlap). → `stacks/albok-service/README.md`
+30
View File
@@ -0,0 +1,30 @@
# nh3-pve-2: Minisforum MS-03, Proxmox VE 9 (NH3)
The NH3 MS-03, twin of esh-pve-2. **Purpose is TBD on purpose** (Prime, 2026-10-02): a high-powered PVE host,
possibly a dev environment for security-software work. Standalone node, not clustered with nh3-pve.
- **Host:** `10.100.250.62` on nh3-mgmt (`nh3-pve-2.nh3.internal`, `nh3-pve-2.phasefinal.com` in /etc/hosts). Static
in `/etc/network/interfaces`, plus a UDM reservation on `38:05:25:3b:a0:a8`. `ssh infra-ops@10.100.250.62`
(fleet key, NOPASSWD sudo with `log_output`).
- PVE 9.2.21, installed 2026-10-03 on site to the right disk (the 256 GB Toshiba).
| NIC | MAC | Chip | Use |
|---|---|---|---|
| `nic0` | `:a6` | I226-LM (igc) | **Intel AMT only.** UDM port 5, native nh3-mgmt, AMT DHCP `.63`. `auto nic0`, no address, `arp_ignore=8`, IPv6 off (`/etc/sysctl.d/90-amt-port.conf`) |
| `nic1` | `:a9` | RTL8127 10G | unused |
| `nic2` | `:a7` | X710 SFP+ | unused |
| `nic3` | `:a8` | X710 SFP+ | **vmbr0**. USW Pro 24 port 25 (native nh3-mgmt, tagged VLANs allowed) |
⚠ **Keep `auto nic0`.** If Linux leaves the AMT port down, igc powers off the PHY and AMT goes dark. That
happened 2026-10-03 when the bridge moved to the SFP+.
⚠ `vmbr0` is a plain (not VLAN-aware) bridge. Making it VLAN-aware is a separate change: apply it with a reboot,
not a live `ifreload`. On 2026-10-03 a live ifreload that also re-IP'd read the new vmbr0 as "not a bridge" and
downed it (auto-memory `feedback_no_remote_reip_thrash_when_operator_has_console`).
**Disks:** boot = Toshiba 256 GB (`69CA12VQK65N`). **VM storage = LVM-thin `vmstore`, 913 GiB**, on the Crucial 1 TB
(`2610EAD05AB7`). It held the old wrong-drive install, which was wiped 2026-10-03 by `playbooks/nh3-pve-2-disk-prep.yaml`,
along with its stale firmware boot entry. Firmware boot order: `0000 proxmox` (256 GB), then its fallback `0006`.
Address disks by `/dev/disk/by-id/`, because nvme numbering can swap between boots (seen on esh-pve-2).
**AMT:** MeshCentral device `nh3-pve-2-amt` (phone-home, AMT 21.0.6, ACM; password vaulted `nh3-pve-2/amt-admin`).
If HW Connect sticks at "Setup…", `cira-tunnel-watchdog` on pfi-tacticalrmm clears it within ~4 min.
+1
View File
@@ -0,0 +1 @@
10.100.250.62