From 62254023f34ffaf2974c1089afe715491c9668fb Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Fri, 2 Oct 2026 19:20:28 -0700 Subject: [PATCH] docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding --- .../esh-pve-cluster-pve9-upgrade-plan.md | 118 ++++++++++++++++++ persistent-memory.md | 1 + 2 files changed, 119 insertions(+) create mode 100644 docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md diff --git a/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md b/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md new file mode 100644 index 0000000..94a3f47 --- /dev/null +++ b/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md @@ -0,0 +1,118 @@ +# esh-pve-cluster: Proxmox VE 8 → 9 upgrade plan (PLAN ONLY — not executed) + +Requested by Prime via Miranda, 2026-10-02 (thread `01M3YX5QGQ5V8BQXYM46E5T3PB`). **Do not execute +without a further green light from Prime.** Facts below were read live on 2026-10-02 ~1930 PT, +read-only (including `pve8to9 --full` on both nodes). Re-run `pve8to9 --full` and check the +official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything that disagrees wins. + +## The cluster today + +| | `pve` (fleet: esh-pve) | `esh-nas-pve` (fleet: esh-pve-nas) | +|---|---|---| +| Mgmt IP | 10.0.250.35 | 10.0.50.55 | +| HW | Minisforum MS-01, i9-13900H, 62 GiB | Xeon W-1250, 125 GiB | +| PVE / kernel | 8.4.20 / 6.8.12-42 (‑43 pending) | 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending) | +| Root | ext4 on LVM `pve/root` (96 G, 69 G free), VG free 16 G | ZFS `nvme/ROOT/pve-1` (pool `nvme`) | +| Boot | UEFI, **GRUB** (proxmox-boot-tool not in use) | UEFI, **GRUB reading ZFS** (proxmox-boot-tool not in use) | +| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2, **DEGRADED**) | +| Guests | 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter | 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin | + +Repos: `pve-no-subscription` + Debian bookworm main/contrib, no Ceph, no HA resources. +Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`. + +## ⚠ Findings that come before the upgrade + +1. **`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk + `wwn-0x5000c500c91df554` shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20, + 3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated, + 0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. **Fix first:** + `zpool clear tank` + a full scrub; if errors return, replace the disk. Do not upgrade this node + with the pool degraded. +2. **Remote access to ESH depends on CT 108 (esh-scale) on `pve`.** Measured path from NH3: + `10.100.50.46 → 100.64.0.2 (esh-scale) → 10.0.50.1 → ESH`. When `pve` reboots, NH3/ANA lose ESH + until 108 is back. If `pve` does not come back, there is **no remote recovery path** (esh-pve's + AMT is not set up). Either someone is at ESH for `pve`'s reboot, or set up esh-pve's AMT + (+ phone-home, as done for nh3-pve) first. +3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot + via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays). + +## Pre-flight (no guest downtime) + +1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the + window, **including 108 and 110**, which the nightly job does not back up. Check each new snapshot is + listed. Plus a host config backup per node: `/etc` and `/etc/pve` tar to PBS or the NAS. +2. **Bring both nodes to the latest 8.4** (`apt update && apt full-upgrade`), then + `pve8to9 --full` must show no FAIL. +3. **Clear the pve8to9 items:** + - both: `apt remove systemd-boot`; + - `pve`: `echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u`; + - both, optional but recommended: enable `non-free-firmware`, `apt install intel-microcode`. +4. **Pin NIC names by MAC** (systemd `.link`) on both nodes before the upgrade. Debian 13's systemd may + change predictable names, and a renamed uplink leaves `vmbr0` with no link: the whole node is off the + network. pve8to9 did not flag it; this is a precaution. Template: `playbooks/nh3-pve-pin-nic-names.yaml`. + Uplinks today: `pve` `enp3s0f0np0` (X710), `esh-nas-pve` `enp5s0f0`. +5. **NVIDIA on `pve`:** 580.178.04 is DKMS-built for 6.8. After the dist-upgrade and **before** + rebooting, `dkms status` must show the module built for the new kernel; if not, reinstall the + driver for trixie before rebooting. esh-ml1 (CT 110) is the fleet's **only** embed/rerank backend + behind LiteLLM, and albok-service now depends on it. **Put nh3-ml1 (its twin) behind the gateway + first**, so embeddings survive `pve`'s window. +6. **Root snapshots for rollback:** + - `esh-nas-pve`: `zfs snapshot -r nvme/ROOT@pre-pve9`. + - `pve`: `lvcreate -s -L 15G -n root-pre-pve9 pve/root` (VG has 16 G free). + +## Order and window + +**Recommended: `esh-nas-pve` first, then `pve`, in one evening window of about 2–2.5 h.** +Doing the NAS node first keeps the remote path (CT 108 on `pve`) up while the data-heavy node is +worked on, so anything that goes wrong there is still diagnosable from outside. + +1. On `pve`, cleanly **shut down VM 101 (esh-vm-db, disk on NFS from 10.0.50.50) and VM 100 + (esh-vm-docker, hard NFS mounts)**. A NAS reboot under them means I/O hangs or unkillable D-state + (2026 incident). +2. `esh-nas-pve`: switch repos, dist-upgrade, reboot. Verify that all pools import, `tank` state, NFS + exports, and that guests 103–107 are up. +3. Start VMs 100 and 101 on `pve`; verify the DB and Docker services. +4. `pve`: switch repos, dist-upgrade, check `dkms status`, reboot. Verify CT 108 (mesh), CT 110 (TEI + answering through LiteLLM), CT 111 (Matter), VMs 100/101. +5. `pve8to9 --full` on both (it has post-upgrade checks), `pvecm status` (quorate), guest list. + +**Repo switch, per node** (verify against the wiki on the day): +``` +sed -i 's/bookworm/trixie/g' /etc/apt/sources.list # Debian main/updates/security +# replace the bookworm pve-no-subscription line with deb822 /etc/apt/sources.list.d/proxmox.sources: +# Types: deb +# URIs: http://download.proxmox.com/debian/pve +# Suites: trixie +# Components: pve-no-subscription +# Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg +apt update && apt dist-upgrade # keep local config where asked unless the wiki says otherwise +``` + +## Expected downtime and guest behaviour + +- **The apt phase** (15–30 min per node): guests keep running. +- **The reboot phase:** that node's guests stop and restart on boot (all are `onboot=1`): + - roughly 5–10 min for `pve`; + - 10–20 min for `esh-nas-pve` (the 175 T `tank` import, NFS); + - VMs 100/101 are down from step 1 until step 3, roughly 30–45 min. +- **No migration:** no HA and no live migration (local storage on both nodes, `cpu=host`, different CPUs). + Guests are stopped, not moved. +- **Quorum:** while either node reboots, the other loses quorum. Its running guests keep running, but + you cannot start, stop or reconfigure guests and `/etc/pve` is read-only until the node is back. + Do not use `pvecm expected 1` unless something is stuck. +- **During `pve`'s reboot:** + - ESH loses its mesh router (no remote path; the Matter/HA bridge drops); + - fleet embeddings go down unless nh3-ml1 is behind the gateway; + - the ESH AdGuard resolver (on esh-vm-docker) is down; the DNS ring falls back to the other sites. + +## Rollback: honest version + +There is **no supported in-place downgrade** from PVE 9 / Debian 13 back to 8 / 12. What exists: +- **`esh-nas-pve`:** `zfs rollback nvme/ROOT/pve-1@pre-pve9` from a rescue shell (PVE ISO, or console). + ⚠ Never run `zpool upgrade` on `nvme` after the upgrade. GRUB reads that pool directly, and new + feature flags can make it unbootable *and* break this rollback. +- **`pve`:** merge the LVM snapshot (`lvconvert --merge pve/root-pre-pve9`) from rescue, then reboot. + The snapshot only holds while its 15 G covers the changes; drop it once 9 is confirmed good. +- **Last resort:** reinstall 8.4 and restore `/etc` and guests from the pre-flight PBS backups (hours). +- **Guest data is not touched by the host upgrade**, so the risk is host availability, not data loss. + Both rollbacks need console access, which is why finding 2 matters. diff --git a/persistent-memory.md b/persistent-memory.md index fddce49..d47298b 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -290,6 +290,7 @@ _As of 2026-10-01 ~0446 PT._ - `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423. - `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md` - `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today. +- `[2026-10-02]` ⚠ **ESH `tank` (esh-nas-pve, 175 T raidz2×2) is DEGRADED since ~2026-08-20 — found while planning PVE 9.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. Next: `zpool clear` + scrub, replace if errors return (Prime's call). **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package. - `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md` - `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug. - `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.