docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding

This commit is contained in:
vh
2026-10-02 19:20:28 -07:00
parent 5439d5fdc1
commit 62254023f3
2 changed files with 119 additions and 0 deletions
@@ -0,0 +1,118 @@
# esh-pve-cluster: Proxmox VE 8 → 9 upgrade plan (PLAN ONLY — not executed)
Requested by Prime via Miranda, 2026-10-02 (thread `01M3YX5QGQ5V8BQXYM46E5T3PB`). **Do not execute
without a further green light from Prime.** Facts below were read live on 2026-10-02 ~1930 PT,
read-only (including `pve8to9 --full` on both nodes). Re-run `pve8to9 --full` and check the
official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything that disagrees wins.
## The cluster today
| | `pve` (fleet: esh-pve) | `esh-nas-pve` (fleet: esh-pve-nas) |
|---|---|---|
| Mgmt IP | 10.0.250.35 | 10.0.50.55 |
| HW | Minisforum MS-01, i9-13900H, 62 GiB | Xeon W-1250, 125 GiB |
| PVE / kernel | 8.4.20 / 6.8.12-42 (‑43 pending) | 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending) |
| Root | ext4 on LVM `pve/root` (96 G, 69 G free), VG free 16 G | ZFS `nvme/ROOT/pve-1` (pool `nvme`) |
| Boot | UEFI, **GRUB** (proxmox-boot-tool not in use) | UEFI, **GRUB reading ZFS** (proxmox-boot-tool not in use) |
| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2, **DEGRADED**) |
| Guests | 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter | 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin |
Repos: `pve-no-subscription` + Debian bookworm main/contrib, no Ceph, no HA resources.
Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`.
## ⚠ Findings that come before the upgrade
1. **`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk
`wwn-0x5000c500c91df554` shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20,
3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated,
0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. **Fix first:**
`zpool clear tank` + a full scrub; if errors return, replace the disk. Do not upgrade this node
with the pool degraded.
2. **Remote access to ESH depends on CT 108 (esh-scale) on `pve`.** Measured path from NH3:
`10.100.50.46 → 100.64.0.2 (esh-scale) → 10.0.50.1 → ESH`. When `pve` reboots, NH3/ANA lose ESH
until 108 is back. If `pve` does not come back, there is **no remote recovery path** (esh-pve's
AMT is not set up). Either someone is at ESH for `pve`'s reboot, or set up esh-pve's AMT
(+ phone-home, as done for nh3-pve) first.
3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot
via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays).
## Pre-flight (no guest downtime)
1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the
window, **including 108 and 110**, which the nightly job does not back up. Check each new snapshot is
listed. Plus a host config backup per node: `/etc` and `/etc/pve` tar to PBS or the NAS.
2. **Bring both nodes to the latest 8.4** (`apt update && apt full-upgrade`), then
`pve8to9 --full` must show no FAIL.
3. **Clear the pve8to9 items:**
- both: `apt remove systemd-boot`;
- `pve`: `echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u`;
- both, optional but recommended: enable `non-free-firmware`, `apt install intel-microcode`.
4. **Pin NIC names by MAC** (systemd `.link`) on both nodes before the upgrade. Debian 13's systemd may
change predictable names, and a renamed uplink leaves `vmbr0` with no link: the whole node is off the
network. pve8to9 did not flag it; this is a precaution. Template: `playbooks/nh3-pve-pin-nic-names.yaml`.
Uplinks today: `pve` `enp3s0f0np0` (X710), `esh-nas-pve` `enp5s0f0`.
5. **NVIDIA on `pve`:** 580.178.04 is DKMS-built for 6.8. After the dist-upgrade and **before**
rebooting, `dkms status` must show the module built for the new kernel; if not, reinstall the
driver for trixie before rebooting. esh-ml1 (CT 110) is the fleet's **only** embed/rerank backend
behind LiteLLM, and albok-service now depends on it. **Put nh3-ml1 (its twin) behind the gateway
first**, so embeddings survive `pve`'s window.
6. **Root snapshots for rollback:**
- `esh-nas-pve`: `zfs snapshot -r nvme/ROOT@pre-pve9`.
- `pve`: `lvcreate -s -L 15G -n root-pre-pve9 pve/root` (VG has 16 G free).
## Order and window
**Recommended: `esh-nas-pve` first, then `pve`, in one evening window of about 2–2.5 h.**
Doing the NAS node first keeps the remote path (CT 108 on `pve`) up while the data-heavy node is
worked on, so anything that goes wrong there is still diagnosable from outside.
1. On `pve`, cleanly **shut down VM 101 (esh-vm-db, disk on NFS from 10.0.50.50) and VM 100
(esh-vm-docker, hard NFS mounts)**. A NAS reboot under them means I/O hangs or unkillable D-state
(2026 incident).
2. `esh-nas-pve`: switch repos, dist-upgrade, reboot. Verify that all pools import, `tank` state, NFS
exports, and that guests 103–107 are up.
3. Start VMs 100 and 101 on `pve`; verify the DB and Docker services.
4. `pve`: switch repos, dist-upgrade, check `dkms status`, reboot. Verify CT 108 (mesh), CT 110 (TEI
answering through LiteLLM), CT 111 (Matter), VMs 100/101.
5. `pve8to9 --full` on both (it has post-upgrade checks), `pvecm status` (quorate), guest list.
**Repo switch, per node** (verify against the wiki on the day):
```
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list # Debian main/updates/security
# replace the bookworm pve-no-subscription line with deb822 /etc/apt/sources.list.d/proxmox.sources:
# Types: deb
# URIs: http://download.proxmox.com/debian/pve
# Suites: trixie
# Components: pve-no-subscription
# Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
apt update && apt dist-upgrade # keep local config where asked unless the wiki says otherwise
```
## Expected downtime and guest behaviour
- **The apt phase** (15–30 min per node): guests keep running.
- **The reboot phase:** that node's guests stop and restart on boot (all are `onboot=1`):
- roughly 5–10 min for `pve`;
- 10–20 min for `esh-nas-pve` (the 175 T `tank` import, NFS);
- VMs 100/101 are down from step 1 until step 3, roughly 30–45 min.
- **No migration:** no HA and no live migration (local storage on both nodes, `cpu=host`, different CPUs).
Guests are stopped, not moved.
- **Quorum:** while either node reboots, the other loses quorum. Its running guests keep running, but
you cannot start, stop or reconfigure guests and `/etc/pve` is read-only until the node is back.
Do not use `pvecm expected 1` unless something is stuck.
- **During `pve`'s reboot:**
- ESH loses its mesh router (no remote path; the Matter/HA bridge drops);
- fleet embeddings go down unless nh3-ml1 is behind the gateway;
- the ESH AdGuard resolver (on esh-vm-docker) is down; the DNS ring falls back to the other sites.
## Rollback: honest version
There is **no supported in-place downgrade** from PVE 9 / Debian 13 back to 8 / 12. What exists:
- **`esh-nas-pve`:** `zfs rollback nvme/ROOT/pve-1@pre-pve9` from a rescue shell (PVE ISO, or console).
⚠ Never run `zpool upgrade` on `nvme` after the upgrade. GRUB reads that pool directly, and new
feature flags can make it unbootable *and* break this rollback.
- **`pve`:** merge the LVM snapshot (`lvconvert --merge pve/root-pre-pve9`) from rescue, then reboot.
The snapshot only holds while its 15 G covers the changes; drop it once 9 is confirmed good.
- **Last resort:** reinstall 8.4 and restore `/etc` and guests from the pre-flight PBS backups (hours).
- **Guest data is not touched by the host upgrade**, so the risk is host availability, not data loss.
Both rollbacks need console access, which is why finding 2 matters.
+1
View File
@@ -290,6 +290,7 @@ _As of 2026-10-01 ~0446 PT._
- `[2026-10-02]` **Demo outage ~13 min, ROLLED BACK (worldtree-dev URGENT 6684).** Their b193 release commit c2d87263 swept 60 staged deletions into the tree, so the demo api crash-looped. I recreated `worldtree-api` only (`compose up -d --no-deps`) on the `.env`-pinned fa8bc51cc064 (compose.yaml identical at both shas): healthy in 35 s at 15:16Z. ⚠ Their deploy health gate does NOT restore the old container on failure (it only withholds `:latest`); flagged to them. Fix-forward 7ab6ae40 (tag v1.0.0b193 moved onto it) deployed via CI 15:22Z and verified healthy, same seven retired-key WARNINGs; the deploy-gate gap is Worldtree #423.
- `[2026-10-02]` **nh3-pve's AMT is in TacticalRMM's MeshCentral (group `PFI-AMT`, device `nh3-pve-amt`).** The adds had failed silently because TRMM installs MeshCentral `WANonly` (meshuser.js:2682 drops AMT adds). Prime ruled hybrid; switched 0812, 14/14 agents back, AMT connected at once (16.1.25, TLS fine). **UPDATE 0859: it now PHONES HOME (CIRA)** — MeshCentral mpsPass, ana-gw VIP/policy 76 on 4433, AMT settings via `scripts/amt-cira-setup.py`, AMT moved static → DHCP (Intel: CIRA needs DHCP; reservation keeps .61). Tunnel is independent of nh3-pve. ⚠ While phoning home the AMT REFUSES LAN management (:16993 dark), so manage it via MeshCentral; revert body `servers/nh3-pve/amt-ethernet-static-revert.xml`. 2FA still not forced. Prime's MeshCentral login token is vaulted `pfi-tacticalrmm/meshcentral-login-token`. → `servers/pfi-tacticalrmm/README.md`
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
- `[2026-10-02]` ⚠ **ESH `tank` (esh-nas-pve, 175 T raidz2×2) is DEGRADED since ~2026-08-20 — found while planning PVE 9.** raidz2-0 disk `wwn-0x5000c500c91df554`: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. Next: `zpool clear` + scrub, replace if errors return (Prime's call). **PVE 8→9 plan for esh-pve-cluster written, NOT executed** (Prime via Miranda): `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. Also found: remote access to ESH rides CT 108 on `pve` (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.
- `[2026-10-02]` **albok-service 0.1.0 LIVE on nh3-docker (albok-dev ask, operator-approved): `http://albok.nh3.internal:8392`** (8390 there is the post office). Store/private on local ext4 under /srv/albok, uid 1500, read groups 1510/1511 (container gets a mounted /etc/group + group_add so its getgrnam/chgrp work). Scoped LiteLLM key `albok-service` (qwen3-embedding only) + bootstrap admin token in the vault under `albok/`. Upgraded to **0.1.1** (e349d50) at 1805: /health now `ok` (0.1.0's always-degraded canary bug fixed); albok-dev smoke-tested ingest/search/readback (token vaulted `albok/token-albok-dev`). → `stacks/albok-service/README.md`
- `[2026-10-02]` **nh3-pve-2 status at end of day: AMT done, Proxmox on the WRONG NVMe; Prime reinstalls ON SITE 2026-10-03.** IDE-R from MeshCentral (ISO `proxmox-ve_9.2-1.iso` in lkraven's My Files, kept) installed it but then stalled across the internet; the 4K dummy plug blacks the AMT console once Linux takes the display (8-bit/grayscale encoding or a 1080p plug fixes it); the Realtek 10G (`…:a0:a9`, USW port 1) only links during firmware, so it likely has no driver. **On-site checklist:** USB-stick install → pick the right disk under Target Harddisk → Options; host NIC = i226 (shared with AMT) or an X710 SFP+, not the Realtek; then I wipe the old disk (`wipefs` + `zpool labelclear` if ZFS: two `rpool`s collide), reserve .62, onboard `infra-ops`, keep the AMT port admin-UP in Linux (MS-01 lesson), swap in a 1080p plug.
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand. One MS-03 is being deployed as `nh3-pve-2` at NH3 (Prime, ~0910).** Plan: its AMT on DHCP (needed for phone-home) with a UDM reservation, proposed 10.100.250.63 / host 10.100.250.62 (static in PVE + reservation); configure KVM/opt-in over LAN FIRST, then `scripts/amt-cira-setup.py` (LAN goes dark after). MS-03 = i226-LM vPro 2.5G + RTL8127 10G RJ45 + 2× X710 SFP+. **Cabled 2026-10-02 1435 on UDM port 5: AMT 21.0.6 answers (TLS-only, :16993/:664), MAC 38:05:25:3b:a0:a6, sharing the factory Windows' DHCP address (WIN-94HJ50P1LUE).** Port 5 is now native nh3-mgmt (copy of port 6); reservation nh3-pve-2-amt = 10.100.250.63; DNS added. It moved to .63 at 14:47 (replug/reboot). Then (MEBx password = nh3-pve's, vaulted `nh3-pve-2/amt-admin`): KVM on, listener on, OptIn 0 (ACM), `amt-cira-setup.py --apply` → **phoning home at once**; MeshCentral device `nh3-pve-2-amt` (creds + tls=1, needed a MeshCentral restart to log in) shows AMT 21.0.6, power on. LAN :16993 dark by design. Remote install: MeshCentral's embedded MeshCommander (device → Intel AMT tab) has IDE-R; AMT_RedirectionService 32771 = IDER+SOL enabled. **Roles (Prime ~0915): both MS-03 run Proxmox. The other one hosts `esh-dev` at ESH, which will INHERIT MOST OF nh3-dev's SESSIONS (a migration, not yet planned). nh3-pve-2's purpose is TBD ON PURPOSE (high-powered PVE host; possibly a dev environment for security software).** Do not assign it a role. When the esh-dev move is planned, inventory what is anchored to nh3-dev first: the althing herald, svos/hermes-gateway (Miranda's channel), the Booth, the fleet TLS caddy and the `*.nh3.phasefinal.com` rewrite to 10.100.10.50, dev-backup, ttyd/zellij seats, and the `nh3-dev/` vault namespace. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.