docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding

This commit is contained in:
vh
2026-10-02 19:20:28 -07:00
parent 5439d5fdc1
commit 62254023f3
2 changed files with 119 additions and 0 deletions
@@ -0,0 +1,118 @@
# esh-pve-cluster: Proxmox VE 8 → 9 upgrade plan (PLAN ONLY — not executed)
Requested by Prime via Miranda, 2026-10-02 (thread `01M3YX5QGQ5V8BQXYM46E5T3PB`). **Do not execute
without a further green light from Prime.** Facts below were read live on 2026-10-02 ~1930 PT,
read-only (including `pve8to9 --full` on both nodes). Re-run `pve8to9 --full` and check the
official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything that disagrees wins.
## The cluster today
| | `pve` (fleet: esh-pve) | `esh-nas-pve` (fleet: esh-pve-nas) |
|---|---|---|
| Mgmt IP | 10.0.250.35 | 10.0.50.55 |
| HW | Minisforum MS-01, i9-13900H, 62 GiB | Xeon W-1250, 125 GiB |
| PVE / kernel | 8.4.20 / 6.8.12-42 (‑43 pending) | 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending) |
| Root | ext4 on LVM `pve/root` (96 G, 69 G free), VG free 16 G | ZFS `nvme/ROOT/pve-1` (pool `nvme`) |
| Boot | UEFI, **GRUB** (proxmox-boot-tool not in use) | UEFI, **GRUB reading ZFS** (proxmox-boot-tool not in use) |
| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2, **DEGRADED**) |
| Guests | 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter | 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin |
Repos: `pve-no-subscription` + Debian bookworm main/contrib, no Ceph, no HA resources.
Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`.
## ⚠ Findings that come before the upgrade
1. **`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk
`wwn-0x5000c500c91df554` shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20,
3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated,
0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. **Fix first:**
`zpool clear tank` + a full scrub; if errors return, replace the disk. Do not upgrade this node
with the pool degraded.
2. **Remote access to ESH depends on CT 108 (esh-scale) on `pve`.** Measured path from NH3:
`10.100.50.46 → 100.64.0.2 (esh-scale) → 10.0.50.1 → ESH`. When `pve` reboots, NH3/ANA lose ESH
until 108 is back. If `pve` does not come back, there is **no remote recovery path** (esh-pve's
AMT is not set up). Either someone is at ESH for `pve`'s reboot, or set up esh-pve's AMT
(+ phone-home, as done for nh3-pve) first.
3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot
via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays).
## Pre-flight (no guest downtime)
1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the
window, **including 108 and 110**, which the nightly job does not back up. Check each new snapshot is
listed. Plus a host config backup per node: `/etc` and `/etc/pve` tar to PBS or the NAS.
2. **Bring both nodes to the latest 8.4** (`apt update && apt full-upgrade`), then
`pve8to9 --full` must show no FAIL.
3. **Clear the pve8to9 items:**
- both: `apt remove systemd-boot`;
- `pve`: `echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u`;
- both, optional but recommended: enable `non-free-firmware`, `apt install intel-microcode`.
4. **Pin NIC names by MAC** (systemd `.link`) on both nodes before the upgrade. Debian 13's systemd may
change predictable names, and a renamed uplink leaves `vmbr0` with no link: the whole node is off the
network. pve8to9 did not flag it; this is a precaution. Template: `playbooks/nh3-pve-pin-nic-names.yaml`.
Uplinks today: `pve` `enp3s0f0np0` (X710), `esh-nas-pve` `enp5s0f0`.
5. **NVIDIA on `pve`:** 580.178.04 is DKMS-built for 6.8. After the dist-upgrade and **before**
rebooting, `dkms status` must show the module built for the new kernel; if not, reinstall the
driver for trixie before rebooting. esh-ml1 (CT 110) is the fleet's **only** embed/rerank backend
behind LiteLLM, and albok-service now depends on it. **Put nh3-ml1 (its twin) behind the gateway
first**, so embeddings survive `pve`'s window.
6. **Root snapshots for rollback:**
- `esh-nas-pve`: `zfs snapshot -r nvme/ROOT@pre-pve9`.
- `pve`: `lvcreate -s -L 15G -n root-pre-pve9 pve/root` (VG has 16 G free).
## Order and window
**Recommended: `esh-nas-pve` first, then `pve`, in one evening window of about 2–2.5 h.**
Doing the NAS node first keeps the remote path (CT 108 on `pve`) up while the data-heavy node is
worked on, so anything that goes wrong there is still diagnosable from outside.
1. On `pve`, cleanly **shut down VM 101 (esh-vm-db, disk on NFS from 10.0.50.50) and VM 100
(esh-vm-docker, hard NFS mounts)**. A NAS reboot under them means I/O hangs or unkillable D-state
(2026 incident).
2. `esh-nas-pve`: switch repos, dist-upgrade, reboot. Verify that all pools import, `tank` state, NFS
exports, and that guests 103–107 are up.
3. Start VMs 100 and 101 on `pve`; verify the DB and Docker services.
4. `pve`: switch repos, dist-upgrade, check `dkms status`, reboot. Verify CT 108 (mesh), CT 110 (TEI
answering through LiteLLM), CT 111 (Matter), VMs 100/101.
5. `pve8to9 --full` on both (it has post-upgrade checks), `pvecm status` (quorate), guest list.
**Repo switch, per node** (verify against the wiki on the day):
```
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list # Debian main/updates/security
# replace the bookworm pve-no-subscription line with deb822 /etc/apt/sources.list.d/proxmox.sources:
# Types: deb
# URIs: http://download.proxmox.com/debian/pve
# Suites: trixie
# Components: pve-no-subscription
# Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
apt update && apt dist-upgrade # keep local config where asked unless the wiki says otherwise
```
## Expected downtime and guest behaviour
- **The apt phase** (15–30 min per node): guests keep running.
- **The reboot phase:** that node's guests stop and restart on boot (all are `onboot=1`):
- roughly 5–10 min for `pve`;
- 10–20 min for `esh-nas-pve` (the 175 T `tank` import, NFS);
- VMs 100/101 are down from step 1 until step 3, roughly 30–45 min.
- **No migration:** no HA and no live migration (local storage on both nodes, `cpu=host`, different CPUs).
Guests are stopped, not moved.
- **Quorum:** while either node reboots, the other loses quorum. Its running guests keep running, but
you cannot start, stop or reconfigure guests and `/etc/pve` is read-only until the node is back.
Do not use `pvecm expected 1` unless something is stuck.
- **During `pve`'s reboot:**
- ESH loses its mesh router (no remote path; the Matter/HA bridge drops);
- fleet embeddings go down unless nh3-ml1 is behind the gateway;
- the ESH AdGuard resolver (on esh-vm-docker) is down; the DNS ring falls back to the other sites.
## Rollback: honest version
There is **no supported in-place downgrade** from PVE 9 / Debian 13 back to 8 / 12. What exists:
- **`esh-nas-pve`:** `zfs rollback nvme/ROOT/pve-1@pre-pve9` from a rescue shell (PVE ISO, or console).
⚠ Never run `zpool upgrade` on `nvme` after the upgrade. GRUB reads that pool directly, and new
feature flags can make it unbootable *and* break this rollback.
- **`pve`:** merge the LVM snapshot (`lvconvert --merge pve/root-pre-pve9`) from rescue, then reboot.
The snapshot only holds while its 15 G covers the changes; drop it once 9 is confirmed good.
- **Last resort:** reinstall 8.4 and restore `/etc` and guests from the pre-flight PBS backups (hours).
- **Guest data is not touched by the host upgrade**, so the risk is host availability, not data loss.
Both rollbacks need console access, which is why finding 2 matters.