Files
esh-pfi-infrastructure/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md
T

9.2 KiB
Raw Blame History

esh-pve-cluster: Proxmox VE 8 → 9 upgrade plan (PLAN ONLY — not executed)

Requested by Prime via Miranda, 2026-10-02 (thread 01M3YX5QGQ5V8BQXYM46E5T3PB). Do not execute without a further green light from Prime. Facts below were read live on 2026-10-02 ~1930 PT, read-only (including pve8to9 --full on both nodes). Re-run pve8to9 --full and check the official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything that disagrees wins.

The cluster today

pve (fleet: esh-pve) esh-nas-pve (fleet: esh-pve-nas)
Mgmt IP 10.0.250.35 10.0.50.55
HW Minisforum MS-01, i9-13900H, 62 GiB Xeon W-1250, 125 GiB
PVE / kernel 8.4.20 / 6.8.12-42 (‑43 pending) 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending)
Root ext4 on LVM pve/root (96 G, 69 G free), VG free 16 G ZFS nvme/ROOT/pve-1 (pool nvme)
Boot UEFI, GRUB (proxmox-boot-tool not in use) UEFI, GRUB reading ZFS (proxmox-boot-tool not in use)
Extras NVIDIA 580.178.04 DKMS (RTX 2000E Ada → CT 110 esh-ml1) ZFS pools nvme, ssd, tank (175 T raidz2×2; healthy since 2026-10-03, was DEGRADED)
Guests 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin

Repos: pve-no-subscription + Debian bookworm main/contrib, no Ceph, no HA resources. Quorum: 2 nodes × 1 vote, no QDevice, no two_node.

⚠ Findings that come before the upgrade

  1. ✅ RESOLVED 2026-10-03 0629, so no longer a blocker. Two full scrubs ran. The first repaired 198 G, the residue of the Aug 20 resilver, which had never been followed by a full scrub. The second repaired 0 B with 0 errors. The disk was kept, and zpool status -x says healthy. The original finding follows. tank on esh-nas-pve is DEGRADED, and has been since about 2026-08-20. In raidz2-0, disk wwn-0x5000c500c91df554 shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20, 3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated, 0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. Fix first: zpool clear tank + a full scrub; if errors return, replace the disk. Do not upgrade this node with the pool degraded.
  2. Remote access to ESH depends on CT 108 (esh-scale) on pve. Measured path from NH3: 10.100.50.46 → 100.64.0.2 (esh-scale) → 10.0.50.1 → ESH. When pve reboots, NH3/ANA lose ESH until 108 is back. If pve does not come back, there is no remote recovery path (esh-pve's AMT is not set up). Either someone is at ESH for pve's reboot, or set up esh-pve's AMT (+ phone-home, as done for nh3-pve) first.
  3. pve8to9 FAIL on both nodes: the systemd-boot meta-package is installed. Both nodes boot via GRUB, so remove it (apt remove systemd-boot; systemd-boot-efi stays).

esh-nas-pve facts re-read 2026-10-03 (Prime asked whether infra-ops can upgrade it unattended)

  • QNAP hardware (QNAP Systems, Inc., Comet Lake HECI present, but no AMT/IPMI we can use), so there is no remote console.
  • PVE 8.4.20, but still running kernel 6.8.12-13, up since 2026-08-18. Its first reboot in ~7 weeks is itself a risk: do a reboot test on PVE 8 first, so a reboot problem is not mistaken for an upgrade problem.
  • No NIC pins (/etc/systemd/network/ is empty); the uplink is enp5s0f0. Pin before the upgrade.
  • No DKMS on this node (the NVIDIA headers lesson applies to pve, not here).
  • Infra-ops can do the prep remotely. The upgrade reboot should happen only with someone able to reach ESH, or with Prime's explicit acceptance that a failed boot means an outage until someone is on site.

Pre-flight (no guest downtime)

  1. Backups: one-off vzdump of all ten guests to pbs-ana in snapshot mode, right before the window, including 108 and 110, which the nightly job does not back up. Check each new snapshot is listed. Plus a host config backup per node: /etc and /etc/pve tar to PBS or the NAS.
  2. Bring both nodes to the latest 8.4 (apt update && apt full-upgrade), then pve8to9 --full must show no FAIL.
  3. Clear the pve8to9 items:
    • both: apt remove systemd-boot;
    • pve: echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u;
    • both, optional but recommended: enable non-free-firmware, apt install intel-microcode.
  4. Pin NIC names by MAC (systemd .link) on both nodes before the upgrade. Debian 13's systemd may change predictable names, and a renamed uplink leaves vmbr0 with no link: the whole node is off the network. pve8to9 did not flag it; this is a precaution. Template: playbooks/nh3-pve-pin-nic-names.yaml. Uplinks today: pve enp3s0f0np0 (X710), esh-nas-pve enp5s0f0.
  5. NVIDIA on pve: first make sure proxmox-default-headers is installed. On nh3-pve (2026-10-03) the upgrade put in kernel 7.0 with NO 7.0 headers, so DKMS never built the module and the GPU container did not start. The fix was apt install proxmox-default-headers proxmox-headers-<new> + dkms autoinstall. 580.178.04 does compile on 7.0.14. (playbooks/nh3-pve-pve9-postfix.yaml.) 580.178.04 is DKMS-built for 6.8. After the dist-upgrade and before rebooting, dkms status must show the module built for the new kernel; if not, reinstall the driver for trixie before rebooting. esh-ml1 (CT 110) is the fleet's only embed/rerank backend behind LiteLLM, and albok-service now depends on it. Put nh3-ml1 (its twin) behind the gateway first, so embeddings survive pve's window.
  6. Root snapshots for rollback:
    • esh-nas-pve: zfs snapshot -r nvme/ROOT@pre-pve9.
    • pve: lvcreate -s -L 15G -n root-pre-pve9 pve/root (VG has 16 G free).

Order and window

Recommended: esh-nas-pve first, then pve, in one evening window of about 2–2.5 h. Doing the NAS node first keeps the remote path (CT 108 on pve) up while the data-heavy node is worked on, so anything that goes wrong there is still diagnosable from outside.

  1. On pve, cleanly shut down VM 101 (esh-vm-db, disk on NFS from 10.0.50.50) and VM 100 (esh-vm-docker, hard NFS mounts). A NAS reboot under them means I/O hangs or unkillable D-state (2026 incident).
  2. esh-nas-pve: switch repos, dist-upgrade, reboot. Verify that all pools import, tank state, NFS exports, and that guests 103–107 are up.
  3. Start VMs 100 and 101 on pve; verify the DB and Docker services.
  4. pve: switch repos, dist-upgrade, check dkms status, reboot. Verify CT 108 (mesh), CT 110 (TEI answering through LiteLLM), CT 111 (Matter), VMs 100/101.
  5. pve8to9 --full on both (it has post-upgrade checks), pvecm status (quorate), guest list.

Repo switch, per node (verify against the wiki on the day):

sed -i 's/bookworm/trixie/g' /etc/apt/sources.list            # Debian main/updates/security
# replace the bookworm pve-no-subscription line with deb822 /etc/apt/sources.list.d/proxmox.sources:
#   Types: deb
#   URIs: http://download.proxmox.com/debian/pve
#   Suites: trixie
#   Components: pve-no-subscription
#   Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
apt update && apt dist-upgrade        # keep local config where asked unless the wiki says otherwise

Expected downtime and guest behaviour

  • The apt phase (15–30 min per node): guests keep running.
  • The reboot phase: that node's guests stop and restart on boot (all are onboot=1):
    • roughly 5–10 min for pve;
    • 10–20 min for esh-nas-pve (the 175 T tank import, NFS);
    • VMs 100/101 are down from step 1 until step 3, roughly 30–45 min.
  • No migration: no HA and no live migration (local storage on both nodes, cpu=host, different CPUs). Guests are stopped, not moved.
  • Quorum: while either node reboots, the other loses quorum. Its running guests keep running, but you cannot start, stop or reconfigure guests and /etc/pve is read-only until the node is back. Do not use pvecm expected 1 unless something is stuck.
  • During pve's reboot:
    • ESH loses its mesh router (no remote path; the Matter/HA bridge drops);
    • fleet embeddings go down unless nh3-ml1 is behind the gateway;
    • the ESH AdGuard resolver (on esh-vm-docker) is down; the DNS ring falls back to the other sites.

Rollback: honest version

There is no supported in-place downgrade from PVE 9 / Debian 13 back to 8 / 12. What exists:

  • esh-nas-pve: zfs rollback nvme/ROOT/pve-1@pre-pve9 from a rescue shell (PVE ISO, or console). ⚠ Never run zpool upgrade on nvme after the upgrade. GRUB reads that pool directly, and new feature flags can make it unbootable and break this rollback.
  • pve: merge the LVM snapshot (lvconvert --merge pve/root-pre-pve9) from rescue, then reboot. The snapshot only holds while its 15 G covers the changes; drop it once 9 is confirmed good.
  • Last resort: reinstall 8.4 and restore /etc and guests from the pre-flight PBS backups (hours).
  • Guest data is not touched by the host upgrade, so the risk is host availability, not data loss. Both rollbacks need console access, which is why finding 2 matters.