Files
esh-pfi-infrastructure/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md
T

119 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# esh-pve-cluster: Proxmox VE 8 → 9 upgrade plan (PLAN ONLY — not executed)
Requested by Prime via Miranda, 2026-10-02 (thread `01M3YX5QGQ5V8BQXYM46E5T3PB`). **Do not execute
without a further green light from Prime.** Facts below were read live on 2026-10-02 ~1930 PT,
read-only (including `pve8to9 --full` on both nodes). Re-run `pve8to9 --full` and check the
official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything that disagrees wins.
## The cluster today
| | `pve` (fleet: esh-pve) | `esh-nas-pve` (fleet: esh-pve-nas) |
|---|---|---|
| Mgmt IP | 10.0.250.35 | 10.0.50.55 |
| HW | Minisforum MS-01, i9-13900H, 62 GiB | Xeon W-1250, 125 GiB |
| PVE / kernel | 8.4.20 / 6.8.12-42 (‑43 pending) | 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending) |
| Root | ext4 on LVM `pve/root` (96 G, 69 G free), VG free 16 G | ZFS `nvme/ROOT/pve-1` (pool `nvme`) |
| Boot | UEFI, **GRUB** (proxmox-boot-tool not in use) | UEFI, **GRUB reading ZFS** (proxmox-boot-tool not in use) |
| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2, **DEGRADED**) |
| Guests | 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter | 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin |
Repos: `pve-no-subscription` + Debian bookworm main/contrib, no Ceph, no HA resources.
Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`.
## ⚠ Findings that come before the upgrade
1. **`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk
`wwn-0x5000c500c91df554` shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20,
3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated,
0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. **Fix first:**
`zpool clear tank` + a full scrub; if errors return, replace the disk. Do not upgrade this node
with the pool degraded.
2. **Remote access to ESH depends on CT 108 (esh-scale) on `pve`.** Measured path from NH3:
`10.100.50.46 → 100.64.0.2 (esh-scale) → 10.0.50.1 → ESH`. When `pve` reboots, NH3/ANA lose ESH
until 108 is back. If `pve` does not come back, there is **no remote recovery path** (esh-pve's
AMT is not set up). Either someone is at ESH for `pve`'s reboot, or set up esh-pve's AMT
(+ phone-home, as done for nh3-pve) first.
3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot
via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays).
## Pre-flight (no guest downtime)
1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the
window, **including 108 and 110**, which the nightly job does not back up. Check each new snapshot is
listed. Plus a host config backup per node: `/etc` and `/etc/pve` tar to PBS or the NAS.
2. **Bring both nodes to the latest 8.4** (`apt update && apt full-upgrade`), then
`pve8to9 --full` must show no FAIL.
3. **Clear the pve8to9 items:**
- both: `apt remove systemd-boot`;
- `pve`: `echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u`;
- both, optional but recommended: enable `non-free-firmware`, `apt install intel-microcode`.
4. **Pin NIC names by MAC** (systemd `.link`) on both nodes before the upgrade. Debian 13's systemd may
change predictable names, and a renamed uplink leaves `vmbr0` with no link: the whole node is off the
network. pve8to9 did not flag it; this is a precaution. Template: `playbooks/nh3-pve-pin-nic-names.yaml`.
Uplinks today: `pve` `enp3s0f0np0` (X710), `esh-nas-pve` `enp5s0f0`.
5. **NVIDIA on `pve`:** 580.178.04 is DKMS-built for 6.8. After the dist-upgrade and **before**
rebooting, `dkms status` must show the module built for the new kernel; if not, reinstall the
driver for trixie before rebooting. esh-ml1 (CT 110) is the fleet's **only** embed/rerank backend
behind LiteLLM, and albok-service now depends on it. **Put nh3-ml1 (its twin) behind the gateway
first**, so embeddings survive `pve`'s window.
6. **Root snapshots for rollback:**
- `esh-nas-pve`: `zfs snapshot -r nvme/ROOT@pre-pve9`.
- `pve`: `lvcreate -s -L 15G -n root-pre-pve9 pve/root` (VG has 16 G free).
## Order and window
**Recommended: `esh-nas-pve` first, then `pve`, in one evening window of about 2–2.5 h.**
Doing the NAS node first keeps the remote path (CT 108 on `pve`) up while the data-heavy node is
worked on, so anything that goes wrong there is still diagnosable from outside.
1. On `pve`, cleanly **shut down VM 101 (esh-vm-db, disk on NFS from 10.0.50.50) and VM 100
(esh-vm-docker, hard NFS mounts)**. A NAS reboot under them means I/O hangs or unkillable D-state
(2026 incident).
2. `esh-nas-pve`: switch repos, dist-upgrade, reboot. Verify that all pools import, `tank` state, NFS
exports, and that guests 103–107 are up.
3. Start VMs 100 and 101 on `pve`; verify the DB and Docker services.
4. `pve`: switch repos, dist-upgrade, check `dkms status`, reboot. Verify CT 108 (mesh), CT 110 (TEI
answering through LiteLLM), CT 111 (Matter), VMs 100/101.
5. `pve8to9 --full` on both (it has post-upgrade checks), `pvecm status` (quorate), guest list.
**Repo switch, per node** (verify against the wiki on the day):
```
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list # Debian main/updates/security
# replace the bookworm pve-no-subscription line with deb822 /etc/apt/sources.list.d/proxmox.sources:
# Types: deb
# URIs: http://download.proxmox.com/debian/pve
# Suites: trixie
# Components: pve-no-subscription
# Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
apt update && apt dist-upgrade # keep local config where asked unless the wiki says otherwise
```
## Expected downtime and guest behaviour
- **The apt phase** (15–30 min per node): guests keep running.
- **The reboot phase:** that node's guests stop and restart on boot (all are `onboot=1`):
- roughly 5–10 min for `pve`;
- 10–20 min for `esh-nas-pve` (the 175 T `tank` import, NFS);
- VMs 100/101 are down from step 1 until step 3, roughly 30–45 min.
- **No migration:** no HA and no live migration (local storage on both nodes, `cpu=host`, different CPUs).
Guests are stopped, not moved.
- **Quorum:** while either node reboots, the other loses quorum. Its running guests keep running, but
you cannot start, stop or reconfigure guests and `/etc/pve` is read-only until the node is back.
Do not use `pvecm expected 1` unless something is stuck.
- **During `pve`'s reboot:**
- ESH loses its mesh router (no remote path; the Matter/HA bridge drops);
- fleet embeddings go down unless nh3-ml1 is behind the gateway;
- the ESH AdGuard resolver (on esh-vm-docker) is down; the DNS ring falls back to the other sites.
## Rollback: honest version
There is **no supported in-place downgrade** from PVE 9 / Debian 13 back to 8 / 12. What exists:
- **`esh-nas-pve`:** `zfs rollback nvme/ROOT/pve-1@pre-pve9` from a rescue shell (PVE ISO, or console).
⚠ Never run `zpool upgrade` on `nvme` after the upgrade. GRUB reads that pool directly, and new
feature flags can make it unbootable *and* break this rollback.
- **`pve`:** merge the LVM snapshot (`lvconvert --merge pve/root-pre-pve9`) from rescue, then reboot.
The snapshot only holds while its 15 G covers the changes; drop it once 9 is confirmed good.
- **Last resort:** reinstall 8.4 and restore `/etc` and guests from the pre-flight PBS backups (hours).
- **Guest data is not touched by the host upgrade**, so the risk is host availability, not data loss.
Both rollbacks need console access, which is why finding 2 matters.