feat(esh-pve-nas): install the 225-package backlog; reboot deferred
pve-manager 8.4.11 -> 8.4.20, corosync 3.1.9 -> 3.1.10-pve2, and kernel
6.8.12-42 staged on the /boot LV. dpkg clean, nothing outstanding for
apt -f install, all PVE services active, cluster quorate, no unapplied
conffiles. Reboot deliberately deferred at operator request, so the host
still runs 6.8.12-13 until a chosen window.
This validates the GRUB fix from 061c4b7 under the exact condition it
was written for. update-grub regenerated entries for the new kernel and
entry 0 -- what GRUB_DEFAULT=0 selects -- is now
/vmlinuz-6.8.12-42-pve with root=ZFS=nvme/ROOT/pve-1, supplied by the
grub.d drop-in since grub-mkconfig cannot derive the pool name itself.
The old kernel keeps correct entries as a fallback and the ext4 rollback
entry is untouched. Had the fix not landed first, saved_entry would
still be pinned to 6.8.12-13 and the host would boot the old kernel
indefinitely -- 161 security updates installed and never run.
/boot holds both kernel sets at 176M used of 488M, confirming the 512M
LV carved out of swap was sized correctly.
Adds a ZFS snapshot step to the upgrade playbook, taken automatically on
ZFS-root nodes before any package lands. That is the first real use of
the boot-environment upside the migration was meant to unlock: rollback
for this upgrade is now `zfs rollback -r
nvme/ROOT/pve-1@pre-upgrade-20260818T141652Z && reboot` rather than
archaeology in dpkg. Also documents that the corosync bump restarts
corosync mid-upgrade, which on a 2-node cluster is a brief quorum event.
This commit is contained in:
@@ -15,6 +15,12 @@
|
||||
# FIRST. Its boot default used to pin a single kernel, so installing a new one
|
||||
# would either break the default entry or silently keep booting the old kernel.
|
||||
#
|
||||
# ⚠ The corosync bump RESTARTS corosync mid-upgrade, which on a 2-node cluster is
|
||||
# a brief quorum event — both nodes' /etc/pve go read-only for a few seconds and
|
||||
# then recover. Guests are unaffected and PVE does this routinely, but do not run
|
||||
# it concurrently with anything that writes cluster config, and check
|
||||
# `pvecm status` afterwards rather than assuming.
|
||||
#
|
||||
# Conffile policy: --force-confdef + --force-confold, i.e. keep the on-disk
|
||||
# version wherever a package ships a changed conffile. That is the right default
|
||||
# for these hosts (hand-tuned /etc/default/grub, grub.d drop-ins, storage.cfg),
|
||||
@@ -58,6 +64,20 @@ steps:
|
||||
ls -la {{ backup_dir }}/
|
||||
creates: "{{ backup_dir }}/pre-upgrade-backup.done"
|
||||
|
||||
# On a ZFS-root node this is the cheapest insurance available: an instant,
|
||||
# space-free snapshot of the entire userspace before 200+ packages land. If the
|
||||
# upgrade goes wrong, the recovery is a rollback and a reboot rather than an
|
||||
# archaeology session in dpkg. Skipped automatically on non-ZFS roots.
|
||||
- name: Snapshot the root dataset (ZFS-root nodes only)
|
||||
shell: |
|
||||
ds=$(findmnt -no SOURCE /)
|
||||
snap="${ds}@pre-upgrade-$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
zfs snapshot "$snap"
|
||||
echo " created $snap"
|
||||
echo " rollback if needed: zfs rollback -r $snap && reboot"
|
||||
zfs list -t snapshot -o name,used,creation -s creation "$ds" 2>/dev/null | tail -4
|
||||
when: "test \"$(findmnt -no FSTYPE /)\" = zfs"
|
||||
|
||||
- name: Refresh package lists
|
||||
shell: apt-get update -qq
|
||||
changed_when: "true"
|
||||
|
||||
Reference in New Issue
Block a user