061c4b7712
The cutover left saved_entry=pve-zfs-root, a hand-authored 40_custom entry hardcoding /vmlinuz-6.8.12-13-pve. The pending upgrade installs proxmox-kernel-6.8.12-42, which made that a trap with two exits: if -13 were autoremoved the default entry would point at a missing kernel and the host would need console recovery it has no IPMI for; if -13 survived the host would silently keep booting the old kernel, so 161 security updates including a kernel would install and never run. That entry was written as a one-time cutover target. It was never fit to be the standing default across kernel upgrades, and this is remediation of that, caught before the upgrade rather than after. Fix is to stop hand-authoring the ZFS entry: GRUB_DEFAULT=0 boots the first auto-generated entry, which grub-mkconfig regenerates for the newest kernel on every install, and which /etc/default/grub.d/zfs-root.cfg already corrects to the pool-qualified root=ZFS=nvme/ROOT/pve-1. grubenv is cleared so nothing overrides it. The rollback entry stays pinned, which is correct rather than an oversight: it boots the untouched ext4 root on the DOM, whose /boot is never regenerated because update-initramfs writes only to the /boot LV. That kernel genuinely never changes. Also adds a reusable safe-reboot playbook for this host, carrying the constraints that are easy to forget: quiesce the hard-NFS clients first, the other cluster node goes read-only while this one is down (quorum 2, no qdevice), and a failed boot has no auto-fallback and no remote console.