feat(esh-pve-nas): cut PVE root over to ZFS on the mirrored NVMe

Root is now nvme/ROOT/pve-1. The USB DOM keeps the ESP and /boot but is
out of the runtime I/O path, so a bus reset can no longer drop root from
under a running hypervisor. All five guests healthy, three pools ONLINE,
system running, ext4 pve-root intact and unmounted as the rollback with
its own kernel and initrd. zfs-import-cache is now the active import
path -- the all-three-pools cachefile fix doing its job.

The window cost an unplanned outage, and the cause was this repo's own
tooling rather than the migration.

The staging chroot ran `mount --rbind /dev` and /sys with no
--make-rslave. On systemd `/` has shared propagation, so the cutover's
`umount -R` propagated back into the live host and removed the real
/sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone systemd-logind
could not create a session: ping fine, TCP fine, SSH authentication
succeeded, resident daemons kept serving -- and every new exec hung,
including /sbin/reboot, so the reboot never ran at all.

It impersonates failing root-disk I/O almost perfectly, and I called it
as the DOM dying. That was wrong. dmesg had the answer throughout: the
DOM attached cleanly with no errors, and the last log timestamp was
12114881s -- 140 days -- meaning this was still the original boot. A
down-detector had also never reported the host down, which I read as a
fast reboot rather than as no reboot.

Fixes and guards:

- --make-rslave after every rbind, plus a guard that refuses to proceed
  while any chroot bind still reports shared propagation.
- Confirm a reboot by observing the host DOWN, not by watching for it to
  come back. Those two states are indistinguishable otherwise.
- Blast radius now measured from the server: `ss` inside CT 103 found
  five NFS clients, not the two documented. The new one that mattered is
  esh-vm-db, hard-mounted and unreachable by ssh. Left mounted on
  purpose and it came through read-write.
- grub-reboot's one-shot does NOT work here: grubenv sits on an LVM LV
  which GRUB reads but cannot write, so next_entry survived the boot
  that consumed it. Steady state is saved_entry=pve-zfs-root with no
  next_entry. There is no auto-fallback on this host and no IPMI.

Recovery needed no console: an idempotent cgroup2/devpts/shm remount
landed in the brief windows where exec succeeded. No data was lost, and
neither the DOM nor any pool was ever at risk.
This commit is contained in:
vh
2026-08-18 06:39:34 -07:00
parent b637947ffd
commit 5f11d1b3cb
8 changed files with 480 additions and 24 deletions
@@ -0,0 +1,55 @@
# esh-pve-nas cutover, step 1 of 5 — quiesce esh-docker-vm's hard NFS mounts.
#
# Run: scripts/elway infra-ops@10.0.50.45 --playbook playbooks/esh-cutover-1-quiesce-docker-vm.yaml
#
# Why this is first and why it is not optional: /mnt/books and /mnt/backup are
# `hard` NFS from CT 103 on esh-pve-nas. A hard mount does not fail when the
# server goes away — it blocks forever in D-state, and the only known remedy is
# rebooting THIS host. /mnt/books was deliberately left hard because calibre's
# SQLite risks corruption under `soft`, so the mount option is not the fix; the
# quiesce is.
#
# Measured 2026-08-18: exactly one container binds these paths
# (calibre-web-automated -> /mnt/books/calibre/{ingest,calibre_library}) and
# /mnt/backup has no container consumers at all. The blast radius is one service,
# not the seventeen containers on this host.
#
# Reversed by playbooks/esh-cutover-5-restore.yaml.
steps:
# No --format here: elway substitutes {{ ... }}, so Go template braces in a
# shell command are a booby trap. --filter + -q avoids them entirely.
- name: Stop the only container holding the NFS mounts
shell: sudo -n docker stop calibre-web-automated
when: "test -n \"$(sudo -n docker ps -q --filter name=^calibre-web-automated$)\""
- name: Confirm nothing else has files open under the mounts
shell: |
busy=$(sudo -n lsof +D /mnt/books +D /mnt/backup 2>/dev/null | tail -n +2 | wc -l)
if [ "$busy" -ne 0 ]; then
echo "STILL BUSY — refusing to unmount:"
sudo -n lsof +D /mnt/books +D /mnt/backup 2>/dev/null | head -20
exit 1
fi
echo "no open files under either mount"
changed_when: "false"
- name: Unmount /mnt/books
shell: sudo -n umount /mnt/books
when: "mountpoint -q /mnt/books"
- name: Unmount /mnt/backup
shell: sudo -n umount /mnt/backup
when: "mountpoint -q /mnt/backup"
verify:
- name: Neither NFS mount remains
shell: "! findmnt -t nfs,nfs4 -o TARGET | grep -qE '/mnt/(books|backup)'"
changed_when: "false"
- name: The other sixteen containers are still up
shell: |
n=$(sudo -n docker ps -q | wc -l)
echo "$n containers still running"
test "$n" -ge 10
changed_when: "false"
@@ -0,0 +1,51 @@
# esh-pve-nas cutover, step 2 of 5 — quiesce esh-pve's hard NFS storages.
#
# Run: scripts/elway root@10.0.250.35 --playbook playbooks/esh-cutover-2-quiesce-esh-pve.yaml
#
# esh-pve mounts two `hard` NFS storages from CT 103 on esh-pve-nas:
# esh-nas -> 10.0.50.50:/mnt/pvestore at /mnt/pve/esh-nas
# tank-vmbu -> 10.0.50.50:/mnt/tank-vmbu at /mnt/pve/tank-vmbu
#
# Disabling the storage first matters: if the storage stays enabled, pvestatd
# keeps stat()ing the path and will re-trigger the mount (and then block on it)
# the moment the server disappears. Disable, THEN unmount.
#
# Measured 2026-08-18: esh-nas holds 2.9 MB of 96 TB and no running guest has a
# disk on either storage — all three (100 esh-vm-docker, 101 esh-vm-db,
# 102 esh-vm-workstation) live on local-lvm. So this quiesce costs backup targets
# for the duration, not guest availability. Guests are deliberately left running.
#
# Reversed by playbooks/esh-cutover-5-restore.yaml.
steps:
- name: Disable the esh-nas storage so pvestatd stops touching it
shell: pvesm set esh-nas --disable 1
when: "pvesm status 2>/dev/null | awk '$1==\"esh-nas\"{print $3}' | grep -q active"
- name: Disable the tank-vmbu storage
shell: pvesm set tank-vmbu --disable 1
when: "grep -q '^nfs: tank-vmbu' /etc/pve/storage.cfg && ! grep -A8 '^nfs: tank-vmbu' /etc/pve/storage.cfg | grep -q 'disable'"
- name: Give pvestatd a moment to let go before unmounting
shell: sleep 5
changed_when: "false"
- name: Unmount /mnt/pve/esh-nas
shell: umount /mnt/pve/esh-nas || umount -l /mnt/pve/esh-nas
when: "mountpoint -q /mnt/pve/esh-nas"
- name: Unmount /mnt/pve/tank-vmbu
shell: umount /mnt/pve/tank-vmbu || umount -l /mnt/pve/tank-vmbu
when: "mountpoint -q /mnt/pve/tank-vmbu"
verify:
- name: Neither esh-nas-backed NFS mount remains
shell: "! findmnt -t nfs,nfs4 -o SOURCE | grep -q '10\\.0\\.50\\.50'"
changed_when: "false"
- name: All three guests are still running
shell: |
n=$(qm list | awk 'NR>1 && $3=="running"' | wc -l)
echo "$n VMs running"
test "$n" -eq 3
changed_when: "false"
+154
View File
@@ -0,0 +1,154 @@
# esh-pve-nas cutover, step 3 of 5 — point the ESP at the new /boot and reboot.
#
# Run: scripts/elway root@esh-pve-nas --playbook playbooks/esh-cutover-3-esh-pve-nas.yaml
#
# PRECONDITION: steps 1 and 2 must have run. Both NFS clients hold `hard` mounts
# from CT 103 which lives on this host; taking it down with them mounted wedges
# esh-docker-vm in unkillable D-state. A guard below refuses to proceed if either
# client is still mounted.
#
# ⚠ ORDERING TRAP, and it is the reason this is a playbook and not four commands:
# `zfs set mountpoint=/` on a dataset that is CURRENTLY MOUNTED makes ZFS unmount
# and REMOUNT it at the new location — i.e. it would try to mount the ZFS root
# over the live ext4 root of a running hypervisor. canmount=noauto does not save
# you; that governs automatic mounting at import, not an explicit property change
# on a mounted dataset. The dataset must be UNMOUNTED first, which means the
# chroot binds have to come down first, which means grub-install and grub-reboot
# have to happen BEFORE any of that. Hence the sequence below is not negotiable.
#
# This playbook ENDS BY REBOOTING THE HOST. elway will lose the connection; that
# is expected, not a failure.
vars:
newroot: /mnt/newroot
root_dataset: nvme/ROOT/pve-1
quiesced: "no" # caller MUST pass --var quiesced=yes after verifying both clients
steps:
# ---------- guards ----------
- name: GUARD — still on the ext4 root (not already cut over)
shell: |
test "$(findmnt -no FSTYPE /)" = "ext4" || { echo "already on ZFS; refusing"; exit 1; }
changed_when: "false"
# This host has no ssh keys to the NFS clients, so the caller verifies their
# mount tables and attests via --var quiesced=yes.
#
# ⚠ THE RUNBOOK'S BLAST RADIUS WAS WRONG. It named two dependents. `ss` on CT 103
# showed FIVE distinct clients on 2026-08-18:
# 10.0.50.45 esh-docker-vm hard -> quiesced by step 1
# 10.0.250.35 esh-pve hard -> quiesced by step 2
# 10.0.50.60 esh-vm-db hard -> DELIBERATELY LEFT MOUNTED (see below)
# 10.0.50.154 vm-esh-nas n/a -> is VM 104 on THIS host; dies with it
# 10.100.10.50 nh3-dev soft,ro -> errors instead of blocking; safe
#
# esh-vm-db is left mounted on purpose. It is a backup TARGET with no live user:
# resticprofile-backup and postgresql-dump next fire ~19h out, and a hard mount
# with nothing actively using it blocks and then resumes when the server returns
# — that is what `hard` is for. Unmounting it would mean an unmount/remount cycle
# over the qemu guest agent on a host with no ssh access, where a failed remount
# breaks backups silently. Leaving it is the lower-risk branch, not the lazy one.
# The gate is `quiesced`, which the CALLER sets only after checking each client's
# mount table directly (this host has no ssh to them; see step 1/2 playbooks).
#
# ⚠ It deliberately does NOT gate on server-side NFS session count. Measured
# 2026-08-18: esh-docker-vm's sessions drained within ~90s, but esh-pve held 11
# established connections to :2049 indefinitely with NO mounts in either
# `findmnt` or `/proc/mounts` and nothing holding a cwd there. That is the Linux
# NFSv4 client keeping its transport alive past the last unmount, and it is the
# wrong thing to gate on: the failure this whole runbook exists to prevent is a
# process blocking on a MOUNTED hard filesystem when the server vanishes. With no
# mount there is nothing to block on — an idle socket to a departing server just
# resets. Gating on sessions would have stalled the window forever on a condition
# that never clears and never mattered.
- name: GUARD — caller has confirmed both hard-NFS clients are unmounted
shell: |
test "{{ quiesced }}" = "yes" || {
echo "run playbooks 1 and 2 and confirm client mount tables first"; exit 1; }
echo "caller attests: esh-docker-vm and esh-pve carry no esh-nas mounts"
echo "--- server-side sessions, informational only ---"
pct exec 103 -- ss -tnH state established '( sport = :2049 )' 2>/dev/null \
| awk '{print $4}' | sed 's/:[0-9]*$//' | sort | uniq -c || true
changed_when: "false"
- name: GUARD — staging artifacts are all present
shell: |
mountpoint -q {{ newroot }} || { echo "{{ newroot }} not mounted"; exit 1; }
mountpoint -q {{ newroot }}/boot || { echo "boot LV not in the chroot"; exit 1; }
grep -q pve-zfs-root {{ newroot }}/boot/grub/grub.cfg || { echo "no ZFS entry"; exit 1; }
grep -q 'saved_entry=pve-ext4-rollback' {{ newroot }}/boot/grub/grubenv || { echo "grubenv not pinned to rollback"; exit 1; }
changed_when: "false"
# ---------- stop the guests, NAS last ----------
- name: Stop the guests (reverse of startup order — CT 103, the NAS, goes last)
shell: |
for v in 105 106 107; do pct status $v 2>/dev/null | grep -q running && pct shutdown $v --timeout 90 || true; done
qm status 104 2>/dev/null | grep -q running && qm shutdown 104 --timeout 90 || true
for i in $(seq 1 30); do
running=$( (pct list | awk 'NR>1 && $2=="running"'; qm list | awk 'NR>1 && $3=="running"') | wc -l )
[ "$running" -le 1 ] && break
sleep 3
done
pct status 103 2>/dev/null | grep -q running && pct shutdown 103 --timeout 90 || true
sleep 3
echo "--- remaining ---"; pct list; qm list
changed_when: "true"
# ---------- the actual cutover ----------
- name: Point the ESP at the new /boot LV
shell: |
chroot {{ newroot }} grub-install --target=x86_64-efi \
--efi-directory=/boot/efi --bootloader-id=proxmox
changed_when: "true"
- name: Verify the ESP stub now points at the /boot LV, not the ext4 root
shell: |
BOOT_UUID=$(blkid -s UUID -o value /dev/mapper/pve-boot)
grep -q "$BOOT_UUID" {{ newroot }}/boot/efi/EFI/proxmox/grub.cfg || {
echo "ESP stub does NOT reference the boot LV — aborting before reboot"; exit 1; }
echo "ESP stub -> boot LV $BOOT_UUID"
changed_when: "false"
- name: Arm the ONE-SHOT ZFS boot (default stays pinned to the ext4 rollback)
shell: |
chroot {{ newroot }} grub-reboot pve-zfs-root
grep -o 'next_entry=.*' {{ newroot }}/boot/grub/grubenv
grep -o 'saved_entry=.*' {{ newroot }}/boot/grub/grubenv
changed_when: "true"
# ---------- tear the chroot down so the dataset can be unmounted ----------
- name: Unmount the chroot, innermost first
shell: |
for m in proc/sys/fs/binfmt_misc proc sys dev/pts dev/shm dev/mqueue dev/hugepages dev boot/efi boot; do
mountpoint -q {{ newroot }}/$m && umount -R {{ newroot }}/$m 2>/dev/null || true
done
findmnt -R {{ newroot }} -o TARGET | tail -n +2 || echo " (nothing left under {{ newroot }})"
changed_when: "true"
- name: Unmount the ZFS root dataset BEFORE changing its mountpoint
shell: zfs unmount {{ root_dataset }}
when: "mountpoint -q {{ newroot }}"
- name: Set the dataset's final mountpoint (safe only now that it is unmounted)
shell: |
zfs set mountpoint=/ {{ root_dataset }}
zfs get -H -o value mountpoint,canmount {{ root_dataset }} | tr '\n' ' '; echo
# paranoia: the live root must STILL be the ext4 LV at this instant
test "$(findmnt -no SOURCE /)" = "/dev/mapper/pve-root" || {
echo "ZFS MOUNTED OVER THE LIVE ROOT — do not reboot, investigate"; exit 1; }
changed_when: "true"
- name: Final pre-reboot assertion
shell: |
echo "root now: $(findmnt -no SOURCE,FSTYPE /)"
echo "dataset: $(zfs get -H -o value mounted {{ root_dataset }}) mounted, canmount=$(zfs get -H -o value canmount {{ root_dataset }})"
echo "next_entry: $(grep -o 'next_entry=.*' {{ newroot }}/boot/grub/grubenv 2>/dev/null || echo '(grubenv not readable — boot LV is unmounted, expected)')"
changed_when: "false"
- name: REBOOT — connection loss here is expected
shell: systemd-run --on-active=3 --timer-property=AccuracySec=1s /sbin/reboot
changed_when: "true"
@@ -0,0 +1,45 @@
# esh-pve-nas cutover, step 5a of 5 — restore esh-docker-vm's NFS mounts.
#
# Run: scripts/elway infra-ops@10.0.50.45 --playbook playbooks/esh-cutover-5-restore-docker-vm.yaml
#
# Reverses playbooks/esh-cutover-1-quiesce-docker-vm.yaml. Mount first, THEN start
# the container: calibre opens its SQLite library on startup, and starting it
# against an unmounted /mnt/books would have it create a fresh empty library on
# the local disk underneath the mountpoint — which then gets shadowed the moment
# the real mount lands, and looks exactly like data loss.
steps:
- name: Mount /mnt/books
shell: sudo -n mount /mnt/books
when: "! mountpoint -q /mnt/books"
- name: Mount /mnt/backup
shell: sudo -n mount /mnt/backup
when: "! mountpoint -q /mnt/backup"
- name: Confirm the library is actually there before starting calibre
shell: |
test -f /mnt/books/calibre/calibre_library/metadata.db || {
echo "calibre library NOT visible — refusing to start the container"; exit 1; }
echo "metadata.db present: $(stat -c %s /mnt/books/calibre/calibre_library/metadata.db) bytes"
changed_when: "false"
- name: Start calibre-web-automated
shell: sudo -n docker start calibre-web-automated
when: "test -z \"$(sudo -n docker ps -q --filter name=^calibre-web-automated$)\""
verify:
- name: Both NFS mounts are back
shell: |
mountpoint -q /mnt/books && mountpoint -q /mnt/backup
findmnt -no SOURCE,OPTIONS /mnt/books | grep -q hard
changed_when: "false"
- name: calibre-web-automated is running
shell: test -n "$(sudo -n docker ps -q --filter name=^calibre-web-automated$)"
changed_when: "false"
- name: Full container count is back
shell: |
n=$(sudo -n docker ps -q | wc -l); echo "$n containers running"; test "$n" -ge 17
changed_when: "false"
+29 -2
View File
@@ -86,13 +86,40 @@ steps:
mount --bind /boot/efi {{ newroot }}/boot/efi
when: "! mountpoint -q {{ newroot }}/boot"
- name: Bind the kernel filesystems into the chroot
# ⚠⚠ --make-rslave IS LOAD-BEARING. Without it this cost a production outage on
# 2026-08-18.
#
# On a systemd host `/` has SHARED mount propagation, so `mount --rbind /dev`
# creates a bind that shares propagation with the original. Every later
# `umount -R` of the chroot copy then propagates BACK to the live system and
# unmounts the REAL /sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone,
# systemd-logind cannot create a session: sshd still completes authentication
# and already-resident daemons keep serving from memory, but every new exec
# hangs forever. The host looks alive and is unusable, and — this is the part
# that wasted the most time — it looks exactly like failing root-disk I/O.
#
# --make-rslave makes propagation one-way: host -> chroot only. Teardown then
# cannot reach back.
- name: Bind the kernel filesystems into the chroot (SLAVE propagation)
shell: |
for d in dev dev/pts proc sys; do
for d in dev proc sys; do
mountpoint -q {{ newroot }}/$d || mount --rbind /$d {{ newroot }}/$d
mount --make-rslave {{ newroot }}/$d
done
echo "--- propagation (must NOT say shared) ---"
findmnt -o TARGET,PROPAGATION {{ newroot }}/dev {{ newroot }}/sys {{ newroot }}/proc
changed_when: "true"
- name: GUARD — refuse to continue if any chroot bind is still shared
shell: |
if findmnt -no PROPAGATION -R {{ newroot }}/dev {{ newroot }}/sys {{ newroot }}/proc \
2>/dev/null | grep -q shared; then
echo "chroot binds are SHARED — teardown would unmount the live host's /sys and /dev"
exit 1
fi
echo "all chroot binds are private/slave — teardown cannot propagate back"
changed_when: "false"
# ---------- build the boot artifacts inside the chroot ----------
- name: Pin the default boot entry to the rollback, not to ZFS