Files
esh-pfi-infrastructure/playbooks/esh-cutover-5-restore-docker-vm.yaml
vh 5f11d1b3cb feat(esh-pve-nas): cut PVE root over to ZFS on the mirrored NVMe
Root is now nvme/ROOT/pve-1. The USB DOM keeps the ESP and /boot but is
out of the runtime I/O path, so a bus reset can no longer drop root from
under a running hypervisor. All five guests healthy, three pools ONLINE,
system running, ext4 pve-root intact and unmounted as the rollback with
its own kernel and initrd. zfs-import-cache is now the active import
path -- the all-three-pools cachefile fix doing its job.

The window cost an unplanned outage, and the cause was this repo's own
tooling rather than the migration.

The staging chroot ran `mount --rbind /dev` and /sys with no
--make-rslave. On systemd `/` has shared propagation, so the cutover's
`umount -R` propagated back into the live host and removed the real
/sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone systemd-logind
could not create a session: ping fine, TCP fine, SSH authentication
succeeded, resident daemons kept serving -- and every new exec hung,
including /sbin/reboot, so the reboot never ran at all.

It impersonates failing root-disk I/O almost perfectly, and I called it
as the DOM dying. That was wrong. dmesg had the answer throughout: the
DOM attached cleanly with no errors, and the last log timestamp was
12114881s -- 140 days -- meaning this was still the original boot. A
down-detector had also never reported the host down, which I read as a
fast reboot rather than as no reboot.

Fixes and guards:

- --make-rslave after every rbind, plus a guard that refuses to proceed
  while any chroot bind still reports shared propagation.
- Confirm a reboot by observing the host DOWN, not by watching for it to
  come back. Those two states are indistinguishable otherwise.
- Blast radius now measured from the server: `ss` inside CT 103 found
  five NFS clients, not the two documented. The new one that mattered is
  esh-vm-db, hard-mounted and unreachable by ssh. Left mounted on
  purpose and it came through read-write.
- grub-reboot's one-shot does NOT work here: grubenv sits on an LVM LV
  which GRUB reads but cannot write, so next_entry survived the boot
  that consumed it. Steady state is saved_entry=pve-zfs-root with no
  next_entry. There is no auto-fallback on this host and no IPMI.

Recovery needed no console: an idempotent cgroup2/devpts/shm remount
landed in the brief windows where exec succeeded. No data was lost, and
neither the DOM nor any pool was ever at risk.
2026-08-18 06:39:34 -07:00

46 lines
1.8 KiB
YAML

# esh-pve-nas cutover, step 5a of 5 — restore esh-docker-vm's NFS mounts.
#
# Run: scripts/elway infra-ops@10.0.50.45 --playbook playbooks/esh-cutover-5-restore-docker-vm.yaml
#
# Reverses playbooks/esh-cutover-1-quiesce-docker-vm.yaml. Mount first, THEN start
# the container: calibre opens its SQLite library on startup, and starting it
# against an unmounted /mnt/books would have it create a fresh empty library on
# the local disk underneath the mountpoint — which then gets shadowed the moment
# the real mount lands, and looks exactly like data loss.
steps:
- name: Mount /mnt/books
shell: sudo -n mount /mnt/books
when: "! mountpoint -q /mnt/books"
- name: Mount /mnt/backup
shell: sudo -n mount /mnt/backup
when: "! mountpoint -q /mnt/backup"
- name: Confirm the library is actually there before starting calibre
shell: |
test -f /mnt/books/calibre/calibre_library/metadata.db || {
echo "calibre library NOT visible — refusing to start the container"; exit 1; }
echo "metadata.db present: $(stat -c %s /mnt/books/calibre/calibre_library/metadata.db) bytes"
changed_when: "false"
- name: Start calibre-web-automated
shell: sudo -n docker start calibre-web-automated
when: "test -z \"$(sudo -n docker ps -q --filter name=^calibre-web-automated$)\""
verify:
- name: Both NFS mounts are back
shell: |
mountpoint -q /mnt/books && mountpoint -q /mnt/backup
findmnt -no SOURCE,OPTIONS /mnt/books | grep -q hard
changed_when: "false"
- name: calibre-web-automated is running
shell: test -n "$(sudo -n docker ps -q --filter name=^calibre-web-automated$)"
changed_when: "false"
- name: Full container count is back
shell: |
n=$(sudo -n docker ps -q | wc -l); echo "$n containers running"; test "$n" -ge 17
changed_when: "false"