5f11d1b3cb
Root is now nvme/ROOT/pve-1. The USB DOM keeps the ESP and /boot but is out of the runtime I/O path, so a bus reset can no longer drop root from under a running hypervisor. All five guests healthy, three pools ONLINE, system running, ext4 pve-root intact and unmounted as the rollback with its own kernel and initrd. zfs-import-cache is now the active import path -- the all-three-pools cachefile fix doing its job. The window cost an unplanned outage, and the cause was this repo's own tooling rather than the migration. The staging chroot ran `mount --rbind /dev` and /sys with no --make-rslave. On systemd `/` has shared propagation, so the cutover's `umount -R` propagated back into the live host and removed the real /sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone systemd-logind could not create a session: ping fine, TCP fine, SSH authentication succeeded, resident daemons kept serving -- and every new exec hung, including /sbin/reboot, so the reboot never ran at all. It impersonates failing root-disk I/O almost perfectly, and I called it as the DOM dying. That was wrong. dmesg had the answer throughout: the DOM attached cleanly with no errors, and the last log timestamp was 12114881s -- 140 days -- meaning this was still the original boot. A down-detector had also never reported the host down, which I read as a fast reboot rather than as no reboot. Fixes and guards: - --make-rslave after every rbind, plus a guard that refuses to proceed while any chroot bind still reports shared propagation. - Confirm a reboot by observing the host DOWN, not by watching for it to come back. Those two states are indistinguishable otherwise. - Blast radius now measured from the server: `ss` inside CT 103 found five NFS clients, not the two documented. The new one that mattered is esh-vm-db, hard-mounted and unreachable by ssh. Left mounted on purpose and it came through read-write. - grub-reboot's one-shot does NOT work here: grubenv sits on an LVM LV which GRUB reads but cannot write, so next_entry survived the boot that consumed it. Steady state is saved_entry=pve-zfs-root with no next_entry. There is no auto-fallback on this host and no IPMI. Recovery needed no console: an idempotent cgroup2/devpts/shm remount landed in the brief windows where exec succeeded. No data was lost, and neither the DOM nor any pool was ever at risk.
46 lines
1.8 KiB
YAML
46 lines
1.8 KiB
YAML
# esh-pve-nas cutover, step 5a of 5 — restore esh-docker-vm's NFS mounts.
|
|
#
|
|
# Run: scripts/elway infra-ops@10.0.50.45 --playbook playbooks/esh-cutover-5-restore-docker-vm.yaml
|
|
#
|
|
# Reverses playbooks/esh-cutover-1-quiesce-docker-vm.yaml. Mount first, THEN start
|
|
# the container: calibre opens its SQLite library on startup, and starting it
|
|
# against an unmounted /mnt/books would have it create a fresh empty library on
|
|
# the local disk underneath the mountpoint — which then gets shadowed the moment
|
|
# the real mount lands, and looks exactly like data loss.
|
|
|
|
steps:
|
|
- name: Mount /mnt/books
|
|
shell: sudo -n mount /mnt/books
|
|
when: "! mountpoint -q /mnt/books"
|
|
|
|
- name: Mount /mnt/backup
|
|
shell: sudo -n mount /mnt/backup
|
|
when: "! mountpoint -q /mnt/backup"
|
|
|
|
- name: Confirm the library is actually there before starting calibre
|
|
shell: |
|
|
test -f /mnt/books/calibre/calibre_library/metadata.db || {
|
|
echo "calibre library NOT visible — refusing to start the container"; exit 1; }
|
|
echo "metadata.db present: $(stat -c %s /mnt/books/calibre/calibre_library/metadata.db) bytes"
|
|
changed_when: "false"
|
|
|
|
- name: Start calibre-web-automated
|
|
shell: sudo -n docker start calibre-web-automated
|
|
when: "test -z \"$(sudo -n docker ps -q --filter name=^calibre-web-automated$)\""
|
|
|
|
verify:
|
|
- name: Both NFS mounts are back
|
|
shell: |
|
|
mountpoint -q /mnt/books && mountpoint -q /mnt/backup
|
|
findmnt -no SOURCE,OPTIONS /mnt/books | grep -q hard
|
|
changed_when: "false"
|
|
|
|
- name: calibre-web-automated is running
|
|
shell: test -n "$(sudo -n docker ps -q --filter name=^calibre-web-automated$)"
|
|
changed_when: "false"
|
|
|
|
- name: Full container count is back
|
|
shell: |
|
|
n=$(sudo -n docker ps -q | wc -l); echo "$n containers running"; test "$n" -ge 17
|
|
changed_when: "false"
|