Everything but the reboot. Two rerunnable elway playbooks; the host is
still running from the ext4 root and its boot path is byte-identical to
the last 140 days, because grub-install is deliberately held back to the
cutover window.
Phase 1 (esh-pve-nas-stage-zfs-root.yaml): carve a 512 MB /boot LV out
of the 768 MB swap LV, populate it, rsync the 4.3 GB ext4 root into
nvme/ROOT/pve-1, write the copy's fstab.
Phase 2 (esh-pve-nas-stage-bootloader.yaml): ZFS initramfs, grub.cfg,
explicit pve-zfs-root and pve-ext4-rollback entries with stable ids,
grubenv pinned to the rollback so cutover's grub-reboot is a one-shot.
Three landmines the plan did not predict, all caught by verify steps
asserting effective state rather than by reading the plan:
- The /boot LV had nowhere to live. VG pve had 4 MB free and mounted
ext4 cannot shrink; freeing space from root needs a rescue boot, which
costs the one-reboot property. Space came from swap (768M -> 256M).
- The runbook's `zpool set cachefile=... nvme` would have broken the
NAS. Populating a cachefile flips the host from import-by-scan to
import-by-cache, so a one-pool cache leaves ssd and tank unimported --
and CT 103 esh-nas has twelve bind mounts spanning all three pools.
Set on all three instead, verified in the resulting cache.
- update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, which
boots to an initramfs prompt. Debian's 10_linux builds ${rpool}${bootfs}
and rpool comes from grub-probe --target=fs_label, which returns empty
because GRUB's ZFS reader cannot open a pool with encryption,
large_dnode and zstd_compress -- the same feature set that forced /boot
to stay ext4. The probe failure is swallowed by `2>/dev/null || true`.
Fixed with a /etc/default/grub.d drop-in plus explicit menu entries.
The transferable lesson: the original verify grepped for the correct
root= string appearing somewhere in grub.cfg, which passes while every
menu entry is still broken. Assert the effective value, not the presence
of a substring.
11 KiB
esh-pve-nas — PVE root on a USB DOM: diagnosis, mitigation, migration plan
The finding
esh-pve-nas (esh-nas-pve.esteban.net, 10.0.50.55) runs PVE root off a USB
Disk-on-Module — sdq, 7.3 GB, ID_BUS=usb, ID_VENDOR=NORELSYS, model 1081 —
carved into a 512 MB ESP + 768 MB swap + a 6 GB ext4 root that was at 90%
(571 MB free).
⚠ Operator corrected my first read: it is a DOM, not a thumb drive. DOMs use SLC/pSLC with a real controller, so the 284 GB written since boot is unremarkable and wear is NOT the driver. I had framed it as a clock ticking; that was wrong and the correction matters. What actually justifies the work:
- It is on the USB bus — a reset or re-enumeration drops the root filesystem out from under a running hypervisor whose guests keep executing. NAND quality is irrelevant to that.
- 6 GB has no headroom —
/usralone is 3.7 GB. - Unmirrored, while 928 GB of mirrored NVMe sits 96% empty.
- It has blocked patching for months — the operator-visible symptom and the real urgency.
The patching blockage (measured)
apt-get -s dist-upgrade: 225 packages pending, 161 carrying deb12uN /
Debian-Security bumps including ssh 1:9.2p1-2+deb12u10. Host sits on
pve-manager/8.4.11 vs sibling esh-pve's 8.4.14, with 20 weeks uptime
because it cannot take a kernel.
⚠ Ordering is load-bearing: migrate FIRST, patch after. The pending set
includes proxmox-kernel-6.8.12-42-pve-signed — ~250 MB of kernel + initramfs
landing in /boot, which is on root. Unpacking 225 packages (dpkg, perl,
glibc-adjacent) into 1.3 GB of headroom risks filling the disk mid-transaction
and wedging dpkg on a hypervisor running five guests. Partial escape hatch if
patching truly cannot wait: apt-get -o Dir::Cache::Archives=/nvme/tmp/apt-archives
keeps downloads off root, but the kernel still lands in /boot.
Mitigation applied 2026-08-17 — root 90% → 76%
| step | effect |
|---|---|
capped journald (SystemMaxUse=64M; was fully default/uncapped) |
stops unbounded growth |
| vacuumed the journal | freed 446 MB |
apt-get clean |
79 MB |
/root/neo (2024 Intel NEO OpenCL debs) → /nvme/tmp/root-neo-20260817/ |
259 MB — moved, not deleted |
/var/log/journal relocated onto ZFS (nvme/varlog) |
dominant writer off the DOM |
571 MB → 1.4 GB free. All five guests stayed up; a fresh logger round-tripped
through the ZFS-backed journal.
⚠ Stopping journald over SSH kills your own session — it takes the
connection's logging path with it. The first attempt died mid-swap, leaving the
dataset staged and the move incomplete (host was never at risk; journald
socket-activated straight back). Redo as a detached systemd-run transient unit.
Script + reason live at root@10.0.50.55:/root/move-journal-to-zfs.sh.
Deliberately not done: moving /var/lib/rrdcached. With the DOM correction
the wear argument no longer justifies touching a service pvestatd depends on.
The plan — split boot from root (operator's proposal, strictly better)
My first plan was a full reinstall to a mirrored-NVMe ZFS root. The operator proposed keeping boot on the DOM with a fallback image and putting all its files on ZFS. That is better and I should have gotten there myself — I had assumed boot and root must share a device.
| device | contents | written when | |
|---|---|---|---|
| boot | DOM sdq |
ESP + /boot (ext4) |
only on kernel/GRUB updates |
| root | nvme pool |
nvme/ROOT/pve-1 |
constantly, on mirrored NVMe |
Keeping /boot on ext4 is the point, not a compromise: GRUB never has to read
ZFS, which matters because the nvme pool has encryption, large_dnode and
zstd_compress enabled and GRUB cannot read those.
Why it beats the reinstall: the nvme pool survives (no guest migration, no
ssd/tank export-import, no reinstall); downtime is one reboot not half a
day; rollback is a GRUB menu entry because the ext4 root stays untouched on
the DOM; and it retires the actual top risk — with root on NVMe a USB bus reset
mid-run no longer kills the running system. Free upside: boot environments
(zfs snapshot nvme/ROOT/pve-1@pre-upgrade).
Preconditions verified already met: UEFI + grub-efi-amd64 2.06-13+pmx7;
zfs-initramfs 2.2.8-pve1 already installed with 76 ZFS files in the running
initrd; root only 4.3 GB to copy; swap 767 MB / 123 MB used against 125 GB RAM
(leave it on the DOM LV — never swap on a zvol).
Two traps: canmount=noauto on the root dataset or ZFS mounts over the live
root; and cachefile is none with a 0-byte /etc/zfs/zpool.cache — pools
import by scan today, which is a coin-flip when the initramfs must find root.
Set the cachefile before rebuilding the initramfs.
Operator ruled a cloned DOM image is sufficient boot-path insurance (no
mirrored boot needed). dd it off-box before anything else; refresh after kernel
updates.
⚠ Blast radius — the gating constraint, invisible from the host itself
CT 103 esh-nas (10.0.50.50) IS the NAS, and it runs on this host. Two
dependents mount it over hard NFS — they do not fail, they hang unkillably:
- esh-docker-vm (10.0.50.45):
/mnt/books,/mnt/backup - esh-pve (10.0.250.35):
/mnt/pve/esh-nas,/mnt/pve/tank-vmbu
Known incident shape — the only remedy for esh-docker-vm's D-state is a host
reboot, and /mnt/books was deliberately left hard because calibre's SQLite
risks corruption under soft. Quiesce both before any reboot of this host.
Recorded in servers/esh-pve-nas/README.md as a never-reboot-casually warning.
Also identified
esh-nasis CT 103 on esh-pve-nas — structurally the same shape as ana-nas being CT 109 on pfi-pve.ESH-FileBot(CT 106, 10.0.50.70) is an empty shell — 80 GB rootfs, six passthrough mounts (books/documents/music/share/pvestore/ssd-pvestore), and nothing running but base systemd, sshd, cron, postfix since 30 March. That resolves the dashboard's long-standing "role TBC". Retire rather than migrate.- Both ESH hypervisors have 20 weeks uptime and differing PVE patch levels.
Staging executed 2026-08-18 — everything but the reboot
Two rerunnable elway playbooks, 0 failed steps, 17/17 verify green:
playbooks/esh-pve-nas-stage-zfs-root.yaml (LV surgery, /boot populate,
4.3 GB root rsync in 228 s, fstab) and playbooks/esh-pve-nas-stage-bootloader.yaml
(ZFS initramfs, grub.cfg, both menu entries, grubenv).
grub-install is deliberately NOT run. The ESP stub still points at the old
/boot inside the ext4 root, so the host's boot path is byte-identical to the
last 140 days and an unplanned reboot mid-staging is a non-event. Cutover is
grub-install + grub-reboot pve-zfs-root + zfs set mountpoint=/ + reboot.
Final DOM layout: pve-root 6.04 G (untouched, the rollback) + pve-boot 512 M
(new) + pve-swap 256 M (was 768 M).
The three landmines staging found
- The
/bootLV had nowhere to live. VGpvehad 4 MB free, and mounted ext4 cannot shrink — freeing space from root needs a rescue boot, which costs the "one reboot" property the design rests on. Only live source was the swap LV. Operator chose shrink-to-256M over drop-entirely. - The one-pool cachefile would have broken the NAS.
zpool set cachefile=… nvmelooks scoped and safe; it is the opposite. Populating a cachefile flips the host fromzfs-import-scantozfs-import-cache(verified: scan active, cache inactive beforehand), so a cache holding onlynvmeleavesssdandtankunimported at boot — and CT 103 has twelve bind mounts spanning all three pools. Every export would come up empty and bothhardNFS clients would hang. update-grubsilently emitted a pool-lessroot=ZFS=/ROOT/pve-1. Debian's10_linuxbuilds${rpool}${bootfs};rpoolcomes fromgrub-probe --target=fs_label, which returns empty because GRUB's ZFS reader cannot open a pool withencryption/large_dnode/zstd_compress— and the probe failure is swallowed by2>/dev/null || true. The same feature set that forced/bootto stay ext4 also corrupts the kernel command line, which the design did not anticipate. Fixed with a/etc/default/grub.d/zfs-root.cfgdrop-in (lastroot=wins) plus explicitpve-zfs-rootandpve-ext4-rollbackentries carrying stable ids — the auto-generated ids are derived from pool member device paths and would shift if the mirror changed.
The transferable lesson from (3): the original verify grepped for
root=ZFS=nvme/ROOT/pve-1 appearing somewhere in grub.cfg. Once the drop-in
was added that grep passes — while pool-less entries sit in the menu untouched.
The check that holds walks every linux line, takes the last root=, and
asserts it against a known-good set. Assert the effective value, not the
presence of a substring.
One-shot boot, not a new default
GRUB_DEFAULT=saved with grubenv pinned to pve-ext4-rollback, and cutover uses
grub-reboot pve-zfs-root so ZFS is tried exactly once. A failed ZFS boot
returns to ext4 by itself on the next reboot — no console, no hands. That matters
more here than on a normal host: a hang at an initramfs prompt takes CT 103 down
and the NFS clients hang rather than fail. Only after a second clean ZFS boot
should the saved default move.
Off-box artifacts (nh3-dev:~/backups/esh-pve-nas/)
dom-sdq-20260818.img.zst— full DOM image, 7,837,450,240 B raw / 2.38 GiB compressed, zstd XXH64 verified. ⚠ Crash-consistent, not clean — the root LV was live during the read, so a restore replays the ext4 journal. Not fixable with an LVM snapshot: the VG has no free extents.bootchain-20260818.tar.gz— clean, consistent tar of/boot+ ESP (88 MB, 644 entries, full proxmox shim/grub EFI chain). This is the higher-quality boot-chain artifact; the dd image is the belt-and-braces full-device restore.pve-config-snapshot-20260818T051*.tar.gz— 147 entries incl. the new grub.cfg, fstab, LVM/ZFS/blkid state. ⚠ Building this the first time produced a corrupt archive:pvs; vgs; lvs > fileredirects only the last command, sopvs/vgsoutput leaked into the tar stream on stdout. Group with{ …; } > file.
Runbook: docs/runbooks/esh-pve-nas-boot-migration.md. Earlier config snapshot at
nh3-dev:~/backups/esh-pve-nas/pve-config-snapshot-20260818T043027Z.tar.gz (0600,
sha256 dc312793d027dc43…) — /etc/pve, network, fstab, apt, authorized_keys plus
captured zpool/zfs/disk-by-id/lsblk-with-serials/pvesm/dpkg state and
every guest config. The newest on-disk copy before this was June 2024.
Commits 2275e11, 3e31175, 8ddc87c.