c4b2278e7d
Everything but the reboot. Two rerunnable elway playbooks; the host is
still running from the ext4 root and its boot path is byte-identical to
the last 140 days, because grub-install is deliberately held back to the
cutover window.
Phase 1 (esh-pve-nas-stage-zfs-root.yaml): carve a 512 MB /boot LV out
of the 768 MB swap LV, populate it, rsync the 4.3 GB ext4 root into
nvme/ROOT/pve-1, write the copy's fstab.
Phase 2 (esh-pve-nas-stage-bootloader.yaml): ZFS initramfs, grub.cfg,
explicit pve-zfs-root and pve-ext4-rollback entries with stable ids,
grubenv pinned to the rollback so cutover's grub-reboot is a one-shot.
Three landmines the plan did not predict, all caught by verify steps
asserting effective state rather than by reading the plan:
- The /boot LV had nowhere to live. VG pve had 4 MB free and mounted
ext4 cannot shrink; freeing space from root needs a rescue boot, which
costs the one-reboot property. Space came from swap (768M -> 256M).
- The runbook's `zpool set cachefile=... nvme` would have broken the
NAS. Populating a cachefile flips the host from import-by-scan to
import-by-cache, so a one-pool cache leaves ssd and tank unimported --
and CT 103 esh-nas has twelve bind mounts spanning all three pools.
Set on all three instead, verified in the resulting cache.
- update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, which
boots to an initramfs prompt. Debian's 10_linux builds ${rpool}${bootfs}
and rpool comes from grub-probe --target=fs_label, which returns empty
because GRUB's ZFS reader cannot open a pool with encryption,
large_dnode and zstd_compress -- the same feature set that forced /boot
to stay ext4. The probe failure is swallowed by `2>/dev/null || true`.
Fixed with a /etc/default/grub.d drop-in plus explicit menu entries.
The transferable lesson: the original verify grepped for the correct
root= string appearing somewhere in grub.cfg, which passes while every
menu entry is still broken. Assert the effective value, not the presence
of a substring.
197 lines
11 KiB
Markdown
197 lines
11 KiB
Markdown
# esh-pve-nas — PVE root on a USB DOM: diagnosis, mitigation, migration plan
|
|
|
|
## The finding
|
|
|
|
`esh-pve-nas` (`esh-nas-pve.esteban.net`, 10.0.50.55) runs PVE root off a **USB
|
|
Disk-on-Module** — `sdq`, 7.3 GB, `ID_BUS=usb`, `ID_VENDOR=NORELSYS`, model 1081 —
|
|
carved into a 512 MB ESP + 768 MB swap + a **6 GB ext4 root** that was at **90%
|
|
(571 MB free)**.
|
|
|
|
⚠ **Operator corrected my first read: it is a DOM, not a thumb drive.** DOMs use
|
|
SLC/pSLC with a real controller, so the **284 GB written since boot is
|
|
unremarkable and wear is NOT the driver**. I had framed it as a clock ticking;
|
|
that was wrong and the correction matters. What actually justifies the work:
|
|
|
|
1. **It is on the USB bus** — a reset or re-enumeration drops the *root
|
|
filesystem* out from under a running hypervisor whose guests keep executing.
|
|
NAND quality is irrelevant to that.
|
|
2. **6 GB has no headroom** — `/usr` alone is 3.7 GB.
|
|
3. **Unmirrored**, while 928 GB of mirrored NVMe sits 96% empty.
|
|
4. **It has blocked patching for months** — the operator-visible symptom and the
|
|
real urgency.
|
|
|
|
## The patching blockage (measured)
|
|
|
|
`apt-get -s dist-upgrade`: **225 packages pending, 161 carrying `deb12uN` /
|
|
Debian-Security bumps** including `ssh 1:9.2p1-2+deb12u10`. Host sits on
|
|
`pve-manager/8.4.11` vs sibling esh-pve's **8.4.14**, with 20 weeks uptime
|
|
because it cannot take a kernel.
|
|
|
|
⚠ **Ordering is load-bearing: migrate FIRST, patch after.** The pending set
|
|
includes `proxmox-kernel-6.8.12-42-pve-signed` — ~250 MB of kernel + initramfs
|
|
landing in `/boot`, **which is on root**. Unpacking 225 packages (dpkg, perl,
|
|
glibc-adjacent) into 1.3 GB of headroom risks filling the disk mid-transaction
|
|
and wedging dpkg on a hypervisor running five guests. Partial escape hatch if
|
|
patching truly cannot wait: `apt-get -o Dir::Cache::Archives=/nvme/tmp/apt-archives`
|
|
keeps downloads off root, but the kernel still lands in `/boot`.
|
|
|
|
## Mitigation applied 2026-08-17 — root 90% → 76%
|
|
|
|
| step | effect |
|
|
|---|---|
|
|
| capped journald (`SystemMaxUse=64M`; was **fully default/uncapped**) | stops unbounded growth |
|
|
| vacuumed the journal | **freed 446 MB** |
|
|
| `apt-get clean` | 79 MB |
|
|
| `/root/neo` (2024 Intel NEO OpenCL debs) → `/nvme/tmp/root-neo-20260817/` | 259 MB — **moved, not deleted** |
|
|
| **`/var/log/journal` relocated onto ZFS** (`nvme/varlog`) | dominant writer off the DOM |
|
|
|
|
571 MB → **1.4 GB free**. All five guests stayed up; a fresh `logger` round-tripped
|
|
through the ZFS-backed journal.
|
|
|
|
⚠ **Stopping journald over SSH kills your own session** — it takes the
|
|
connection's logging path with it. The first attempt died mid-swap, leaving the
|
|
dataset staged and the move incomplete (host was never at risk; journald
|
|
socket-activated straight back). Redo as a detached `systemd-run` transient unit.
|
|
Script + reason live at `root@10.0.50.55:/root/move-journal-to-zfs.sh`.
|
|
|
|
Deliberately **not** done: moving `/var/lib/rrdcached`. With the DOM correction
|
|
the wear argument no longer justifies touching a service `pvestatd` depends on.
|
|
|
|
## The plan — split boot from root (operator's proposal, strictly better)
|
|
|
|
My first plan was a full reinstall to a mirrored-NVMe ZFS root. **The operator
|
|
proposed keeping boot on the DOM with a fallback image and putting all its files
|
|
on ZFS. That is better and I should have gotten there myself** — I had assumed
|
|
boot and root must share a device.
|
|
|
|
| | device | contents | written when |
|
|
|---|---|---|---|
|
|
| boot | DOM `sdq` | ESP + `/boot` (ext4) | only on kernel/GRUB updates |
|
|
| root | `nvme` pool | `nvme/ROOT/pve-1` | constantly, on mirrored NVMe |
|
|
|
|
Keeping `/boot` on **ext4** is the point, not a compromise: GRUB never has to read
|
|
ZFS, which matters because the `nvme` pool has `encryption`, `large_dnode` and
|
|
`zstd_compress` enabled and **GRUB cannot read those**.
|
|
|
|
**Why it beats the reinstall:** the `nvme` pool survives (no guest migration, no
|
|
`ssd`/`tank` export-import, no reinstall); downtime is **one reboot** not half a
|
|
day; **rollback is a GRUB menu entry** because the ext4 root stays untouched on
|
|
the DOM; and it retires the actual top risk — with root on NVMe a USB bus reset
|
|
mid-run no longer kills the running system. Free upside: boot environments
|
|
(`zfs snapshot nvme/ROOT/pve-1@pre-upgrade`).
|
|
|
|
**Preconditions verified already met:** UEFI + `grub-efi-amd64 2.06-13+pmx7`;
|
|
**`zfs-initramfs 2.2.8-pve1` already installed with 76 ZFS files in the running
|
|
initrd**; root only 4.3 GB to copy; swap 767 MB / 123 MB used against 125 GB RAM
|
|
(leave it on the DOM LV — **never** swap on a zvol).
|
|
|
|
**Two traps:** `canmount=noauto` on the root dataset or ZFS mounts over the live
|
|
root; and `cachefile` is `none` with a **0-byte `/etc/zfs/zpool.cache`** — pools
|
|
import by scan today, which is a coin-flip when the initramfs must find root.
|
|
Set the cachefile before rebuilding the initramfs.
|
|
|
|
Operator ruled a **cloned DOM image is sufficient** boot-path insurance (no
|
|
mirrored boot needed). `dd` it off-box before anything else; refresh after kernel
|
|
updates.
|
|
|
|
## ⚠ Blast radius — the gating constraint, invisible from the host itself
|
|
|
|
**CT 103 `esh-nas` (10.0.50.50) IS the NAS, and it runs on this host.** Two
|
|
dependents mount it over **`hard`** NFS — they do not fail, they hang unkillably:
|
|
|
|
- **esh-docker-vm** (10.0.50.45): `/mnt/books`, `/mnt/backup`
|
|
- **esh-pve** (10.0.250.35): `/mnt/pve/esh-nas`, `/mnt/pve/tank-vmbu`
|
|
|
|
Known incident shape — the only remedy for esh-docker-vm's D-state is a host
|
|
reboot, and `/mnt/books` was *deliberately* left `hard` because calibre's SQLite
|
|
risks corruption under `soft`. Quiesce both before any reboot of this host.
|
|
Recorded in `servers/esh-pve-nas/README.md` as a never-reboot-casually warning.
|
|
|
|
## Also identified
|
|
|
|
- **`esh-nas` is CT 103** on esh-pve-nas — structurally the same shape as ana-nas
|
|
being CT 109 on pfi-pve.
|
|
- **`ESH-FileBot` (CT 106, 10.0.50.70) is an empty shell** — 80 GB rootfs, six
|
|
passthrough mounts (`books`/`documents`/`music`/`share`/`pvestore`/`ssd-pvestore`),
|
|
and **nothing running but base systemd, sshd, cron, postfix** since 30 March.
|
|
That resolves the dashboard's long-standing "role TBC". Retire rather than
|
|
migrate.
|
|
- Both ESH hypervisors have **20 weeks uptime** and differing PVE patch levels.
|
|
|
|
## Staging executed 2026-08-18 — everything but the reboot
|
|
|
|
Two rerunnable elway playbooks, 0 failed steps, 17/17 verify green:
|
|
`playbooks/esh-pve-nas-stage-zfs-root.yaml` (LV surgery, `/boot` populate,
|
|
4.3 GB root rsync in 228 s, fstab) and `playbooks/esh-pve-nas-stage-bootloader.yaml`
|
|
(ZFS initramfs, grub.cfg, both menu entries, grubenv).
|
|
|
|
**`grub-install` is deliberately NOT run.** The ESP stub still points at the old
|
|
`/boot` inside the ext4 root, so the host's boot path is byte-identical to the
|
|
last 140 days and an unplanned reboot mid-staging is a non-event. Cutover is
|
|
`grub-install` + `grub-reboot pve-zfs-root` + `zfs set mountpoint=/` + reboot.
|
|
|
|
Final DOM layout: `pve-root` 6.04 G (untouched, the rollback) + `pve-boot` 512 M
|
|
(new) + `pve-swap` 256 M (was 768 M).
|
|
|
|
### The three landmines staging found
|
|
|
|
1. **The `/boot` LV had nowhere to live.** VG `pve` had **4 MB free**, and
|
|
mounted ext4 cannot shrink — freeing space from root needs a rescue boot,
|
|
which costs the "one reboot" property the design rests on. Only live source
|
|
was the swap LV. Operator chose shrink-to-256M over drop-entirely.
|
|
2. **The one-pool cachefile would have broken the NAS.** `zpool set
|
|
cachefile=… nvme` looks scoped and safe; it is the opposite. Populating a
|
|
cachefile flips the host from `zfs-import-scan` to `zfs-import-cache`
|
|
(verified: scan active, cache inactive beforehand), so a cache holding only
|
|
`nvme` leaves `ssd` and `tank` unimported at boot — and CT 103 has twelve
|
|
bind mounts spanning all three pools. Every export would come up empty and
|
|
both `hard` NFS clients would hang.
|
|
3. **`update-grub` silently emitted a pool-less `root=ZFS=/ROOT/pve-1`.**
|
|
Debian's `10_linux` builds `${rpool}${bootfs}`; `rpool` comes from
|
|
`grub-probe --target=fs_label`, which returns empty because GRUB's ZFS reader
|
|
cannot open a pool with `encryption`/`large_dnode`/`zstd_compress` — and the
|
|
probe failure is swallowed by `2>/dev/null || true`. The same feature set
|
|
that forced `/boot` to stay ext4 also corrupts the kernel command line, which
|
|
the design did not anticipate. Fixed with a `/etc/default/grub.d/zfs-root.cfg`
|
|
drop-in (last `root=` wins) plus explicit `pve-zfs-root` and
|
|
`pve-ext4-rollback` entries carrying stable ids — the auto-generated ids are
|
|
derived from pool member device paths and would shift if the mirror changed.
|
|
|
|
**The transferable lesson from (3):** the original verify grepped for
|
|
`root=ZFS=nvme/ROOT/pve-1` *appearing somewhere* in grub.cfg. Once the drop-in
|
|
was added that grep passes — while pool-less entries sit in the menu untouched.
|
|
The check that holds walks every `linux` line, takes the **last** `root=`, and
|
|
asserts it against a known-good set. **Assert the effective value, not the
|
|
presence of a substring.**
|
|
|
|
### One-shot boot, not a new default
|
|
|
|
`GRUB_DEFAULT=saved` with grubenv pinned to `pve-ext4-rollback`, and cutover uses
|
|
`grub-reboot pve-zfs-root` so ZFS is tried **exactly once**. A failed ZFS boot
|
|
returns to ext4 by itself on the next reboot — no console, no hands. That matters
|
|
more here than on a normal host: a hang at an initramfs prompt takes CT 103 down
|
|
and the NFS clients hang rather than fail. Only after a second clean ZFS boot
|
|
should the saved default move.
|
|
|
|
### Off-box artifacts (`nh3-dev:~/backups/esh-pve-nas/`)
|
|
|
|
- `dom-sdq-20260818.img.zst` — full DOM image, 7,837,450,240 B raw / 2.38 GiB
|
|
compressed, zstd XXH64 verified. ⚠ **Crash-consistent, not clean** — the root
|
|
LV was live during the read, so a restore replays the ext4 journal. Not
|
|
fixable with an LVM snapshot: the VG has no free extents.
|
|
- `bootchain-20260818.tar.gz` — clean, consistent tar of `/boot` + ESP (88 MB,
|
|
644 entries, full proxmox shim/grub EFI chain). This is the higher-quality
|
|
boot-chain artifact; the dd image is the belt-and-braces full-device restore.
|
|
- `pve-config-snapshot-20260818T051*.tar.gz` — 147 entries incl. the new
|
|
grub.cfg, fstab, LVM/ZFS/blkid state.
|
|
⚠ Building this the first time produced a **corrupt archive**: `pvs; vgs; lvs >
|
|
file` redirects only the last command, so `pvs`/`vgs` output leaked into the
|
|
tar stream on stdout. Group with `{ …; } > file`.
|
|
|
|
Runbook: `docs/runbooks/esh-pve-nas-boot-migration.md`. Earlier config snapshot at
|
|
`nh3-dev:~/backups/esh-pve-nas/pve-config-snapshot-20260818T043027Z.tar.gz` (0600,
|
|
sha256 `dc312793d027dc43…`) — `/etc/pve`, network, fstab, apt, authorized_keys plus
|
|
captured `zpool`/`zfs`/`disk-by-id`/`lsblk`-with-serials/`pvesm`/`dpkg` state and
|
|
every guest config. **The newest on-disk copy before this was June 2024.**
|
|
Commits `2275e11`, `3e31175`, `8ddc87c`.
|