From c4b2278e7ddcfe5778c0c9292ea3a36efa929fba Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Mon, 17 Aug 2026 22:11:14 -0700 Subject: [PATCH] feat(esh-pve-nas): stage the PVE root migration off the USB DOM Everything but the reboot. Two rerunnable elway playbooks; the host is still running from the ext4 root and its boot path is byte-identical to the last 140 days, because grub-install is deliberately held back to the cutover window. Phase 1 (esh-pve-nas-stage-zfs-root.yaml): carve a 512 MB /boot LV out of the 768 MB swap LV, populate it, rsync the 4.3 GB ext4 root into nvme/ROOT/pve-1, write the copy's fstab. Phase 2 (esh-pve-nas-stage-bootloader.yaml): ZFS initramfs, grub.cfg, explicit pve-zfs-root and pve-ext4-rollback entries with stable ids, grubenv pinned to the rollback so cutover's grub-reboot is a one-shot. Three landmines the plan did not predict, all caught by verify steps asserting effective state rather than by reading the plan: - The /boot LV had nowhere to live. VG pve had 4 MB free and mounted ext4 cannot shrink; freeing space from root needs a rescue boot, which costs the one-reboot property. Space came from swap (768M -> 256M). - The runbook's `zpool set cachefile=... nvme` would have broken the NAS. Populating a cachefile flips the host from import-by-scan to import-by-cache, so a one-pool cache leaves ssd and tank unimported -- and CT 103 esh-nas has twelve bind mounts spanning all three pools. Set on all three instead, verified in the resulting cache. - update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, which boots to an initramfs prompt. Debian's 10_linux builds ${rpool}${bootfs} and rpool comes from grub-probe --target=fs_label, which returns empty because GRUB's ZFS reader cannot open a pool with encryption, large_dnode and zstd_compress -- the same feature set that forced /boot to stay ext4. The probe failure is swallowed by `2>/dev/null || true`. Fixed with a /etc/default/grub.d drop-in plus explicit menu entries. The transferable lesson: the original verify grepped for the correct root= string appearing somewhere in grub.cfg, which passes while every menu entry is still broken. Assert the effective value, not the presence of a substring. --- docs/runbooks/esh-pve-nas-boot-migration.md | 196 +++++++++++-- .../2026-08-17-esh-pve-nas-dom.md | 72 ++++- persistent-memory.md | 10 +- playbooks/esh-pve-nas-stage-bootloader.yaml | 271 ++++++++++++++++++ playbooks/esh-pve-nas-stage-zfs-root.yaml | 178 ++++++++++++ servers/esh-pve-nas/README.md | 33 ++- 6 files changed, 724 insertions(+), 36 deletions(-) create mode 100644 playbooks/esh-pve-nas-stage-bootloader.yaml create mode 100644 playbooks/esh-pve-nas-stage-zfs-root.yaml diff --git a/docs/runbooks/esh-pve-nas-boot-migration.md b/docs/runbooks/esh-pve-nas-boot-migration.md index 5bff988..13beb8b 100644 --- a/docs/runbooks/esh-pve-nas-boot-migration.md +++ b/docs/runbooks/esh-pve-nas-boot-migration.md @@ -1,6 +1,16 @@ # esh-pve-nas — moving PVE root off the USB DOM -**Status: PLANNED, not executed.** Mitigation landed 2026-08-17 (root 90% → 76%). +**Status: STAGED — everything but the reboot is done.** Mitigation landed +2026-08-17 (root 90% → 76%); staging landed 2026-08-18. The host is still +running from the ext4 root on the DOM and its boot path is unchanged; what +remains is the cutover window (§ Cutover). + +Staging is two playbooks, both rerunnable: + +| phase | playbook | what it did | +|---|---|---| +| 1 | `playbooks/esh-pve-nas-stage-zfs-root.yaml` | carved the `/boot` LV out of swap, populated it, rsynced the root into `nvme/ROOT/pve-1`, wrote the copy's fstab | +| 2 | `playbooks/esh-pve-nas-stage-bootloader.yaml` | ZFS initramfs, grub.cfg, rollback entry, grubenv — **without** `grub-install` | **Plan revised 2026-08-17** from "reinstall to a mirrored-NVMe ZFS root" to **"split the boot chain from the root filesystem"** — operator's proposal, and it @@ -91,9 +101,23 @@ Verified on the host 2026-08-17: on the DOM LV. **Do not put swap on a zvol** (deadlock risk) ⚠ **`cachefile` is `none` and `/etc/zfs/zpool.cache` is 0 bytes** — pools import -by scan today. For root-on-ZFS this must be deterministic: set -`zpool set cachefile=/etc/zfs/zpool.cache nvme` and rebuild the initramfs, or the -pool may not be imported early enough to find root. +by scan today (verified: `zfs-import-scan.service` active, +`zfs-import-cache.service` inactive). For root-on-ZFS this must be +deterministic, or the pool may not be imported early enough to find root. + +⚠⚠ **Set the cachefile on ALL THREE pools, not just `nvme`.** An earlier draft +of this runbook said `zpool set cachefile=/etc/zfs/zpool.cache nvme`, and that +one-pool form is a trap. Populating a cachefile flips the host from +import-by-scan to import-by-cache — so a cache containing only `nvme` means +**`ssd` and `tank` never get imported at boot.** CT 103 `esh-nas` has twelve +bind mounts spanning all three pools (`/tank/media`, `/ssd/compose`, +`/nvme/nvme-pvestore`, …), so the NAS would come up with every export empty and +both NFS clients would hang on `hard` mounts. The scoped-looking command is more +dangerous than the broad one. + +Done 2026-08-18 for `nvme`, `ssd` and `tank`; verified all three present in the +resulting 11,976-byte cache via `zdb -C -U /etc/zfs/zpool.cache`. Phase 1's +third guard step re-asserts this on every run. ## ⚠ Blast radius — unchanged, and still the gating constraint @@ -112,36 +136,154 @@ SQLite risks corruption under `soft`. The reboot in this plan is brief, but it is still a reboot — quiesce the clients first. +## Where the `/boot` LV came from — the VG was full + +The original step 7 said `/boot` could "stay inside the DOM's existing LVM as +its own small ext4 LV, or reuse the freed space once root moves off." Neither +was available: **VG `pve` had 4 MB free**, and the 6 GB root is mounted ext4, +which cannot shrink online — freeing space from it needs a rescue boot, which +would have cost the "one reboot" property the whole design rests on. + +The only space reclaimable live was the **768 MB swap LV** (123 MB in use +against 125 GB of RAM). Operator's call 2026-08-17: **shrink swap rather than +drop it.** Final layout: + +| LV | size | role | +|---|---|---| +| `pve-root` | 6.04 G | ext4 — **untouched**, the rollback root | +| `pve-boot` | 512 M | ext4 — the new `/boot` (NEW) | +| `pve-swap` | 256 M | swap (was 768 M) | + +Rejected alternatives: dropping swap outright (more kernel headroom, no OOM +cushion); `proxmox-boot-tool` on the 512 MB ESP (PVE-native and no LVM surgery, +but it reformats the ESP and downgrades rollback from "pick a menu entry" to +"restore the DOM image"); rescue-boot to shrink root (keeps swap whole, costs a +second reboot and an offline resize of the filesystem we are fleeing). + ## Sequence -**Pre-flight (no downtime)** -1. `dd` the DOM to an image, off-box. This is the rollback for the boot chain. +**Pre-flight (no downtime)** — done 2026-08-18 +1. `dd` the DOM to an off-box image. **Crash-consistent, not clean** — the root + LV is live during the read, so a restore replays the ext4 journal. That is + fine for its purpose (boot-chain insurance) and is what a snapshot backup + does anyway. Not fixable with an LVM snapshot: the VG has no free extents. 2. Refresh the config snapshot (`nh3-dev:~/backups/esh-pve-nas/`). -3. `zpool set cachefile=/etc/zfs/zpool.cache nvme`. +3. `zpool set cachefile=/etc/zfs/zpool.cache` on **`nvme`, `ssd` AND `tank`** + (see the precondition warning above — the one-pool form breaks the NAS). -**Build the ZFS root (live, no disruption)** -4. `zfs create -o mountpoint=none nvme/ROOT` - `zfs create -o mountpoint=/ -o canmount=noauto nvme/ROOT/pve-1` - `canmount=noauto` is essential — otherwise ZFS tries to mount it over the - running root. -5. Mount it somewhere temporary and `rsync -aHAX --numeric-ids` the live root - into it, excluding `/proc /sys /dev /run /mnt /nvme /ssd /tank /var/log/journal`. -6. In the copy: edit `/etc/fstab` to drop the root line, keep the ESP and swap - lines, and add the new `/boot` entry. +**Phase 1 — `playbooks/esh-pve-nas-stage-zfs-root.yaml`** (live, no disruption) +4. `zfs create -o mountpoint=none nvme/ROOT`, then `nvme/ROOT/pve-1` with + `canmount=noauto`, `compression=zstd`, `xattr=sa`, `acltype=posixacl`. + Create it with `mountpoint=none` and only set `/` at the very end — + `canmount=noauto` alone is the documented guard, but never having a dataset + that claims `/` while the ext4 root is live is the guard that cannot misfire. +5. Reclaim the swap LV into `pve-boot`, mkfs, populate from `/boot`. +6. Mount the dataset at `/mnt/newroot` and rsync the live root in. + `--one-file-system` does the exclusion work: every path the old plan listed + by hand (`/proc /sys /dev /run /nvme /ssd /tank /var/log/journal /boot`) is + already a separate mount, so it is skipped structurally rather than by a + list that can drift. +7. Write the copy's `/etc/fstab`: no root line (the initramfs mounts it), plus + `/dev/pve/boot /boot ext4`, the ESP, and swap. -**Split `/boot` onto the DOM** -7. Shrink/repartition is not needed — `/boot` can stay inside the DOM's existing - LVM as its own small ext4 LV, or reuse the freed space once root moves off. - Copy `/boot` into it and mount at `/boot`. +**Phase 2 — `playbooks/esh-pve-nas-stage-bootloader.yaml`** (live, no disruption) +8. Chroot into the copy with the boot LV and ESP mounted, then + `update-initramfs -u -k all` + `update-grub`. +9. ⚠⚠ **`grub-mkconfig` gets the ZFS root WRONG here, silently. Override it.** + See § The pool-name bug below — this is the single most dangerous thing + found during staging. +10. `GRUB_DEFAULT=saved`, plus `40_custom` carrying **both** boot paths as + hand-authored entries with stable ids (`pve-zfs-root`, `pve-ext4-rollback`), + with grubenv pinned to the **rollback**, not to ZFS (see § Cutover for why). -**Bootloader** -8. `GRUB_CMDLINE_LINUX="root=ZFS=nvme/ROOT/pve-1"`, `update-initramfs -u -k all`, - `update-grub`. **Keep the existing ext4 entry in the menu** and leave - `GRUB_TIMEOUT=5` so it is selectable. +## ⚠ The pool-name bug — the near-miss worth reading -**Cut over** -9. Quiesce the NFS clients (see blast radius), shut down guests, reboot. -10. Verify (below). If it fails, reboot and choose the ext4 entry. +Left to itself, `update-grub` on this host produces: + +``` +linux /vmlinuz-6.8.12-13-pve root=ZFS=/ROOT/pve-1 ro quiet intel_iommu=on +``` + +**The pool name is missing.** It should be `root=ZFS=nvme/ROOT/pve-1`. That +boots to an initramfs prompt — with CT 103 `esh-nas` down and both NFS clients +hanging on `hard` mounts, at whatever hour the window happens to be. + +It is not a typo, and it is not random. Debian's `/etc/grub.d/10_linux` builds +the ZFS root as `${rpool}${bootfs}`: + +| part | from | value here | +|---|---|---| +| `rpool` | `grub-probe --device --target=fs_label` | **empty** | +| `bootfs` | `make_system_path_relative_to_its_root /` | `/ROOT/pve-1` | + +`grub-probe --target=fs /` fails outright on this pool — `grub-probe: error: +unknown filesystem` — because **GRUB's own ZFS reader cannot open a pool with +`encryption`, `large_dnode` and `zstd_compress` enabled.** So `rpool` comes back +empty and concatenates to nothing. + +That is the *same* feature set that forced `/boot` to stay ext4 on the DOM. The +design already accounted for GRUB being unable to read the pool; what was missed +is that the same limitation also corrupts the kernel command line — and does it +**without an error**, because `grub-probe`'s failure is swallowed by +`2>/dev/null || true`. + +**The fix, in two layers:** + +1. `/etc/default/grub.d/zfs-root.cfg` sets + `GRUB_CMDLINE_LINUX="root=ZFS=nvme/ROOT/pve-1 boot=zfs"`. This is appended + *after* the bogus value, and both the kernel and the zfs initramfs script + take the **last** `root=` on the line — so every auto-generated entry becomes + correct. A drop-in, not an edit to `/etc/default/grub`, so a grub package + upgrade cannot revert it in a conffile merge. +2. `40_custom` carries an explicit `pve-zfs-root` entry with a single clean + `root=` and a stable id. That is what cutover's `grub-reboot` targets — the + auto-generated ids are derived from pool member device paths + (`gnulinux-simple-/dev/nvme0n1p1_/dev/nvme1n1p1`) and would shift if the + mirror ever changed. + +**The general lesson, which is the transferable part:** the phase-2 verify step +originally grepped for `root=ZFS=nvme/ROOT/pve-1` *appearing somewhere* in +grub.cfg. Once the drop-in was added that grep passes — while pool-less entries +sit in the menu untouched. The check that actually holds walks every `linux` +line, takes the **last** `root=` on it, and asserts it against a known-good set. +Assert the effective value, not the presence of a substring. +10. **`grub-install` is deliberately NOT run during staging.** The ESP stub + still points at the old `/boot` inside the ext4 root, so the host's boot + path stays byte-identical to what it has been for 140 days. Everything + error-prone is built and verified in advance; the ESP rewrite is a + two-second idempotent command held back to the window. + +**Cutover** — the remaining work, § Cutover below. + +## Cutover + +The only remaining work. Everything below the quiesce is minutes. + +1. **Quiesce the NFS clients** (see § Blast radius). On **esh-docker-vm** + (10.0.50.45) stop whatever holds `/mnt/books` and `/mnt/backup` and unmount + them; on **esh-pve** (10.0.250.35) disable the `esh-nas` and `tank-vmbu` + storages. Do this first and confirm it — a `hard` mount left live turns a + brief reboot into an unkillable D-state needing a reboot of *that* host too. +2. Shut down the five guests. +3. Point the ESP at the new `/boot` and arm the one-shot: + ``` + chroot /mnt/newroot grub-install --target=x86_64-efi \ + --efi-directory=/boot/efi --bootloader-id=proxmox + chroot /mnt/newroot grub-reboot '' + ``` +4. Set the dataset's final mountpoint, then reboot: + ``` + zfs set mountpoint=/ nvme/ROOT/pve-1 # canmount stays noauto + reboot + ``` + +**Why `grub-reboot` and not a new default.** `GRUB_DEFAULT=saved` with grubenv +pinned to the ext4 rollback means the ZFS entry is tried **exactly once**. If it +fails, the next reboot returns to ext4 *by itself* — no console, no hands. That +matters more here than on a normal host: a boot that hangs at an initramfs +prompt takes CT 103 `esh-nas` down with it, and the NFS clients hang rather than +fail. Only after the second successful ZFS boot (§ Verification) should the +saved default move to the ZFS entry with `grub-set-default`. ## Verification diff --git a/persistent-memory.d/2026-08-17-esh-pve-nas-dom.md b/persistent-memory.d/2026-08-17-esh-pve-nas-dom.md index 9119f14..fd87833 100644 --- a/persistent-memory.d/2026-08-17-esh-pve-nas-dom.md +++ b/persistent-memory.d/2026-08-17-esh-pve-nas-dom.md @@ -118,7 +118,77 @@ Recorded in `servers/esh-pve-nas/README.md` as a never-reboot-casually warning. migrate. - Both ESH hypervisors have **20 weeks uptime** and differing PVE patch levels. -Runbook: `docs/runbooks/esh-pve-nas-boot-migration.md`. Config snapshot off-box at +## Staging executed 2026-08-18 — everything but the reboot + +Two rerunnable elway playbooks, 0 failed steps, 17/17 verify green: +`playbooks/esh-pve-nas-stage-zfs-root.yaml` (LV surgery, `/boot` populate, +4.3 GB root rsync in 228 s, fstab) and `playbooks/esh-pve-nas-stage-bootloader.yaml` +(ZFS initramfs, grub.cfg, both menu entries, grubenv). + +**`grub-install` is deliberately NOT run.** The ESP stub still points at the old +`/boot` inside the ext4 root, so the host's boot path is byte-identical to the +last 140 days and an unplanned reboot mid-staging is a non-event. Cutover is +`grub-install` + `grub-reboot pve-zfs-root` + `zfs set mountpoint=/` + reboot. + +Final DOM layout: `pve-root` 6.04 G (untouched, the rollback) + `pve-boot` 512 M +(new) + `pve-swap` 256 M (was 768 M). + +### The three landmines staging found + +1. **The `/boot` LV had nowhere to live.** VG `pve` had **4 MB free**, and + mounted ext4 cannot shrink — freeing space from root needs a rescue boot, + which costs the "one reboot" property the design rests on. Only live source + was the swap LV. Operator chose shrink-to-256M over drop-entirely. +2. **The one-pool cachefile would have broken the NAS.** `zpool set + cachefile=… nvme` looks scoped and safe; it is the opposite. Populating a + cachefile flips the host from `zfs-import-scan` to `zfs-import-cache` + (verified: scan active, cache inactive beforehand), so a cache holding only + `nvme` leaves `ssd` and `tank` unimported at boot — and CT 103 has twelve + bind mounts spanning all three pools. Every export would come up empty and + both `hard` NFS clients would hang. +3. **`update-grub` silently emitted a pool-less `root=ZFS=/ROOT/pve-1`.** + Debian's `10_linux` builds `${rpool}${bootfs}`; `rpool` comes from + `grub-probe --target=fs_label`, which returns empty because GRUB's ZFS reader + cannot open a pool with `encryption`/`large_dnode`/`zstd_compress` — and the + probe failure is swallowed by `2>/dev/null || true`. The same feature set + that forced `/boot` to stay ext4 also corrupts the kernel command line, which + the design did not anticipate. Fixed with a `/etc/default/grub.d/zfs-root.cfg` + drop-in (last `root=` wins) plus explicit `pve-zfs-root` and + `pve-ext4-rollback` entries carrying stable ids — the auto-generated ids are + derived from pool member device paths and would shift if the mirror changed. + +**The transferable lesson from (3):** the original verify grepped for +`root=ZFS=nvme/ROOT/pve-1` *appearing somewhere* in grub.cfg. Once the drop-in +was added that grep passes — while pool-less entries sit in the menu untouched. +The check that holds walks every `linux` line, takes the **last** `root=`, and +asserts it against a known-good set. **Assert the effective value, not the +presence of a substring.** + +### One-shot boot, not a new default + +`GRUB_DEFAULT=saved` with grubenv pinned to `pve-ext4-rollback`, and cutover uses +`grub-reboot pve-zfs-root` so ZFS is tried **exactly once**. A failed ZFS boot +returns to ext4 by itself on the next reboot — no console, no hands. That matters +more here than on a normal host: a hang at an initramfs prompt takes CT 103 down +and the NFS clients hang rather than fail. Only after a second clean ZFS boot +should the saved default move. + +### Off-box artifacts (`nh3-dev:~/backups/esh-pve-nas/`) + +- `dom-sdq-20260818.img.zst` — full DOM image, 7,837,450,240 B raw / 2.38 GiB + compressed, zstd XXH64 verified. ⚠ **Crash-consistent, not clean** — the root + LV was live during the read, so a restore replays the ext4 journal. Not + fixable with an LVM snapshot: the VG has no free extents. +- `bootchain-20260818.tar.gz` — clean, consistent tar of `/boot` + ESP (88 MB, + 644 entries, full proxmox shim/grub EFI chain). This is the higher-quality + boot-chain artifact; the dd image is the belt-and-braces full-device restore. +- `pve-config-snapshot-20260818T051*.tar.gz` — 147 entries incl. the new + grub.cfg, fstab, LVM/ZFS/blkid state. + ⚠ Building this the first time produced a **corrupt archive**: `pvs; vgs; lvs > + file` redirects only the last command, so `pvs`/`vgs` output leaked into the + tar stream on stdout. Group with `{ …; } > file`. + +Runbook: `docs/runbooks/esh-pve-nas-boot-migration.md`. Earlier config snapshot at `nh3-dev:~/backups/esh-pve-nas/pve-config-snapshot-20260818T043027Z.tar.gz` (0600, sha256 `dc312793d027dc43…`) — `/etc/pve`, network, fstab, apt, authorized_keys plus captured `zpool`/`zfs`/`disk-by-id`/`lsblk`-with-serials/`pvesm`/`dpkg` state and diff --git a/persistent-memory.md b/persistent-memory.md index 4b4405c..eb4a043 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -109,7 +109,7 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-08-18 — **gen seat SWAPPED to `absolute-heresy` and operator-confirmed in real use.** Also this session: irv-ml1 cleared of 782 GB of dead weights, Homepage brought under version control, and `esh-pve-nas` diagnosed as running PVE root off a USB DOM — mitigated tonight (90% → 76%), migration planned and about to be staged. No blocking work; the esh-pve-nas reboot window is the next scheduled thing._ +_As of 2026-08-18 — **gen seat SWAPPED to `absolute-heresy` and operator-confirmed in real use**, and **`esh-pve-nas`'s migration off the USB DOM is fully STAGED — only the reboot window remains.** Also this session: irv-ml1 cleared of 782 GB of dead weights, Homepage brought under version control. The esh-pve-nas cutover is the next scheduled thing; nothing else is blocking._ - **🟢 GEN SEAT — SWAPPED to `absolute-heresy` 2026-08-17 (validated, promoted).** Live gen = `/tank/aimodels/qwen38-27b-heresy-nvfp4-mixed` — **MuXodious/Qwen3.8-27B-absolute-heresy** (Heretic v1.4.0 + **SOMPOA**, trial T377, pin `c2374593`) put through our own mixed NVFP4+FP8 recipe. Chosen because it beats the incumbent on **both** axes at once: author refusals 2/101 vs 12/100, first-token KL 0.0759 vs 0.1191. **Gate (probe :8017, pinned nightly, seat-matched flags): MTP 47.2% (inc. 48.2%), decode 103.5 tok/s (96.4), prefill 6618/5403 @6.7k/27k (6334/5085), PPL 6.910 (7.059 — 2.1% BETTER), surface 6/6, abliteration 4/4, and 0/55 refusals on our battery-instruct arm with ZERO EMPTY (no catatonia).** ⚠ speed deltas are **image-confounded** (probe on the pinned nightly, incumbent numbers from an earlier image) — read as "not worse", not a clean win. All 7 LiteLLM aliases verified end-to-end; GPU0 at 91.3/97.9 GB with meromero healthy (more headroom than the old build's 96.8). ⚠ **RC1, 2 days old, ~348 downloads.** **Operator-confirmed "working very well" in real use 2026-08-17**, same evening as the cutover — the signal the synthetic gates structurally cannot give (multi-turn degeneration is stochastic; four synthetic tests once validated three non-fixes). Not yet the 60k-token bar the prior seat cleared, so **keep watching and do NOT delete the rollback weights yet**. **ROLLBACK:** `sudo cp /opt/docker/compose/gen-seat/.env.bak-heresy-20260817 /opt/docker/compose/gen-seat/.env && cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen`; incumbent weights UNTOUCHED at `qwen38-27b-uncensored-nvfp4-mixed` — **do NOT delete** until this holds. Runbook `services/gen-seat-mixed-quant/RUNBOOK-heresy-swap.md`. @@ -133,7 +133,11 @@ _As of 2026-08-18 — **gen seat SWAPPED to `absolute-heresy` and operator-confi - **⏳ WATCHING: DavidAU's HERETIC build of Qwen3.8-27B — the one worth waiting for.** `DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1` examined 2026-08-17 and **NOT adopted**: it is a capability/efficiency finetune of **stock** Qwen3.8 and every bench row on its card is labelled **`[non heretic]`** — adopting it would reintroduce base refusals the current seat does not have. ⚠ **Easy to misread as uncensored** (operator did, and it is a fair mistake): DavidAU's back catalog is almost all `Uncensored-Heretic` builds — Fable-Fusion 711, Qwen3.5-9B Cold-Fusion — so the naming pattern implies it. This one simply has not had that stage run yet; the card's roadmap says the HERETIC version is **IN PROGRESS from base**. **That** is the release to watch, not this one. What makes it worth watching: third-party benches (Nightmedia, mxfp8) beat stock Qwen3.8 by **+0.064 arc/c, +0.056 arc/e, +0.056 obkqa**; claimed **MTP acceptance 55.7% (record 59.9%)** vs our measured 47.2%; **thinking tokens cut to 1/10–1/2**; PPL *dropped* vs base. Same GAIN/Cold-Fusion pipeline that produced Fable-Fusion 711, which we already serve on `char-rp-reasoning` — proven in-fleet, not just claimed. Structurally clean (1199 tensors, 15 mtp in shard 18, 333 visual). ⚠ MTP/speed figures are **GGUF/llama.cpp on a 5090, not vLLM** — may not transfer; and because its MTP head was likely *trained*, the free CPU-hash shortcut would NOT apply (it won't match base) so a real acceptance gate would be needed. -- **🟡 esh-pve-nas — PVE root on a USB DOM; MITIGATED, migration STAGING NEXT.** Root was 90% (571 MB free) on a 6 GB ext4 root on `sdq`, a **NORELSYS USB Disk-on-Module**. Wear is NOT the driver (a DOM is SLC/pSLC — operator corrected my first read); the drivers are the **USB bus** (a reset drops root under a running hypervisor), no headroom, no mirror, and **blocked patching: 225 packages pending, 161 with `deb12uN`/security bumps**, stuck on PVE 8.4.11 vs esh-pve's 8.4.14. **Mitigated 2026-08-17 → 76% / 1.4 GB free** (journald capped, journal moved to ZFS `nvme/varlog`, apt clean, `/root/neo` stashed). **PLAN: split boot from root** — `/boot`+ESP stay ext4 on the DOM (so GRUB never reads ZFS; the pool has `encryption`/`large_dnode`/`zstd_compress`), root moves to `nvme/ROOT/pve-1`. One reboot, rollback is a GRUB entry, `nvme` pool survives. ⚠ **Migrate FIRST, patch after** — a signed kernel would land in `/boot` on the 1.3 GB root. ⚠ **CT 103 `esh-nas` (10.0.50.50) runs on this host and serves `hard` NFS to esh-docker-vm and esh-pve — quiesce both before any reboot** or you wedge esh-docker-vm into D-state. Runbook `docs/runbooks/esh-pve-nas-boot-migration.md`; config snapshot off-box at `nh3-dev:~/backups/esh-pve-nas/`. Detail → `persistent-memory.d/2026-08-17-esh-pve-nas-dom.md`. +- **🟢 esh-pve-nas — migration STAGED 2026-08-18; ONLY THE REBOOT WINDOW REMAINS.** PVE root ran off a **NORELSYS USB Disk-on-Module** (`sdq`). Wear was NOT the driver (a DOM is SLC/pSLC — operator corrected my first read); the drivers are the **USB bus** (a reset drops root under a running hypervisor), no headroom, no mirror, and **blocked patching: 225 packages, 161 with `deb12uN`/security bumps**, stuck on PVE 8.4.11 vs esh-pve's 8.4.14. Mitigated 2026-08-17 (90% → 76%); **staged 2026-08-18 via two rerunnable elway playbooks** (`esh-pve-nas-stage-zfs-root.yaml`, `esh-pve-nas-stage-bootloader.yaml`), 0 failed, 17/17 verify green. + **Host is still on ext4 and its boot path is byte-identical to the last 140 days** — `grub-install` is deliberately held back to the window, so an unplanned reboot mid-staging is a non-event. **CUTOVER = quiesce NFS clients → stop guests → `grub-install` → `grub-reboot pve-zfs-root` → `zfs set mountpoint=/ nvme/ROOT/pve-1` → reboot.** + ⚠⚠ **THE POOL-NAME BUG — the near-miss worth remembering.** `update-grub` silently emits `root=ZFS=/ROOT/pve-1` with **the pool name missing** → boots to an initramfs prompt. Cause: Debian's `10_linux` builds `${rpool}${bootfs}` where `rpool = grub-probe --target=fs_label`, and **GRUB's ZFS reader cannot open a pool with `encryption`/`large_dnode`/`zstd_compress`** — so it returns empty and the failure is swallowed by `2>/dev/null || true`. The *same* feature set that forced `/boot` to stay ext4 also corrupts the kernel cmdline, which the design did not anticipate. Fixed with a `/etc/default/grub.d/zfs-root.cfg` drop-in (last `root=` wins) plus explicit `pve-zfs-root` / `pve-ext4-rollback` entries with stable ids. **Durable lesson: assert the EFFECTIVE value, not the presence of a substring** — the original verify grepped for the right string *appearing somewhere* and would have passed with every menu entry still broken. + ⚠ **Never set the ZFS cachefile on one pool.** The runbook's `zpool set cachefile=… nvme` was a trap: populating a cache flips the host from import-by-scan to import-by-cache, so a cache holding only `nvme` leaves `ssd`+`tank` unimported and empties every CT 103 export. Set on all three 2026-08-18, verified in the 11,976-byte cache. + ⚠ **Migrate FIRST, patch after** — a signed kernel would land in `/boot` on the 1.3 GB root. ⚠ **CT 103 `esh-nas` (10.0.50.50) runs on this host and serves `hard` NFS to esh-docker-vm and esh-pve — quiesce both before any reboot** or you wedge esh-docker-vm into D-state. Off-box at `nh3-dev:~/backups/esh-pve-nas/`: DOM image `dom-sdq-20260818.img.zst` (2.38 GiB, crash-consistent), clean `bootchain-20260818.tar.gz`, config snapshot `…20260818T051*.tar.gz`. Runbook `docs/runbooks/esh-pve-nas-boot-migration.md`; detail → `persistent-memory.d/2026-08-17-esh-pve-nas-dom.md`. - **OPEN FOLLOW-UPS (parked):** move gen seat off pinned-nightly to stable once #51113 ships; Lobe one-time TTS UI pass; delete the 1.8GB litellm dump; `harden-esh-docker-vm` (park id 28, PROMOTED — Tier-1 done, `/mnt/books` stays hard w/ watchdog); chatterbox-fast build-context divergence; #363 research-wing ingest (no deadline); optionally attach our MTP reproducer to vllm#47087 (needs a GitHub identity — operator's call). @@ -143,6 +147,8 @@ _As of 2026-08-18 — **gen seat SWAPPED to `absolute-heresy` and operator-confi ## Recent decisions +- `[2026-08-18]` **esh-pve-nas migration STAGED — and staging is where three landmines surfaced, none of which the plan predicted.** (1) The runbook's `/boot` LV had **nowhere to live**: VG `pve` had 4 MB free and mounted ext4 cannot shrink, so the space came from the 768 MB swap LV (operator's call: shrink to 256 MB, not drop). (2) The runbook's `zpool set cachefile=… nvme` would have **broken the NAS** — populating a cache flips the host to import-by-cache, and a one-pool cache leaves `ssd`+`tank` unimported under CT 103's twelve bind mounts. (3) **`update-grub` silently emitted a pool-less `root=ZFS=/ROOT/pve-1`**, because GRUB's ZFS reader cannot open a pool with `encryption`/`large_dnode`/`zstd_compress` and the probe failure is swallowed. All three were caught by *verify steps that asserted effective state*, not by reading the plan. → `persistent-memory.d/2026-08-17-esh-pve-nas-dom.md` + - `[2026-08-17]` **esh-pve-nas PVE root is on a USB DOM — mitigated, and the migration replanned to split boot from root.** Operator's design beats my reinstall plan; wear was never the issue, blocked patching is. → `persistent-memory.d/2026-08-17-esh-pve-nas-dom.md` - `[2026-08-17]` **irv-ml1 cleared of 782 GB, and Homepage brought under version control.** One dead-looking Gradio app pinned three delete targets at once; `/opt/ComfyUI` is NOT the ComfyUI that serves. → `persistent-memory.d/2026-08-17-irv-ml1-cleanup-homepage.md` diff --git a/playbooks/esh-pve-nas-stage-bootloader.yaml b/playbooks/esh-pve-nas-stage-bootloader.yaml new file mode 100644 index 0000000..151265b --- /dev/null +++ b/playbooks/esh-pve-nas-stage-bootloader.yaml @@ -0,0 +1,271 @@ +# esh-pve-nas — PHASE 2 of the ZFS-root migration: build the boot artifacts. +# +# Runbook: docs/runbooks/esh-pve-nas-boot-migration.md +# Run AFTER playbooks/esh-pve-nas-stage-zfs-root.yaml. +# +# ⚠ THIS PLAYBOOK DELIBERATELY DOES NOT RUN `grub-install`. +# +# That is the whole safety design. Everything expensive and error-prone — the +# ZFS-capable initramfs, the generated grub.cfg, the rollback menu entry, the +# grubenv default — is built and verified here, onto the NEW /boot LV, while the +# ESP stub on the DOM still points at the OLD /boot inside the ext4 root LV. +# +# So until cutover the host's boot path is byte-for-byte what it has been for +# 140 days. An unplanned reboot mid-staging lands exactly where it always did. +# The cutover reduces to one idempotent two-second command plus the reboot: +# +# chroot /mnt/newroot grub-install --target=x86_64-efi \ +# --efi-directory=/boot/efi --bootloader-id=proxmox +# chroot /mnt/newroot grub-reboot '' +# reboot +# +# Why the rollback entry matters here: the ext4 root LV keeps its own /boot +# contents (the new LV is a copy, not a move), and its initrd is never +# regenerated — update-initramfs inside the chroot writes only to the new LV. +# So the rollback path is genuinely independent of anything we build. +# +# Why GRUB_DEFAULT=saved: the default stays pinned to the ext4 rollback entry. +# At cutover `grub-reboot` marks the ZFS entry to be tried EXACTLY ONCE. If the +# ZFS root fails to come up, the next reboot returns to ext4 with nobody at the +# console — which matters because a failed boot here takes CT 103 `esh-nas` +# down and wedges esh-docker-vm into unkillable D-state on hard NFS. + +vars: + newroot: /mnt/newroot + root_dataset: nvme/ROOT/pve-1 + +steps: + # ---------- guards ---------- + + - name: GUARD — host must still be running from the ext4 root on the DOM + shell: | + test "$(findmnt -no FSTYPE /)" = "ext4" || { + echo "root is not ext4 — already cut over; refusing"; exit 1; } + changed_when: "false" + + - name: GUARD — phase 1 must have completed (ZFS copy populated) + shell: | + mountpoint -q {{ newroot }} || { echo "{{ newroot }} not mounted"; exit 1; } + test -x {{ newroot }}/usr/bin/pveversion || { echo "ZFS copy incomplete"; exit 1; } + test -f {{ newroot }}/etc/fstab || { echo "ZFS copy has no fstab"; exit 1; } + changed_when: "false" + + # Tolerates either staging location: /mnt/boot-new before this playbook has + # moved the LV, {{ newroot }}/boot after — so a rerun still passes. + - name: GUARD — the new /boot LV must exist and carry a kernel + shell: | + lvs pve/boot >/dev/null 2>&1 || { echo "pve/boot missing"; exit 1; } + ls /mnt/boot-new/vmlinuz-* >/dev/null 2>&1 || \ + ls {{ newroot }}/boot/vmlinuz-* >/dev/null 2>&1 || { + echo "no kernel on the boot LV at either staging path"; exit 1; } + changed_when: "false" + + # ---------- back up what we are about to regenerate ---------- + + - name: Snapshot the ESP and grub defaults before touching anything + shell: | + mkdir -p /root/pre-zfs-boot-backup + tar czf /root/pre-zfs-boot-backup/esp-and-grub.tar.gz \ + -C / boot/efi etc/default/grub 2>/dev/null + ls -la /root/pre-zfs-boot-backup/ + creates: /root/pre-zfs-boot-backup/esp-and-grub.tar.gz + + # ---------- assemble the chroot ---------- + + - name: Release the staging mount of the boot LV so it can move under the chroot + shell: umount /mnt/boot-new + when: "mountpoint -q /mnt/boot-new" + + # ESP goes in as a BIND of the live /boot/efi rather than a second mount of + # /dev/sdq2 — same filesystem either way, but the bind leaves no ambiguity + # about which superblock grub-install writes through at cutover. + - name: Mount the boot LV and ESP inside the ZFS copy + shell: | + mount /dev/pve/boot {{ newroot }}/boot + mkdir -p {{ newroot }}/boot/efi + mount --bind /boot/efi {{ newroot }}/boot/efi + when: "! mountpoint -q {{ newroot }}/boot" + + - name: Bind the kernel filesystems into the chroot + shell: | + for d in dev dev/pts proc sys; do + mountpoint -q {{ newroot }}/$d || mount --rbind /$d {{ newroot }}/$d + done + changed_when: "true" + + # ---------- build the boot artifacts inside the chroot ---------- + + - name: Pin the default boot entry to the rollback, not to ZFS + shell: | + sed -i -e 's/^GRUB_DEFAULT=.*/GRUB_DEFAULT=saved/' \ + -e 's/^#\?GRUB_SAVEDEFAULT=.*/GRUB_SAVEDEFAULT=false/' \ + {{ newroot }}/etc/default/grub + grep -q '^GRUB_DEFAULT=saved' {{ newroot }}/etc/default/grub + grep -q '^GRUB_TIMEOUT=' {{ newroot }}/etc/default/grub || \ + echo 'GRUB_TIMEOUT=5' >> {{ newroot }}/etc/default/grub + changed_when: "true" + + # ⚠ THE POOL-NAME BUG. Left to itself, grub-mkconfig emits + # root=ZFS=/ROOT/pve-1 + # with the pool name MISSING, which drops the boot at an initramfs prompt. + # + # Cause, and it is worth understanding because it is not a typo: Debian's + # /etc/grub.d/10_linux builds the ZFS root as ${rpool}${bootfs}, where + # rpool = grub-probe --device --target=fs_label + # bootfs = make_system_path_relative_to_its_root / -> /ROOT/pve-1 + # and `grub-probe --target=fs /` on this pool fails outright with "unknown + # filesystem" — GRUB's own ZFS reader cannot open a pool with `encryption`, + # `large_dnode` and `zstd_compress` enabled. So rpool comes back EMPTY and + # concatenates to nothing. It is the very same feature set that forced /boot + # to stay ext4; here it silently corrupts the kernel command line instead of + # erroring, which is why this is caught by a verify step and not by trust. + # + # A drop-in is used rather than editing /etc/default/grub so a future grub + # package upgrade cannot revert it in a conffile merge. + - name: Override the ZFS root on the kernel command line (grub cannot derive it) + shell: | + mkdir -p {{ newroot }}/etc/default/grub.d + cat > {{ newroot }}/etc/default/grub.d/zfs-root.cfg <<'EOF' + # grub-mkconfig cannot resolve this pool's name (GRUB's ZFS reader does not + # support encryption/large_dnode/zstd_compress) and emits a pool-less + # root=ZFS=/ROOT/pve-1. This appends the correct value AFTER it; the kernel + # and the zfs initramfs script both take the LAST root= on the line. + # The explicit `pve-zfs-root` menu entry in 40_custom carries a single + # clean root= and is what cutover targets — this drop-in exists so the + # auto-generated entries are correct too. + GRUB_CMDLINE_LINUX="root=ZFS=nvme/ROOT/pve-1 boot=zfs" + EOF + changed_when: "true" + + # Both entries are hand-authored with STABLE ids. The auto-generated ones get + # ids derived from device paths (`gnulinux-simple-/dev/nvme0n1p1_/dev/nvme1n1p1`) + # which change if the pool's members ever change — not something to aim + # `grub-reboot` at during a downtime window. + - name: Author the explicit ZFS-root and ext4-rollback menu entries + shell: | + ROOT_UUID=$(blkid -s UUID -o value /dev/mapper/pve-root) + BOOT_UUID=$(blkid -s UUID -o value /dev/mapper/pve-boot) + KVER=$(basename $(ls -1 {{ newroot }}/boot/vmlinuz-* | sort -V | tail -1) | sed 's/^vmlinuz-//') + test -n "$ROOT_UUID" && test -n "$BOOT_UUID" && test -n "$KVER" + cat > {{ newroot }}/etc/grub.d/40_custom </dev/null + changed_when: "false" + + - name: ESP is still the ORIGINAL stub pointing at the ext4 root (no grub-install yet) + shell: | + grep -q "$(blkid -s UUID -o value /dev/mapper/pve-root)" \ + {{ newroot }}/boot/efi/EFI/proxmox/grub.cfg + changed_when: "false" + + - name: Show the cutover command and every entry's effective root + shell: | + echo "--- cutover one-shot: chroot {{ newroot }} grub-reboot pve-zfs-root ---" + echo "--- effective root= per menu entry ---" + awk '/^[[:space:]]*menuentry/ { t = $0; sub(/^[[:space:]]*menuentry[[:space:]]*/, "", t) } + /^[[:space:]]*linux[[:space:]]/ { + r = ""; + for (i = 1; i <= NF; i++) if ($i ~ /^root=/) r = $i; + printf " %-46.46s -> %s\n", substr(t, 1, 46), r + }' {{ newroot }}/boot/grub/grub.cfg + changed_when: "false" diff --git a/playbooks/esh-pve-nas-stage-zfs-root.yaml b/playbooks/esh-pve-nas-stage-zfs-root.yaml new file mode 100644 index 0000000..f71f3cd --- /dev/null +++ b/playbooks/esh-pve-nas-stage-zfs-root.yaml @@ -0,0 +1,178 @@ +# esh-pve-nas — STAGE the PVE root migration off the USB DOM onto ZFS. +# +# Runbook: docs/runbooks/esh-pve-nas-boot-migration.md +# Design: boot chain stays ext4 on the DOM; root moves to nvme/ROOT/pve-1. +# +# THIS PLAYBOOK DOES NOT CUT OVER. It leaves the host still running from the +# ext4 root on the DOM. Nothing here changes what the next reboot does — the +# bootloader phase is deliberately a separate playbook. +# +# What it does, all live, no downtime: +# 1. Reclaims 512 MB from the 768 MB swap LV for a dedicated /boot LV +# (operator's call 2026-08-17: shrink swap to 256 MB rather than drop it). +# 2. Populates that LV from the current /boot. +# 3. rsyncs the live ext4 root into the ZFS dataset nvme/ROOT/pve-1. +# 4. Writes the ZFS copy's /etc/fstab for the post-cutover layout. +# +# The ext4 root LV is never modified — it stays byte-intact as the rollback, +# including its own /boot contents, which the new mount only shadows. +# +# Preconditions (verified 2026-08-17, re-asserted as guard steps below): +# - nvme/ROOT/pve-1 exists, canmount=noauto, encryption off +# - /etc/zfs/zpool.cache populated with ALL THREE pools (nvme, ssd, tank). +# ⚠ A cache holding only `nvme` flips the host from import-by-scan to +# import-by-cache and leaves ssd+tank unimported at boot — which breaks +# CT 103 `esh-nas`, whose 12 bind mounts span all three pools. +# +# Rerunnable: every step is guarded, so a second run reports ok/skipped. + +vars: + newroot: /mnt/newroot + bootstage: /mnt/boot-new + boot_lv_size: 512M + swap_lv_size: 256M + root_dataset: nvme/ROOT/pve-1 + +steps: + # ---------- guards: refuse to run against an already-migrated or unprepared host ---------- + + - name: GUARD — host must still be running from the ext4 root on the DOM + shell: | + test "$(findmnt -no FSTYPE /)" = "ext4" || { + echo "root is not ext4 — host already cut over; refusing"; exit 1; } + changed_when: "false" + + - name: GUARD — ZFS root dataset must exist with canmount=noauto + shell: | + test "$(zfs get -H -o value canmount {{ root_dataset }})" = "noauto" || { + echo "{{ root_dataset }} missing or canmount!=noauto; refusing"; exit 1; } + changed_when: "false" + + - name: GUARD — zpool.cache must list all three pools + shell: | + for p in nvme ssd tank; do + zdb -C -U /etc/zfs/zpool.cache 2>/dev/null | grep -q "name: '$p'" || { + echo "pool $p missing from zpool.cache — would not import at boot"; exit 1; } + done + changed_when: "false" + + # ---------- phase 1: carve a /boot LV out of swap ---------- + + # Gated on the ORIGINAL 768M size, not on "is swap on" — otherwise a rerun + # swaps off the new 256M device and never turns it back on. + - name: Disable swap so its LV can be resized + shell: swapoff /dev/pve/swap + when: "lvs --noheadings -o lv_size --units m pve/swap 2>/dev/null | grep -q '768'" + + - name: Remove the oversized swap LV + shell: lvremove -y pve/swap + when: "lvs --noheadings -o lv_size --units m pve/swap 2>/dev/null | grep -q '768'" + + - name: Create the dedicated /boot LV + shell: lvcreate -y -L {{ boot_lv_size }} -n boot pve + when: "! lvs pve/boot >/dev/null 2>&1" + + - name: Recreate swap at the reduced size + shell: lvcreate -y -L {{ swap_lv_size }} -n swap pve + when: "! lvs pve/swap >/dev/null 2>&1" + + - name: Make the /boot filesystem + shell: mkfs.ext4 -q -L pveboot /dev/pve/boot + when: "! blkid -s TYPE -o value /dev/pve/boot 2>/dev/null | grep -q ext4" + + - name: Make and enable the new swap + shell: | + blkid -s TYPE -o value /dev/pve/swap 2>/dev/null | grep -q swap || mkswap -L pveswap /dev/pve/swap + swapon /dev/pve/swap + when: "! swapon --show=NAME --noheadings | grep -q dm-" + + # ---------- phase 2: populate the /boot LV ---------- + + - name: Stage-mount the new /boot LV + shell: mkdir -p {{ bootstage }} && mount /dev/pve/boot {{ bootstage }} + when: "! mountpoint -q {{ bootstage }}" + + - name: Copy the current /boot into it (ESP contents excluded — separate vfat mount) + shell: | + rsync -aHAX --numeric-ids --one-file-system --delete \ + --exclude='/lost+found' \ + /boot/ {{ bootstage }}/ + mkdir -p {{ bootstage }}/efi + changed_when: "true" + + - name: Verify the kernel and initrd landed + shell: | + ls {{ bootstage }}/vmlinuz-* {{ bootstage }}/initrd.img-* >/dev/null + test -f {{ bootstage }}/grub/grub.cfg + changed_when: "false" + + # ---------- phase 3: rsync the live root into the ZFS dataset ---------- + + - name: Point the ZFS root dataset at a staging mountpoint + shell: zfs set mountpoint={{ newroot }} {{ root_dataset }} + when: "test \"$(zfs get -H -o value mountpoint {{ root_dataset }})\" != '{{ newroot }}'" + + - name: Mount the ZFS root dataset for staging + shell: zfs mount {{ root_dataset }} + when: "! mountpoint -q {{ newroot }}" + + - name: rsync the ext4 root into ZFS (one-file-system — every other mount is excluded) + shell: | + rsync -aHAX --numeric-ids --one-file-system --delete \ + --exclude='/proc/*' --exclude='/sys/*' --exclude='/dev/*' \ + --exclude='/run/*' --exclude='/tmp/*' --exclude='/mnt/*' \ + --exclude='/media/*' \ + / {{ newroot }}/ + # mountpoints that --one-file-system skipped still need to exist + mkdir -p {{ newroot }}/proc {{ newroot }}/sys {{ newroot }}/dev \ + {{ newroot }}/run {{ newroot }}/tmp {{ newroot }}/mnt \ + {{ newroot }}/boot {{ newroot }}/boot/efi \ + {{ newroot }}/nvme {{ newroot }}/ssd {{ newroot }}/tank \ + {{ newroot }}/var/log/journal + chmod 1777 {{ newroot }}/tmp + changed_when: "true" + + # ---------- phase 4: fstab for the post-cutover layout ---------- + + - name: Write the ZFS copy's /etc/fstab + shell: | + cat > {{ newroot }}/etc/fstab <<'FSTAB' + # + # root is {{ root_dataset }} (ZFS) — mounted by the initramfs, no entry here. + /dev/pve/boot /boot ext4 defaults 0 2 + UUID=1D32-43A5 /boot/efi vfat defaults 0 2 + /dev/pve/swap none swap sw 0 0 + proc /proc proc defaults 0 0 + FSTAB + changed_when: "true" + +verify: + - name: LVM layout is root + boot + swap + shell: lvs --noheadings -o lv_name pve | tr -d ' ' | sort | tr '\n' ',' | grep -qx 'boot,root,swap,' + changed_when: "false" + + - name: Live root is still the untouched ext4 LV + shell: test "$(findmnt -no SOURCE /)" = "/dev/mapper/pve-root" + changed_when: "false" + + - name: Swap is active at the reduced size + shell: swapon --show=NAME --noheadings | grep -q dm- + changed_when: "false" + + - name: New /boot LV carries a bootable kernel set + shell: ls {{ bootstage }}/vmlinuz-* {{ bootstage }}/initrd.img-* >/dev/null + changed_when: "false" + + - name: ZFS root copy has a populated /usr and /etc + shell: test -x {{ newroot }}/usr/bin/pveversion && test -f {{ newroot }}/etc/fstab + changed_when: "false" + + - name: ZFS root copy's fstab has no root line and does have the boot line + shell: | + ! grep -qE '^\S+\s+/\s+' {{ newroot }}/etc/fstab + grep -q '/dev/pve/boot /boot ext4' {{ newroot }}/etc/fstab + changed_when: "false" + + - name: PVE cluster config copied (guest configs present) + shell: test -d {{ newroot }}/var/lib/pve-cluster + changed_when: "false" diff --git a/servers/esh-pve-nas/README.md b/servers/esh-pve-nas/README.md index 3f97bcf..052f628 100644 --- a/servers/esh-pve-nas/README.md +++ b/servers/esh-pve-nas/README.md @@ -17,11 +17,27 @@ Media services on this box sit at `10.0.50.56` (Plex) and `10.0.50.57` (Jellyfin - **RAM:** 125.6 GB - **Kernel:** `6.8.12-13-pve` (Proxmox 8.x) - **Storage:** ⚠ **PVE root is on a USB Disk-on-Module** (`sdq`, 7.3 GB, `ID_BUS=usb`, - NORELSYS 1081) — 6 GB ext4 root + 768 MB swap + a 512 MB ESP that is the host's - **only** boot path. Not a wear problem (a DOM is SLC/pSLC), but it is on the USB - bus, has no headroom, and is unmirrored. Mitigated 2026-08-17 (90% → 76%: journald - capped, journal relocated to ZFS); migration planned in - [`docs/runbooks/esh-pve-nas-boot-migration.md`](../../docs/runbooks/esh-pve-nas-boot-migration.md). + NORELSYS 1081), whose 512 MB ESP is the host's **only** boot path. Not a wear + problem (a DOM is SLC/pSLC), but it is on the USB bus, has no headroom, and is + unmirrored. Mitigated 2026-08-17 (90% → 76%: journald capped, journal relocated + to ZFS). **Migration STAGED 2026-08-18 — awaiting only the reboot window**; + runbook [`docs/runbooks/esh-pve-nas-boot-migration.md`](../../docs/runbooks/esh-pve-nas-boot-migration.md). + + DOM LVM layout after staging (was root 6.04 G + swap 768 M): + + | LV | size | role | + |---|---|---| + | `pve-root` | 6.04 G | ext4 — the **live root today**, and the rollback after cutover | + | `pve-boot` | 512 M | ext4 — the new `/boot`, carved out of swap | + | `pve-swap` | 256 M | swap, shrunk to make room | + + Post-cutover root is `nvme/ROOT/pve-1` on the mirrored NVMe pool. `/boot` stays + ext4 on the DOM on purpose: GRUB cannot read the `nvme` pool, which has + `encryption`, `large_dnode` and `zstd_compress` enabled. +- ⚠ **Never set the ZFS cachefile on one pool.** `zpool set cachefile=…` flips the + host from import-by-scan to import-by-cache; a cache holding only `nvme` leaves + `ssd` and `tank` unimported at boot, which empties every CT 103 export. Set it on + all three or none. - **Pools:** `nvme` (2× 931 GB NVMe mirror — 32 G used, 867 G free, holds every guest rootfs), `ssd` (4× 894 GB Intel SATA, 2 mirrors — 1.42 T free), `tank` (12× 14.6 TB raidz2 ×2 — 40 T of 175 T). Plus NFS `/mnt/pve/tank-vmbu` for VM backups. @@ -31,7 +47,12 @@ Media services on this box sit at `10.0.50.56` (Plex) and `10.0.50.57` (Jellyfin wedges esh-docker-vm into unkillable D-state unless those clients are quiesced first. Never reboot this host casually. -> **Root-fs pressure:** at 87% used on a 5.9 GB root partition, there's not much room for package upgrades or logs. Worth cleaning up or growing the root if this host is staying in production for a while. +> **Root-fs pressure:** 77% of a 5.9 GB root (1.3 GB free) after the 2026-08-17 +> mitigation. Still not enough for the pending upgrade — **225 packages, 161 +> carrying `deb12uN`/security bumps, including a ~250 MB signed kernel that lands +> in `/boot`.** ⚠ **Migrate first, patch after:** unpacking that into 1.3 GB of +> headroom risks wedging dpkg on a hypervisor running five guests. The host is on +> `pve-manager/8.4.11` vs esh-pve's 8.4.14 for exactly this reason. ## What it runs