1.7 KiB
1.7 KiB
ana-ml2 pool health — three actions deferred to a clean-context session (2026-09-09)
Operator ruling 2026-09-09 ~00:30 PT: "snapshot and we'll do all 3 on clean context." Findings commit 3e18a04.
Findings (measured 2026-09-09 00:00 PT):
tank(raidz2, 8× NVMe): ONLINE, 2 CKSUM errors onnvme7n1, boot-time resilver of 638 GB on 2026-09-05 14:26 (box rebooted at 14:26; nvme7 came up late/dirty). No data errors, 58% full. No scrub since 2026-04-12 — the Debianzfsutils-linuxsecond-Sunday cron scrubbedzrooton 08-09 but nottank; cause unknown (zpool history tankshows trims monthly, last scrub 04-12).- No
nvme-cliorsmartctlon the box → nvme7's media-error counter unread. zrootat 91% (345 G of 379 G):docker system df= images 429 GB (204 GB reclaimable), build cache 74 GB (36 GB reclaimable).- pfi-pve
NASPool7% /ospool19%, scrubbed 09-05 / 08-09, clean.
The three actions, in order:
sudo zpool scrub tankon ana-ml2 (1–2 h of extra I/O; seats keep serving) → on a clean passsudo zpool clear tank; if the scrub finds errors on nvme7n1 → replace path.sudo apt install nvme-cli→sudo nvme smart-log /dev/nvme7(media_errors, critical_warning, percentage_used) andnvme id-ctrlfor model/serial; record in the drive inventory.docker image prune -a? NO —docker image prune(dangling only) +docker builder pruneon ana-ml2; the 47 unused-but-tagged images need a look first (some are rollback seats: e.g.vllm/vllm-openai:v0.26.0, nightlies). Target: zroot back under ~75%. Also worth a look while there: why the scrub cron skipstank(/usr/lib/zfs-linux/scrublogic — it skips pools with an active trim/resilver or those not "healthy"?).