Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md
T

1.7 KiB
Raw Blame History

ana-ml2 pool health — three actions deferred to a clean-context session (2026-09-09)

Operator ruling 2026-09-09 ~00:30 PT: "snapshot and we'll do all 3 on clean context." Findings commit 3e18a04.

Findings (measured 2026-09-09 00:00 PT):

  • tank (raidz2, 8× NVMe): ONLINE, 2 CKSUM errors on nvme7n1, boot-time resilver of 638 GB on 2026-09-05 14:26 (box rebooted at 14:26; nvme7 came up late/dirty). No data errors, 58% full. No scrub since 2026-04-12 — the Debian zfsutils-linux second-Sunday cron scrubbed zroot on 08-09 but not tank; cause unknown (zpool history tank shows trims monthly, last scrub 04-12).
  • No nvme-cli or smartctl on the box → nvme7's media-error counter unread.
  • zroot at 91% (345 G of 379 G): docker system df = images 429 GB (204 GB reclaimable), build cache 74 GB (36 GB reclaimable).
  • pfi-pve NASPool 7% / ospool 19%, scrubbed 09-05 / 08-09, clean.

The three actions, in order:

  1. sudo zpool scrub tank on ana-ml2 (12 h of extra I/O; seats keep serving) → on a clean pass sudo zpool clear tank; if the scrub finds errors on nvme7n1 → replace path.
  2. sudo apt install nvme-clisudo nvme smart-log /dev/nvme7 (media_errors, critical_warning, percentage_used) and nvme id-ctrl for model/serial; record in the drive inventory.
  3. docker image prune -a? NO — docker image prune (dangling only) + docker builder prune on ana-ml2; the 47 unused-but-tagged images need a look first (some are rollback seats: e.g. vllm/vllm-openai:v0.26.0, nightlies). Target: zroot back under ~75%. Also worth a look while there: why the scrub cron skips tank (/usr/lib/zfs-linux/scrub logic — it skips pools with an active trim/resilver or those not "healthy"?).