2.5 KiB
ESH tank (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629
[2026-10-02] ✅ ESH tank (esh-nas-pve, 175 T raidz2×2): was DEGRADED since ~2026-08-20 (found while planning PVE 9); HEALTHY since 2026-10-03 0629. raidz2-0 disk wwn-0x5000c500c91df554: 36 CKSUM errors; Aug 20 resilver logged 6.68 M errors; no known data errors; SMART clean. Prime (via Miranda) 1925: verify, then clear + full scrub. Verified residual: the disk was REPLACED 2026-08-20 09:50 (zpool history), resilver ended with 6.68 M errors, and since the last clear only 21 checksum events (Aug 21 02:05 ×20, Aug 29 02:50 ×1), none in 34 days. zpool clear 1926 → ONLINE; full scrub running (background watcher); replace the disk if errors return. UPDATE 2331: the scrub is REPAIRING that disk: 6.19 M CKSUM, 198 G repaired at 72%, 0 read/write errors, still no known data errors, ETA ~0107. History explains it: after the Aug 20 replace, the resilver ended errors=6676436, someone ran only an error-scrub (zpool scrub -e, 4 s) and then zpool clear 19:32, so NO full scrub ran until today, and those blocks sat bad on the new disk for 6 weeks. SMART still clean (0 realloc/pending/uncorrect/CRC, no errors logged), no kernel I/O errors. Read: residual repair, not a failing disk (inferred). Discriminator: clear + a SECOND full scrub; any CKSUM on it → replace. ⚠ The watcher's error check was blind (could not parse 6.19M as an integer); it still fires on completion. 0106 scrub DONE: repaired 198G, 0 errors, no known data errors; that disk 6,494,530 CKSUM exact (97% of the 6,676,436 the Aug 20 resilver failed on), last checksum ereport 21:26:27 = none in the scrub's final 3h40m. 0118: SECOND full scrub started WITHOUT zpool clear, so any new error shows as count > 6494530 (zpool status -p); watcher tested with +1/null/finished controls; ETA ~0700. RESOLVED 0627: second scrub repaired 0B with 0 errors; the disk's CKSUM stayed at 6494530 (zero new); SMART clean; no kernel I/O errors. Verdict: residue of the Aug 20 resilver, now repaired; DISK KEPT. zpool clear 0629 → pool 'tank' is healthy. Reported to Miranda once. PVE 8→9 plan for esh-pve-cluster written, NOT executed (Prime via Miranda): docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md. Also found: remote access to ESH rides CT 108 on pve (no remote recovery if it fails to boot); both nodes fail pve8to9 on the systemd-boot meta-package.