ops(esh-nas-pve): tank verified healthy after two scrubs; disk kept; PVE 9 blocker cleared

This commit is contained in:
vh
2026-10-03 06:28:48 -07:00
parent 6a274a412a
commit 3a6d997c99
2 changed files with 6 additions and 3 deletions
@@ -14,7 +14,7 @@ official wiki (pve.proxmox.com/wiki/Upgrade_from_8_to_9) on the day; anything th
| PVE / kernel | 8.4.20 / 6.8.12-42 (‑43 pending) | 8.4.20 / 6.8.12-13 (up 6 weeks; ‑43 pending) |
| Root | ext4 on LVM `pve/root` (96 G, 69 G free), VG free 16 G | ZFS `nvme/ROOT/pve-1` (pool `nvme`) |
| Boot | UEFI, **GRUB** (proxmox-boot-tool not in use) | UEFI, **GRUB reading ZFS** (proxmox-boot-tool not in use) |
| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2, **DEGRADED**) |
| Extras | NVIDIA 580.178.04 **DKMS** (RTX 2000E Ada → CT 110 esh-ml1) | ZFS pools `nvme`, `ssd`, `tank` (175 T raidz2×2; healthy since 2026-10-03, was DEGRADED) |
| Guests | 100 esh-vm-docker, 101 esh-vm-db, 108 esh-scale, 110 esh-ml1, 111 esh-matter | 103 esh-nas (NFS server 10.0.50.50), 104 vm-esh-nas, 105 plex, 106 filebot, 107 jellyfin |
Repos: `pve-no-subscription` + Debian bookworm main/contrib, no Ceph, no HA resources.
@@ -22,7 +22,10 @@ Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`.
## ⚠ Findings that come before the upgrade
1. **`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk
1. ✅ **RESOLVED 2026-10-03 0629, so no longer a blocker.** Two full scrubs ran. The first repaired 198 G, the
residue of the Aug 20 resilver, which had never been followed by a full scrub. The second repaired 0 B with 0
errors. The disk was kept, and `zpool status -x` says healthy. The original finding follows.
**`tank` on esh-nas-pve is DEGRADED, and has been since about 2026-08-20.** In raidz2-0, disk
`wwn-0x5000c500c91df554` shows 36 CKSUM errors ("too many errors"). The last resilver (Aug 20,
3.11 T) logged 6,676,313 errors. "No known data errors" and SMART is clean (0 reallocated,
0 pending, 0 CRC, 11,006 h). That vdev is down to one disk of redundancy. **Fix first:**