docs: refresh what today's work made stale — booth asks (inline placement promoted to its own section), ana-ml2 nvme7 settled by the scrub result, nh3-dev booth entry + the CLI-on-PATH fix, run-07 runbook outcome + serving state

This commit is contained in:
vh
2026-09-09 14:18:34 -07:00
parent 78c3a7c170
commit 6e0b85ba27
5 changed files with 137 additions and 29 deletions
+13 -2
View File
@@ -44,8 +44,19 @@ at import. While it was missing `tank` was DEGRADED, and Debian's `zfsutils-linu
(`/usr/lib/zfs-linux/{scrub,trim}`) only touches pools whose health is `ONLINE`, so tank
got **no scrub and no trim from 04-12 to 09-06**. ZED's `ZED_EMAIL_ADDR=root` has no
MTA behind it, so the 4½-month degradation alerted nobody. `media_errors=2084` on
nvme7 is a lifetime counter; the 2026-09-09 scrub is the first fresh measurement
(baseline 2084 at 00:32 PT — compare after any future event, growth = replace).
nvme7 is a lifetime counter.
**Settled by the 2026-09-09 scrub** (00:29–02:02 PT, `scrub repaired 0B in 01:32:44
with 0 errors`, then `zpool clear tank` → CKSUM 2 → 0): `media_errors` read **2084
before and 2084 after** a full 6.84 TiB verify, so the counter is prior-life
history, not an active fault, and the 2 CKSUM were the stale-block artefact of the
09-05 late resilver. **nvme7 stays in service; watch the counter at every visit and
replace on growth** (`zpool replace tank nvme7n1 <new>`; any PM1725b 1.6 TB or
larger). Slot 0-5 itself deserves a reseat / cable check at the next hands-on
visit — a bay that dropped a drive for 4½ months is the likelier fault than the
drive. Playbook: `playbooks/ana-ml2-pool-health.yaml` (idempotent; rerunning is a
no-op). ⚠ **Nothing alerts on this** — see the open follow-up in
`persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`.
## Key paths