ops(ana-ml2): pool-health actions landed — tank scrubbed clean + cleared, nvme-cli SMART inventory, zroot 91→73%; root cause of missed scrubs = nvme7 absent 04-23→09-05 (pool DEGRADED, ZED mail unrouted)

- playbooks/ana-ml2-pool-health.yaml: rerunnable elway play (scrub-if-idle, nvme-cli, dangling-image + builder prune; never prune -a)
- servers/ana-ml2/README.md: 8-drive PM1725b inventory with SMART counters, nvme7 absence + alerting-gap note
- persistent-memory: deferred entry closed, outcome + follow-ups (pool-health alerting, nvme7/slot 0-5 watch, boot import race)
This commit is contained in:
vh
2026-09-09 02:03:11 -07:00
parent 5ad948bf31
commit 5de5583762
4 changed files with 122 additions and 5 deletions
+7 -4
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-09 00:35 PT (fleet-ops: ERP run 7 TRAINING on pfi-gx10 [opening-split slot]; run 6 gated TRANSFERRED after the operator's CSAM adjudication; erp-tune-v6-nvfp4a16 serving on ana-ml2 as `trial`; ESH static-WAN follow-ups + YTVC + webhook repaired; ana-ml2 routes persisted; ana-ml2 tank/zroot actions deferred to the next session)_
_Last updated: 2026-09-09 02:10 PT (fleet-ops: ana-ml2 pool actions LANDED — tank scrubbed clean + cleared, nvme7 SMART read [2084 lifetime media errors, zero growth], zroot 91→73%; ROOT CAUSE of the missed scrubs = nvme7 physically absent 04-23→09-05, pool DEGRADED and ZED mail to nowhere; ERP run 7 TRAINING on pfi-gx10 [~64/542 at 00:42 PT, adapter ~noon]; erp-tune-v6-nvfp4a16 serving as `trial`)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -120,9 +120,11 @@ clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved
--adapter run-07/adapter --out serve/merged-run07 --chat-template <stock>`), copy stock `processor_config.json`
into `merged-run07`, serve `erp-seat-base-ara` for floors → on brokkr's swap cue serve `erp-tune-v7` (run-5 flags,
:8098). Runbook `docs/runbooks/gx10-run-07.md`; run-6 choreography in `gx10-run-06.md`.
- **⏳ ana-ml2 pool actions — NEXT SESSION, operator-approved in principle:** (1) `zpool scrub tank` + `zpool clear`
on a clean pass; (2) `apt install nvme-cli` + read nvme7 SMART; (3) docker image/builder prune to pull zroot back
from 91%. Detail + order: `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`.
- **✅ ana-ml2 pool actions DONE 2026-09-09 02:02 PT** (`playbooks/ana-ml2-pool-health.yaml`): tank scrub clean +
cleared, nvme-cli in, zroot 73%. ⚠ Open follow-ups, operator's call: (a) ZFS pool-health ALERTING — tank sat
DEGRADED 04-23→09-05 with nvme7 physically absent and nobody knew (ZED mails `root`, no MTA); (b) nvme7 / slot 0-5
keep-vs-replace — `media_errors` 2084 lifetime, 0 growth over a full scrub, watch it each visit; (c) boot-time
import race (vdevs UNAVAIL→ONLINE + `no_replicas` every boot). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`.
- **`trial` (LiteLLM) → `erp-tune-v6-nvfp4a16` on ana-ml2 :8021** (stack `stacks/erp-seat`, vLLM nightly
`311b3513`, no gate by operator ruling). Forced tool_choice is prompt-driven (6/9); `response_format: json_schema`
is deterministic. ⚠ `stacks/gemma4-charrp` lacks `--exclude-tools-when-tool-choice-none` (same empty-turn trap);
@@ -144,6 +146,7 @@ clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved
## Recent decisions
- `[2026-09-09]` **ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → `zpool clear`; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, `S47VNY0K600221`) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches `ONLINE` pools, and ZED's alert went to a root mailbox with no MTA.** nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook `playbooks/ana-ml2-pool-health.yaml`; inventory in `servers/ana-ml2/README.md`. → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`
- `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`
- `[2026-09-08]` **ana-ml2 mesh return routes PERSISTED** as `/etc/network/if-up.d/mesh-routes` (Debian 13 ifupdown, no netplan) via `playbooks/ana-ml2-mesh-routes.yaml` (elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot. `f923d6a`.
- `[2026-09-08]` **ERP run 7 LAUNCHED on pfi-gx10 23:06 PT** under `operator-2026-09-08-rnd-run7` — opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). → `persistent-memory.d/2026-09-08-erp-run7-launched.md`