docs(pfi-pve): NASPool rebuilt as raidz2 — runbook log complete, memory entry

This commit is contained in:
vh
2026-09-05 21:04:14 -07:00
parent 0df9312832
commit ab222133e9
3 changed files with 47 additions and 1 deletions
+4
View File
@@ -94,3 +94,7 @@ Keepers from `dump/` (newest per guest, ~295G): qemu 101 102 103 104 105
- 01:46Z VERIFIED: checksum re-pass over the 26 keepers (297G) = 0 diffs (20 min); non-dump dirs = 0 diffs (re-run with correct cwd after a path slip in the first pass). Parked copy is complete and byte-verified. GATE: waiting for operator go before `zpool destroy NASPool`.
- 02:00Z operator GO. Preflight: all six by-id paths resolved to NASPool members (zdb label), serials matched the table, ospool untouched.
- 02:01Z `zpool destroy NASPool` executed; six disks labelcleared + wipefs'd by literal by-id path; new NASPool created as raidz2-0 over the same six disks, ONLINE, no errors. Restore (step 6) launched, log `/root/naspool-rebuild.log`.
- 03:53Z REBUILD DONE: backupStore/db/docker/webdav received (logicalreferenced identical to parked), pve-VMStorage rsynced back (317.8 GB, rc=0, 26 keepers). NASPool raidz2: 2.48T alloc / 30.3T free raw, 21.7T usable, no errors. `.system` deliberately not restored (stays in the parked copy).
- 04:02Z bring-up: pve-truenas re-enabled; CT 109 ana-nas, CT 112 ana-filebot, VM 100 pbs-ana started; NFS exports serving; ana-docker automount restarted, rest-server + backrest healthy (401 on /); `pvesm list pbs-ana` = 116 backups, pve-truenas = 15 entries.
- 04:03Z `zpool scrub NASPool` started. Backup tier offline window: 23:04Z → 04:03Z (5h). Guests on ospool never stopped.
- OPEN: (a) destroy `ospool/naspool-evac` (1.65T) after the scrub is clean AND one Backrest (08:00Z) + one PBS run succeed; (b) `zfs destroy NASPool/*@evac` snapshots then; (c) backplane swap next colo visit → labelclear the two bay-9/10 disks → `zpool add NASPool spare`; (d) PSU1 still dead; (e) old PERC serial 3260688 = spare.
@@ -0,0 +1,41 @@
# 2026-09-06 — pfi-pve NASPool rebuilt as six-wide raidz2 (backplane fault, bays 9/10 dead)
**What happened.** NASPool (TrueNAS-legacy, 4×mirror on 8 disks) had run two mirrors at
single-disk redundancy since ~2026-04-23. Root cause by elimination: R750xs backplane — PERC
PHYs 9 and 10 show no link at all (SAS address zero, no error counters); cables replaced by the
operator and the PERC replaced 2026-09-05 (serial 3260688 → 29R1754) with no change. Old PERC is
a spare, not scrap. Drives in bays 9/10 are healthy but invisible.
**Route taken (operator, 2026-09-05 22:59Z): "Route C hybrid".** Park everything on ospool
(which holds EVERY guest boot disk and was never touched by an irreversible command), destroy,
recreate NASPool as raidz2 over the six linked disks, copy back. Bays 9/10 = labelled cold spares
(pull a failed disk, move a spare into its bay, `zpool replace`). Run on one PSU (PSU1 dead since
08-26) — accepted because a power cut mid-copy costs a re-run, not data.
**Result.** 23:04Z quiesce → 01:20Z parked 1.65T (~200 MB/s, zero source read errors — the copy
doubled as the readability test of the two sole-survivor disks; no pre-scrub) → 01:46Z
byte-verified → 02:00Z operator GO → 02:01Z destroy + raidz2 create (by-id, ashift=12, lz4,
recordsize=1M on backupStore/pve-VMStorage) → 03:53Z restored → 04:03Z backup tier up, scrub
started. Guests on ospool never stopped. NASPool now 21.7T usable, any-two-disk failure tolerance.
**Prune.** Legacy vzdump `dump/` (2.4T, 249 files, last written 2026-04-22, superseded by PBS-ANA)
pruned BY OMISSION: only the newest archive per guest (13 keepers, 295G) was copied out; the rest
died with the old pool. Nothing was `rm`'d.
**Correction to the prior brief.** It claimed guests' disks lived on NASPool and that Route C
meant everything down. Wrong: all 10 VM boot zvols + 3 CT rootfs are on ospool; NASPool is the
backup/bulk tier (PBS-ANA datastore, restic repos, vzdump dir storage). Read `/etc/pve/*.conf`
before believing a storage-placement claim.
**Gotchas recorded.** ana-nas exports pin no `fsid=`, so clients hold stale handles after a
pool rebuild — unmount before, remount after (ana-docker systemd automount; pbs-ana VM reboot).
`logicalreferenced` differed 289K vs 280K on a dataset whose content was `diff -r` identical —
compression-property accounting, not data. rsync `--files-from` paths are relative to the
source arg; a verify pass run from the wrong cwd returned a meaningless "0 diffs" alongside errors
— read the rc, not the count.
**Open.** Parked copy `ospool/naspool-evac` (1.65T) stays until the scrub is clean and one
Backrest + one PBS run succeed, then destroy. `@evac` snapshots to drop then. Backplane swap at
next colo visit (Dell R750xs 12×3.5" SAS backplane; front-half teardown), then
`zpool add NASPool spare <bay9> <bay10>`. PSU1 still dead. Runbook:
`docs/runbooks/pfi-pve-naspool-rebuild.md`. Booth: http://10.100.10.50:8090/b/pfi-pve-naspool-plan/
+2 -1
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-05_
_Last updated: 2026-09-06_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -473,6 +473,7 @@ below is a live commitment or a known-open risk._
## Recent decisions
- `[2026-09-06]` **pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10** (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroy `ospool/naspool-evac` after scrub + one backup cycle; backplane swap next visit; PSU1 still dead. → `persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md`
- `[2026-09-05]` **A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him.** Each floor was `|b0-b1|` from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named *why*: the claim was his and flattering. → `persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md`
- `[2026-09-05]` **vLLM RUNS on sm_121 — the blocker was `ninja` off PATH, not the silicon** — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missing `root_sha256` values I knew, because supplying both sides of a check makes it inert. → `persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md`
- `[2026-09-04]` **ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression.** LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. → `persistent-memory.d/2026-09-04-run3c-trained-and-gated.md`