From de11e00dfb3899869f4decef04723f0d9adcfc45 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 1 Oct 2026 05:59:07 -0700 Subject: [PATCH] fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune rsync -a copies a read-only source dir (0555) as read-only, so the hourly prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned snapshot was left as a 22-entry husk while the run still logged OK. The 1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept). The prune now runs chmod -R u+w before rm -rf, logs its error count and the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a failed rsync) so the systemd unit shows failed instead of passing. --- docs/runbooks/nh3-dev-development-backup.md | 10 +++++++++- persistent-memory.md | 11 +---------- scripts/nh3-dev-development-backup.sh | 19 ++++++++++++++++--- 3 files changed, 26 insertions(+), 14 deletions(-) diff --git a/docs/runbooks/nh3-dev-development-backup.md b/docs/runbooks/nh3-dev-development-backup.md index cf462a7..f548183 100644 --- a/docs/runbooks/nh3-dev-development-backup.md +++ b/docs/runbooks/nh3-dev-development-backup.md @@ -13,7 +13,15 @@ job closes that gap: an hourly, versioned, off-box snapshot of `~/development`. - **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged files hardlink (share inodes, ~0 bytes); only changed files consume new space. `latest` symlink points at the newest snapshot. -- **Retention:** newest **48** hourly snapshots (older pruned each run). +- **Retention:** newest **48** hourly snapshots (older pruned each run). The + prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir + as read-only, and `rm` cannot unlink inside it. Each run logs + `retention prune: error lines, snapshots on the NAS`; a non-zero + error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves + the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the + prune had no chmod and the run logged OK whatever happened: it failed silently + from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir), + all removed 2026-10-01. - **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`, `__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`, `build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`, diff --git a/persistent-memory.md b/persistent-memory.md index bf1a9c6..7ddebc7 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -134,16 +134,6 @@ _As of 2026-10-01 ~0446 PT._ - Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects). - NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` -### dev-backup (nh3-dev `~/development` → nh3-nas hourly): retention BROKEN since 2026-07-18 (found 2026-10-01) - -- **Its retention prune has failed every run since 2026-07-18** with "Permission denied", because read-only dirs were copied from the source. **1,788 snapshots** sit in `nh3-nas:/volume1/Backup/nh3-dev-development` against a design of 48. The script still logged "snapshot OK": an instrument that passes in both states. -- The log had reached 4.6 GB (25.7M error lines) and tripped the nh3-dev 85% disk alert. I rotated it to `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` (99 MB) and made the script COUNT prune errors instead of dumping them (`dev-backup.sh.bak-20261001`). nh3-dev is now at 84%; nh3-nas /volume1 is at 78% of 42 TB. -- **Awaiting Prime (one-way for any deleted history):** - - fix the prune (`chmod -R u+w` before `rm`) and prune to the designed 48; - - OR switch to a tiered retention (48 hourly + 30 daily + 12 weekly) and prune the rest; - - OR keep everything for now. - - Until he rules, the prune keeps failing and its error count shows in the log. Nothing is deleted. - ### Worldtree U11 memory cutover (demo + personal) - **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` @@ -300,6 +290,7 @@ _As of 2026-10-01 ~0446 PT._ ## Recent decisions +- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it. - `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap. - `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them. - `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it. diff --git a/scripts/nh3-dev-development-backup.sh b/scripts/nh3-dev-development-backup.sh index 9949056..0568a04 100755 --- a/scripts/nh3-dev-development-backup.sh +++ b/scripts/nh3-dev-development-backup.sh @@ -38,9 +38,22 @@ echo "rsync rc=$RC" # rc 0 = ok; rc 24 = some files vanished mid-transfer (benign for a live tree) if [ "$RC" -eq 0 ] || [ "$RC" -eq 24 ]; then ssh -o BatchMode=yes "$DEST_HOST" "ln -sfn '$DEST_BASE/$STAMP' '$DEST_BASE/latest'" - # retention: keep newest 48 hourly snapshots - ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null | sort | head -n -48 | xargs -r rm -rf" - echo "=== $(date -Is) snapshot $STAMP OK (rc=$RC) ===" + # retention: keep newest 48 hourly snapshots. + # chmod BEFORE rm: rsync -a copies a read-only source dir (mode 0555) as read-only, and rm cannot + # unlink inside it. Without the chmod this prune failed every run from 2026-07-18 to 2026-10-01 and + # left 1,740 husk dirs behind (removed 2026-10-01), while the run still logged OK. Errors are COUNTED, + # not dumped (one failing run logged ~400 MB), and a failed prune now fails the unit. + PRUNE_LIST="ls -1d $DEST_BASE/20* 2>/dev/null | sort | head -n -48" + PRUNE_ERR="$(ssh -o BatchMode=yes "$DEST_HOST" "$PRUNE_LIST | xargs -r chmod -R u+w 2>&1; $PRUNE_LIST | xargs -r rm -rf 2>&1" | wc -l)" + KEPT="$(ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null | wc -l")" + echo "retention prune: ${PRUNE_ERR:-?} error lines, ${KEPT:-?} snapshots on the NAS (expected 0 and <=48)" + if [ "${PRUNE_ERR:-x}" = 0 ] && [ "${KEPT:-999}" -le 48 ] 2>/dev/null; then + echo "=== $(date -Is) snapshot $STAMP OK (rc=$RC) ===" + else + echo "=== $(date -Is) snapshot $STAMP taken (rc=$RC) but RETENTION FAILED ===" + exit 3 + fi else echo "=== $(date -Is) snapshot $STAMP FAILED rc=$RC — keeping partial for inspection ===" + exit 1 fi