fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune

rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
This commit is contained in:
vh
2026-10-01 05:59:07 -07:00
parent 9af5a16d07
commit de11e00dfb
3 changed files with 26 additions and 14 deletions
+9 -1
View File
@@ -13,7 +13,15 @@ job closes that gap: an hourly, versioned, off-box snapshot of `~/development`.
- **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged
files hardlink (share inodes, ~0 bytes); only changed files consume new space.
`latest` symlink points at the newest snapshot.
- **Retention:** newest **48** hourly snapshots (older pruned each run).
- **Retention:** newest **48** hourly snapshots (older pruned each run). The
prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir
as read-only, and `rm` cannot unlink inside it. Each run logs
`retention prune: <errors> error lines, <N> snapshots on the NAS`; a non-zero
error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves
the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the
prune had no chmod and the run logged OK whatever happened: it failed silently
from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir),
all removed 2026-10-01.
- **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`,
`__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`,
`build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`,
+1 -10
View File
@@ -134,16 +134,6 @@ _As of 2026-10-01 ~0446 PT._
- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects).
- NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
### dev-backup (nh3-dev `~/development` → nh3-nas hourly): retention BROKEN since 2026-07-18 (found 2026-10-01)
- **Its retention prune has failed every run since 2026-07-18** with "Permission denied", because read-only dirs were copied from the source. **1,788 snapshots** sit in `nh3-nas:/volume1/Backup/nh3-dev-development` against a design of 48. The script still logged "snapshot OK": an instrument that passes in both states.
- The log had reached 4.6 GB (25.7M error lines) and tripped the nh3-dev 85% disk alert. I rotated it to `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` (99 MB) and made the script COUNT prune errors instead of dumping them (`dev-backup.sh.bak-20261001`). nh3-dev is now at 84%; nh3-nas /volume1 is at 78% of 42 TB.
- **Awaiting Prime (one-way for any deleted history):**
- fix the prune (`chmod -R u+w` before `rm`) and prune to the designed 48;
- OR switch to a tiered retention (48 hourly + 30 daily + 12 weekly) and prune the rest;
- OR keep everything for now.
- Until he rules, the prune keeps failing and its error count shows in the log. Nothing is deleted.
### Worldtree U11 memory cutover (demo + personal)
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
@@ -300,6 +290,7 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
+15 -2
View File
@@ -38,9 +38,22 @@ echo "rsync rc=$RC"
# rc 0 = ok; rc 24 = some files vanished mid-transfer (benign for a live tree)
if [ "$RC" -eq 0 ] || [ "$RC" -eq 24 ]; then
ssh -o BatchMode=yes "$DEST_HOST" "ln -sfn '$DEST_BASE/$STAMP' '$DEST_BASE/latest'"
# retention: keep newest 48 hourly snapshots
ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null | sort | head -n -48 | xargs -r rm -rf"
# retention: keep newest 48 hourly snapshots.
# chmod BEFORE rm: rsync -a copies a read-only source dir (mode 0555) as read-only, and rm cannot
# unlink inside it. Without the chmod this prune failed every run from 2026-07-18 to 2026-10-01 and
# left 1,740 husk dirs behind (removed 2026-10-01), while the run still logged OK. Errors are COUNTED,
# not dumped (one failing run logged ~400 MB), and a failed prune now fails the unit.
PRUNE_LIST="ls -1d $DEST_BASE/20* 2>/dev/null | sort | head -n -48"
PRUNE_ERR="$(ssh -o BatchMode=yes "$DEST_HOST" "$PRUNE_LIST | xargs -r chmod -R u+w 2>&1; $PRUNE_LIST | xargs -r rm -rf 2>&1" | wc -l)"
KEPT="$(ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null | wc -l")"
echo "retention prune: ${PRUNE_ERR:-?} error lines, ${KEPT:-?} snapshots on the NAS (expected 0 and <=48)"
if [ "${PRUNE_ERR:-x}" = 0 ] && [ "${KEPT:-999}" -le 48 ] 2>/dev/null; then
echo "=== $(date -Is) snapshot $STAMP OK (rc=$RC) ==="
else
echo "=== $(date -Is) snapshot $STAMP taken (rc=$RC) but RETENTION FAILED ==="
exit 3
fi
else
echo "=== $(date -Is) snapshot $STAMP FAILED rc=$RC — keeping partial for inspection ==="
exit 1
fi