fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune

rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
This commit is contained in:
vh
2026-10-01 05:59:07 -07:00
parent 9af5a16d07
commit de11e00dfb
3 changed files with 26 additions and 14 deletions
+9 -1
View File
@@ -13,7 +13,15 @@ job closes that gap: an hourly, versioned, off-box snapshot of `~/development`.
- **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged
files hardlink (share inodes, ~0 bytes); only changed files consume new space.
`latest` symlink points at the newest snapshot.
- **Retention:** newest **48** hourly snapshots (older pruned each run).
- **Retention:** newest **48** hourly snapshots (older pruned each run). The
prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir
as read-only, and `rm` cannot unlink inside it. Each run logs
`retention prune: <errors> error lines, <N> snapshots on the NAS`; a non-zero
error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves
the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the
prune had no chmod and the run logged OK whatever happened: it failed silently
from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir),
all removed 2026-10-01.
- **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`,
`__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`,
`build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`,