feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)

Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
This commit is contained in:
vh
2026-10-02 07:42:45 -07:00
parent 49c8fdf383
commit 902e16630f
4 changed files with 93 additions and 16 deletions
+15 -9
View File
@@ -13,15 +13,21 @@ job closes that gap: an hourly, versioned, off-box snapshot of `~/development`.
- **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged
files hardlink (share inodes, ~0 bytes); only changed files consume new space.
`latest` symlink points at the newest snapshot.
- **Retention:** newest **48** hourly snapshots (older pruned each run). The
prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir
as read-only, and `rm` cannot unlink inside it. Each run logs
`retention prune: <errors> error lines, <N> snapshots on the NAS`; a non-zero
error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves
the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the
prune had no chmod and the run logged OK whatever happened: it failed silently
from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir),
all removed 2026-10-01.
- **Retention (since 2026-10-02, Prime):** the newest **48 hourly** snapshots, plus the newest
snapshot of each of the last **30 days**, plus the newest of each of the last **12 ISO weeks**
(about 80 dirs at steady state; hardlinks keep the extra cost to changed files). Days and weeks
count only those that have snapshots, so an outage does not eat the history.
`~/.config/dev-backup/retention.py` decides what to delete (repo copy:
`scripts/nh3-dev-development-backup-retention.py`). It never selects a name that is not exactly
`YYYY-MM-DD_HHMM`, and the NAS side refuses any path outside that pattern as a second guard.
The prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir as read-only,
and `rm` cannot unlink inside it. Each run logs
`retention prune: deleted <n>, <errors> error lines, <N> snapshots on the NAS (expected ... <K> kept)`;
any error or N ≠ K exits 3, and a failed rsync exits 1, so either one leaves the unit **failed**
(`systemctl --user --failed`). ⚠ Before 2026-10-01 the prune had no chmod and the run logged OK
whatever happened: it failed silently from 2026-07-18 and left 1,740 husk dirs (each holding only
the one 0555 dir), all removed 2026-10-01. History before 2026-09-30 is therefore gone; dailies
and weeklies accumulate from 2026-10-02.
- **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`,
`__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`,
`build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`,