feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)

Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
This commit is contained in:
vh
2026-10-02 07:42:45 -07:00
parent 49c8fdf383
commit 902e16630f
4 changed files with 93 additions and 16 deletions
+15 -9
View File
@@ -13,15 +13,21 @@ job closes that gap: an hourly, versioned, off-box snapshot of `~/development`.
- **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged
files hardlink (share inodes, ~0 bytes); only changed files consume new space.
`latest` symlink points at the newest snapshot.
- **Retention:** newest **48** hourly snapshots (older pruned each run). The
prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir
as read-only, and `rm` cannot unlink inside it. Each run logs
`retention prune: <errors> error lines, <N> snapshots on the NAS`; a non-zero
error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves
the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the
prune had no chmod and the run logged OK whatever happened: it failed silently
from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir),
all removed 2026-10-01.
- **Retention (since 2026-10-02, Prime):** the newest **48 hourly** snapshots, plus the newest
snapshot of each of the last **30 days**, plus the newest of each of the last **12 ISO weeks**
(about 80 dirs at steady state; hardlinks keep the extra cost to changed files). Days and weeks
count only those that have snapshots, so an outage does not eat the history.
`~/.config/dev-backup/retention.py` decides what to delete (repo copy:
`scripts/nh3-dev-development-backup-retention.py`). It never selects a name that is not exactly
`YYYY-MM-DD_HHMM`, and the NAS side refuses any path outside that pattern as a second guard.
The prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir as read-only,
and `rm` cannot unlink inside it. Each run logs
`retention prune: deleted <n>, <errors> error lines, <N> snapshots on the NAS (expected ... <K> kept)`;
any error or N ≠ K exits 3, and a failed rsync exits 1, so either one leaves the unit **failed**
(`systemctl --user --failed`). ⚠ Before 2026-10-01 the prune had no chmod and the run logged OK
whatever happened: it failed silently from 2026-07-18 and left 1,740 husk dirs (each holding only
the one 0555 dir), all removed 2026-10-01. History before 2026-09-30 is therefore gone; dailies
and weeklies accumulate from 2026-10-02.
- **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`,
`__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`,
`build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`,
+2
View File
@@ -291,6 +291,8 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions
- `[2026-10-02]` **Prime: dev-backup gets dailies + weeklies — DONE.** Retention is now 48 hourly + newest of 30 days + newest of 12 ISO weeks (`retention.py` beside the script; second path guard on the NAS side; unit fails if kept ≠ expected). Live run 0739: deleted 1, 0 errors, 48 kept = expected. History before 09-30 was already gone; dailies accumulate from today.
- `[2026-10-02]` **Prime: MS-03s get Proxmox; dummy plugs in hand.** Still open: what the second MS-03 is for. On arrival: check the NIC chipset (I226-LM = keep the AMT port admin-UP), fit a plug on each, then the parked AMT follow-ups.
- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next NATURAL reboot (Prime 2026-10-02: no test reboot).** Trigger: `uptime -s` on nh3-dev later than 2026-10-02 0733. Then check the boot was clean (`swapon --show` = /swapfile; `systemd-analyze` shows no ~30 s stall; `journalctl -b | grep -i resume` has no 'waiting for resume device'), and only then `qm delsnapshot 102 pre-rootgrow-20261002`.
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
@@ -0,0 +1,58 @@
#!/usr/bin/env python3
"""dev-backup retention: decide which snapshot dirs to DELETE (Prime, 2026-10-02: add dailies and weeklies).
Reads snapshot paths on stdin, one per line (`/volume1/Backup/nh3-dev-development/YYYY-MM-DD_HHMM`), and
prints the ones to delete. KEEP is the union of:
- the newest HOURLY snapshots (default 48),
- the newest snapshot of each of the most recent DAILY distinct days (default 30),
- the newest snapshot of each of the most recent WEEKLY distinct ISO weeks (default 12).
Days and weeks are counted among those that HAVE snapshots, so an outage does not eat the history.
A line whose basename is not exactly YYYY-MM-DD_HHMM is never printed, so it is never deleted.
python3 retention.py [--hourly 48] [--daily 30] [--weekly 12] [--keep-count] < listing
--keep-count prints only the number of snapshots that will remain (for the post-prune check).
"""
import argparse
import datetime as dt
import os
import re
import sys
NAME = re.compile(r"^(\d{4})-(\d{2})-(\d{2})_(\d{4})$")
def select_delete(paths, hourly, daily, weekly):
snaps = []
for p in paths:
p = p.strip()
m = NAME.match(os.path.basename(p))
if not p or not m:
continue
try:
day = dt.date(int(m[1]), int(m[2]), int(m[3]))
except ValueError: # e.g. month 13: not one of ours, so never delete it
continue
snaps.append((os.path.basename(p), p, day))
snaps.sort(reverse=True) # newest first; the names sort chronologically
keep = {s[1] for s in snaps[:hourly]}
days, weeks = [], []
for name, path, day in snaps:
if day not in days and len(days) < daily:
days.append(day)
keep.add(path) # first seen = newest of that day
wk = day.isocalendar()[:2]
if wk not in weeks and len(weeks) < weekly:
weeks.append(wk)
keep.add(path)
return [s[1] for s in snaps if s[1] not in keep], len(keep)
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("--hourly", type=int, default=48)
ap.add_argument("--daily", type=int, default=30)
ap.add_argument("--weekly", type=int, default=12)
ap.add_argument("--keep-count", action="store_true")
a = ap.parse_args()
delete, kept = select_delete(sys.stdin.read().splitlines(), a.hourly, a.daily, a.weekly)
print(kept if a.keep_count else "\n".join(delete))
+18 -7
View File
@@ -4,7 +4,7 @@
# dev work now has an hourly, versioned, off-box safety net. Secrets + heavy
# reconstructable dirs are excluded. Snapshots are timestamped dirs on the NAS;
# unchanged files hardlink to the previous snapshot (space-efficient). Retention:
# newest 48 hourly snapshots.
# 48 hourly + 30 daily + 12 weekly (retention.py; dailies and weeklies added 2026-10-02).
set -uo pipefail
SRC="$HOME/development/"
@@ -38,16 +38,27 @@ echo "rsync rc=$RC"
# rc 0 = ok; rc 24 = some files vanished mid-transfer (benign for a live tree)
if [ "$RC" -eq 0 ] || [ "$RC" -eq 24 ]; then
ssh -o BatchMode=yes "$DEST_HOST" "ln -sfn '$DEST_BASE/$STAMP' '$DEST_BASE/latest'"
# retention: keep newest 48 hourly snapshots.
# retention (Prime, 2026-10-02): newest 48 hourly + newest of each of the last 30 days + newest of
# each of the last 12 ISO weeks. retention.py (next to this script) picks what to delete; anything
# whose name is not exactly YYYY-MM-DD_HHMM is never selected, and the remote side refuses any path
# outside $DEST_BASE/20??-??-??_???? as a second guard.
# chmod BEFORE rm: rsync -a copies a read-only source dir (mode 0555) as read-only, and rm cannot
# unlink inside it. Without the chmod this prune failed every run from 2026-07-18 to 2026-10-01 and
# left 1,740 husk dirs behind (removed 2026-10-01), while the run still logged OK. Errors are COUNTED,
# not dumped (one failing run logged ~400 MB), and a failed prune now fails the unit.
PRUNE_LIST="ls -1d $DEST_BASE/20* 2>/dev/null | sort | head -n -48"
PRUNE_ERR="$(ssh -o BatchMode=yes "$DEST_HOST" "$PRUNE_LIST | xargs -r chmod -R u+w 2>&1; $PRUNE_LIST | xargs -r rm -rf 2>&1" | wc -l)"
# not dumped (one failing run logged ~400 MB), and a failed prune fails the unit.
ALL="$(ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null")"
DELETE="$(printf '%s\n' "$ALL" | python3 "$(dirname "$0")/retention.py")"
EXPECT="$(printf '%s\n' "$ALL" | python3 "$(dirname "$0")/retention.py" --keep-count)"
NDEL="$(printf '%s' "$DELETE" | grep -c . || true)"
PRUNE_ERR="$(printf '%s\n' "$DELETE" | grep . | ssh -o BatchMode=yes "$DEST_HOST" "while read -r d; do
case \"\$d\" in
$DEST_BASE/20[0-9][0-9]-[0-9][0-9]-[0-9][0-9]_[0-9][0-9][0-9][0-9]) chmod -R u+w \"\$d\" 2>&1 && rm -rf \"\$d\" 2>&1 || echo \"FAILED \$d\" ;;
*) echo \"REFUSED \$d\" ;;
esac
done" | wc -l)"
KEPT="$(ssh -o BatchMode=yes "$DEST_HOST" "ls -1d $DEST_BASE/20* 2>/dev/null | wc -l")"
echo "retention prune: ${PRUNE_ERR:-?} error lines, ${KEPT:-?} snapshots on the NAS (expected 0 and <=48)"
if [ "${PRUNE_ERR:-x}" = 0 ] && [ "${KEPT:-999}" -le 48 ] 2>/dev/null; then
echo "retention prune: deleted ${NDEL:-?}, ${PRUNE_ERR:-?} error lines, ${KEPT:-?} snapshots on the NAS (expected 0 errors and ${EXPECT:-?} kept)"
if [ "${PRUNE_ERR:-x}" = 0 ] && [ -n "${EXPECT:-}" ] && [ "${KEPT:-x}" = "$EXPECT" ]; then
echo "=== $(date -Is) snapshot $STAMP OK (rc=$RC) ==="
else
echo "=== $(date -Is) snapshot $STAMP taken (rc=$RC) but RETENTION FAILED ==="