Files
esh-pfi-infrastructure/docs/runbooks/nh3-dev-development-backup.md
T
vh de11e00dfb fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
2026-10-01 05:59:07 -07:00

66 lines
3.2 KiB
Markdown

# nh3-dev `~/development` — hourly off-box backup
**Why this exists:** nh3-dev is the dev box where agents do uncommitted work under
`~/development/<project>/`. That tree had **no off-box backup**, so a destructive
mistake (a stray `rm -rf` on a working dir on 2026-07-12) had no safety net. This
job closes that gap: an hourly, versioned, off-box snapshot of `~/development`.
## What it does
- **Source:** `nh3-dev:~/development/` (lkraven's working dirs).
- **Destination (off-box):** `nh3-nas:/volume1/Backup/nh3-dev-development/<YYYY-MM-DD_HHMM>/`
— a timestamped dir per snapshot, over rsync-**over-ssh** (syncuser).
- **Versioning:** `rsync --link-dest` against the previous snapshot → unchanged
files hardlink (share inodes, ~0 bytes); only changed files consume new space.
`latest` symlink points at the newest snapshot.
- **Retention:** newest **48** hourly snapshots (older pruned each run). The
prune runs `chmod -R u+w` before `rm -rf`: rsync copies a read-only source dir
as read-only, and `rm` cannot unlink inside it. Each run logs
`retention prune: <errors> error lines, <N> snapshots on the NAS`; a non-zero
error count or N > 48 exits 3, and a failed rsync exits 1, so either one leaves
the unit **failed** (`systemctl --user --failed`). ⚠ Before 2026-10-01 the
prune had no chmod and the run logged OK whatever happened: it failed silently
from 2026-07-18 and left 1,740 husk dirs (each holding only the one 0555 dir),
all removed 2026-10-01.
- **Excludes:** heavy reconstructable dirs (`node_modules`, `.venv`, `venv`,
`__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.cache`, `dist`,
`build`, `.next`, `target`, `*.pyc`) and secrets (`.env`, `.env.*`, `*.pem`,
`*.key`, `id_*`, `*.sqlite*`). **`.git` is kept** (local commits/stashes = the
uncommitted work that matters). Seed snapshot ≈ **11G**; hourly deltas are MB-scale.
## Where it lives (on nh3-dev)
- Script: `~/.config/dev-backup/dev-backup.sh` (mirror committed at
`scripts/nh3-dev-development-backup.sh`).
- systemd `--user` units: `~/.config/systemd/user/dev-backup.{service,timer}`
(`OnCalendar=hourly`, `Persistent=true`, linger on → fires without a login).
- Log: `~/.config/dev-backup/dev-backup.log`.
```bash
systemctl --user list-timers dev-backup.timer # next run
systemctl --user start dev-backup.service # run now
tail -f ~/.config/dev-backup/dev-backup.log
```
## Restore
Snapshots are plain dir trees — no special tool needed:
```bash
# list snapshots
ssh nh3-nas 'ls -1 /volume1/Backup/nh3-dev-development/'
# restore one file/dir from a chosen snapshot
rsync -a nh3-nas:/volume1/Backup/nh3-dev-development/<STAMP>/<proj>/<path> /tmp/restore/
# or pull a whole project back
rsync -a nh3-nas:/volume1/Backup/nh3-dev-development/latest/<proj>/ ~/development/<proj>/
```
## Notes / future
- **Not encrypted at rest** (plaintext on the trusted internal NAS; secrets are
excluded). Upgrade path: migrate to restic once a repo can be created on
rest-server-nh3 (currently returns 404 on repo-create — likely append-only) or
the Synology sftp subsystem is enabled (currently disabled → restic sftp fails).
- Off-box = off the nh3-dev VM (lands on nh3-nas, same NH3 site). Cross-site
mirroring of this repo is a separate future layer.