# NH3 restic prune ritual The quarterly maintenance window where we drop `--append-only` from `rest-server-nh3`, run `forget --prune` on every NH3-side restic client to reclaim space from aged-out snapshots, and put the safety guard back. **Schedule: every ~90 days**, or sooner if NH3 NAS shows disk-pressure on `/volume1/pbs`. The ANA side is fully automated by `scripts/restic-prune.sh ana`. The NH3 side has to be a half-manual ceremony because: * `rest-server-nh3` runs in **DSM Container Manager** on the Synology NAS. The DSM CLI doesn't expose the docker socket on a path our tooling expects. * `sudo` on `syncuser` (the dedicated automation account) isn't NOPASSWD-configured. Toggling the container env from this workstation would require a sudo password over SSH, which our tooling doesn't pipe. So the toggle parts are DSM web-UI work; the prune commands run from your terminal with interactive sudo. Total wall time when nothing unusual happens: **5–10 minutes**. `scripts/restic-prune.sh nh3` prints these same instructions (without the running narrative) — keep that script as the source of truth for the exact commands. This doc is the explanation + lessons learned. --- ## Why "ritual" `rest-server-nh3` boots with `--append-only` by default, which makes the repo immune to ransomware that compromises a backup client: an attacker holding nh3-docker's restic password can `backup` (add snapshots) but can't `forget`/`prune`/`delete` any existing data. Pruning is a privileged operation that requires the safety lifted briefly, then restored. The "ritual" framing keeps the steps in the right order — drop the guard, do the dangerous work, restore the guard — so we don't accidentally leave the repo wide open. If something goes wrong mid-prune (network blip, client crash), the worst case is that `--append-only` is still off when you walk away. The repo is still passworded; an attacker would need both restic creds AND network access to harm it. Phase 3 is mandatory but not emergency-mandatory. --- ## Phase 1 — Drop `--append-only` via DSM In the **DSM web UI** (Synology): 1. Container Manager → **Container** → `rest-server` 2. **Edit** → **Environment** tab 3. Find the `OPTIONS` env var. It currently reads: ``` --private-repos --append-only --prometheus ``` 4. Change to (drop the `--append-only`): ``` --private-repos --prometheus ``` 5. **Apply**. DSM recreates the container with the new env. ### Verification The container should come back up within ~5 seconds. From the workstation: ```bash curl -sf -o /dev/null -w "root: %{http_code}\n" http://10.100.50.50:8000/ # expect: root: 401 (container up, demanding auth — healthy) ``` A non-401 (000 = no response, 5xx = container crashed) means DSM didn't recreate cleanly; check the container's logs in DSM and try **Apply** again before moving on. --- ## Phase 2 — `forget --prune` on every NH3 client Three clients write to `rest-server-nh3`: | host | what it backs up | |---|---| | `nh3-docker` | `/etc`, `/opt/docker`, `/root`, `/var/lib/docker/volumes` | | `nh3-dev` | `/etc`, `/home/lkraven`, `/root` (workstation) | | `irv-ml1` | `/etc`, `/home/lkraven`, `/opt/docker`, `/root`, `/var/lib/docker/volumes`, `/worktank/comfyui/...`, `/worktank/cosyvoice/voices` | Run from your TTY (each is interactive sudo): ```bash ssh -t nh3-docker "sudo resticprofile -c /etc/restic/profiles.yaml forget --prune --verbose" ssh -t nh3-dev "sudo resticprofile -c /etc/restic/profiles.yaml forget --prune --verbose" ssh -t irv-ml1 "sudo resticprofile -c /etc/restic/profiles.yaml forget --prune --verbose" ``` Order doesn't matter; you can also run them in parallel terminals if you want it done faster. ### What to expect * **`forget`** runs the policy from each profile (typically: keep 7 daily, 4 weekly, 12 monthly, 3 yearly). Output lists every snapshot that's being kept and why. * **`prune`** then walks the repo and removes data referenced only by the (now-forgotten) snapshots. **If `forget` removed nothing, `prune` is silent** — nothing to free, nothing to print. * Total per-client wall time: <1s if nothing to do, single-digit minutes if there's a real prune to perform. ### Likely no-op rounds For the first ~6 months of a fresh repo, nothing ages out — every snapshot is still inside the keep window. The output will list all your snapshots under "keep N snapshots:" and print "finished 'forget'" with no removals + no prune output. **This is correct**; the ritual is a guard-restore exercise, not a data-deletion guarantee. Real prune work starts when: * Daily snapshots older than 7 days exist (typically 2nd week of repo life). * Weekly snapshots older than 4 weeks (typically 2nd month). * Monthly snapshots older than 12 months (typically 2nd year). ### If you hit `repository is configured as append-only` You skipped Phase 1 or DSM didn't apply the env change. Re-do Phase 1, verify with the curl probe, then retry. ### If `nh3-dev` prompts for SSH host key authenticity That's a workstation that was rebuilt or never connected from this particular SSH client. Type `yes` to accept the new key, then proceed. The known-hosts entry will then be cached. --- ## Phase 3 — Restore `--append-only` via DSM Mirror image of Phase 1: 1. Container Manager → **Container** → `rest-server` 2. **Edit** → **Environment** 3. Set `OPTIONS` back to: ``` --private-repos --append-only --prometheus ``` 4. **Apply**. ### Verification ```bash curl -sf -o /dev/null -w "root: %{http_code}\n" http://10.100.50.50:8000/ # expect: root: 401 ``` Optionally, prove `--append-only` is back by attempting a no-op forget from any client — it should refuse with `repository is configured as append-only`. **This is destructive-adjacent** (forget attempts a write the server rejects), so it's safe but disruptive to the client's restic history if logging is verbose. Generally trust the DSM UI showing the env change applied. --- ## When to schedule the next round Set a calendar event for **+90 days** from the last successful run. The ANA-side automation runs the same retention policy on its own schedule, so the two sides stay roughly in sync. If `df -h` on `nh3-nas` shows `/volume1/pbs` over 80%, run the ritual sooner rather than waiting for the calendar. Last run history (append a line each time): | date | duration | notes | |---|---|---| | 2026-04-27 | <5 min | First ritual run since pipeline went live; no-op everywhere (every snapshot still in keep window). irv-ml1 had only 3 snapshots due to the 04-25→27 CUDA stall. | --- ## Future automation hooks Once the DSM workflow gets tedious to repeat, the manual phases can be replaced with `ssh -t nh3-nas` + `sudo docker exec` against the DSM-managed container. Two prerequisites: 1. **Find the docker binary path on DSM**. DSM Container Manager doesn't ship `docker` on the default `$PATH` for syncuser. Probe from an interactive shell first: ``` ssh -t nh3-nas "sudo find / -name docker -executable -type f 2>/dev/null" ``` 2. **Set NOPASSWD sudo on syncuser** for the specific docker recreate command. Edit `/etc/sudoers.d/syncuser-docker` carefully — the rule should be path-locked and command-locked so a compromised syncuser can't escalate beyond container restarts. When both are in place, replace the Phase 1 + Phase 3 manual blocks in `scripts/restic-prune.sh` with shell-driven `docker container update` (or container recreate) calls, and the whole ritual becomes a single `scripts/restic-prune.sh nh3` invocation.