backup pipeline: configs, runbooks, NH3 Synology rest-server, cross-site rsync

Bundles the post-2026-04-21 work that built out the two-layer backup
architecture (PBS for VM images + restic for file/DB), plus the cross-
site mirror and the disaster-recovery runbook.

- configs/restic/esh-docker-vm/profiles.yaml: drop the obsolete
  *_offen_backup_data exclude (offen sidecars retired fleet-wide
  2026-04-23; restic now covers the equivalent scope directly).
- configs/restic/esh-vm-db/: new profile for the dedicated DB VM
  (10.0.50.60), with pre-backup pg_dumpall + mongodump hooks.
- configs/rsync/: ana-nas → nh3-nas (04:00 daily, runs as lkraven)
  and nh3-nas → ana-nas (05:00 daily, runs as root because DSM
  rest-server-nh3 writes mode-400 files only root can read).
- docs/runbooks/pbs-deployment.md: 9-phase PBS rollout runbook,
  refined during the 2026-04-22 deployment with per-hypervisor
  namespaces, NFSv3 + ZFS-case-insensitivity workaround, and the
  Synology syno_acl flatten step.
- docs/runbooks/disaster-recovery.md: blast-radius runbook ordered
  Tier 0 → 5 (ana-nas → hypervisors → Docker hosts → VMs → specialty);
  references incident memory + recovery-step playbooks per consumer.
This commit is contained in:
vh
2026-04-24 21:56:22 -07:00
parent 4971e5ad41
commit 574c72daa5
12 changed files with 1114 additions and 20 deletions
+276
View File
@@ -0,0 +1,276 @@
# Fleet disaster-recovery runbook
What breaks when a given host/service goes down, and how to recover.
Ordered by blast-radius severity — read top-down.
**First action for ANY multi-service outage:** check ana-nas
reachability. Many fleet-wide "things are broken" events trace back
to ana-nas's NFS exports going stale. See [ana-nas SPOF memory](../../../.claude/projects/-home-lkraven-development-eshpfi-management/memory/incident_ana_nas_spof.md)
for details, or just `ping 10.250.50.50` first.
## Governing principle: backups must not risk production
Any backup strategy that could take down the host it's backing up is
worse than the risk it mitigates. Specifically applies to:
- **NFS-serving hosts (CT 109 / ana-nas):** scheduled vzdump of the
container itself is prohibited. The rootfs is trivially rebuildable;
the data lives on ospool bind-mounts captured via ZFS snapshot at
the host level. If rootfs config is worth preserving, capture via
`rsync /etc /root` to a known location — no snapshot, no freeze.
- **DB servers on NFS storage (VM 105 / pfi-postgres):** prefer
`pg_dump` + restic over full vzdump. fs-freeze on NFS-backed
PGDATA is a failure mode waiting to happen.
- **Any host with heavy I/O concurrency during the backup window:**
schedule its own backup outside the fleet window OR use `mode=stop`
(planned brief downtime beats random-freeze risk).
Incident reference: 2026-04-23 — CT 109, 112, 113 all went offline on
pfi-pve during morning backup window. Backup suspected as trigger
(full root cause TBD). See `memory/incident_ana_nas_spof.md`.
---
## Tier 0 — Storage fabric (catastrophic blast radius)
### ana-nas (CT 109 on pfi-pve, 10.250.50.50)
**Blast radius:**
- ~~pfi-postgres (VM 105) — PGDATA on `/mnt/db`~~ — **migrated to local disk 2026-04-23**. vaultwarden, gitea, paperless-ng, zammad no longer cascade on ana-nas outage. Left in history for the recovery pre-migration.
- ana-docker rest-server-ana — repo data on `/mnt/backup` → all ana-side restic clients fail (ana-docker, ana-ml2, esh-docker-vm, vm-esh-nas)
- PBS-ANA datastore — NFS-backed on `/mnt/backup/pbs-ana` → fleet vzdumps fail, PBS-NH3 sync fails
- ana-docker NFS mounts for `/mnt/docker`, `/mnt/compose`, `/mnt/pve-VMStorage` if used → various stack misbehavior
**Recovery** (see full procedure in `memory/incident_ana_nas_spof.md`):
1. `pct start 109` on pfi-pve if stopped; `pct console 109` to see boot if stuck
2. Once responsive, remount NFS on each consumer and restart services:
- pfi-postgres: `umount -lf /mnt/db; mount /mnt/db; systemctl restart postgresql`
- ana-docker: `sudo mount -a; sudo docker restart rest-server` — usually enough. `mount -a` re-attempts all fstab entries and bypasses the "failed" state that `mnt-backup.mount` gets stuck in (fstab uses bare `defaults` without auto-retry). If rest-server still errors with `/data/.htpasswd: permission denied`, you've got a **ghost file on the local mount point** — see incident memory for the clean-up procedure (stop rest-server, unmount NFS, rm the ghost, remount, restart).
- PBS-ANA: `umount -lf /mnt/pbs-datastore; mount /mnt/pbs-datastore; systemctl restart proxmox-backup proxmox-backup-proxy`
3. Re-run failed PBS vzdump job for affected VMs (Datacenter → Backup → Run Now)
4. Restic clients recover automatically on next 01:00 schedule
**Prevention:**
- Re-schedule ana-nas's own vzdump OUTSIDE the 03:00 fleet window (concurrent NFS load during fleet backup + self-backup is the suspected 2026-04-23 crash root cause)
- Consider moving pfi-postgres PGDATA to local VM disk (reduces SPOF)
---
### nh3-nas (Synology RS2418+, 10.100.50.50)
**Blast radius:**
- nh3-docker, nh3-dev restic clients — writes to `rest-server-nh3` fail
- PBS-NH3 datastore — mirror sync can't write new chunks
- ana-nas → nh3-nas rsync (04:00 daily) — fails
**Recovery:**
1. Power-cycle via DSM web UI or physical button if unresponsive
2. DSM boots 3-5 minutes; NFS exports auto-start
3. Verify PBS-NH3 datastore recovered:
```
ssh pbs-nh3 'ls /mnt/pbs-datastore/.chunks | head; systemctl restart proxmox-backup proxmox-backup-proxy'
```
4. ana-nas-side rsync retries at next 04:00; no manual action unless you want to trigger now
**Notable:** nh3-nas outage is LESS severe than ana-nas because:
- No production DB depends on it
- It's a mirror/DR tier, not primary
- Ana-side backups keep running independently
---
## Tier 1 — Hypervisors (regional blast radius)
### pfi-pve (Proxmox VE, 10.250.250.31)
**Blast radius — everything hosted on it:**
- ana-docker (VM), ana-nas (CT 109), pfi-postgres (VM 105), pfi-ana-webhost, ana-filebot, pfi-pteradactyl, pfi-tacticalrmm, ana-wg, PBS-ANA (VM), + others
Essentially **all Anaheim primary services** go offline. Because ana-nas lives here too, Tier-0 cascade applies simultaneously.
**Recovery:**
1. Is the Proxmox host itself up? iDRAC/IPMI console for hardware diagnosis.
2. If the host is up but VMs/CTs aren't starting: `systemctl status pve-cluster qemu-server pve-container` on the host.
3. Start VMs/CTs in dependency order:
- **First**: CT 109 (ana-nas) — everything else on the host that uses NFS depends on it
- **Second**: PBS-ANA VM (resume backup target)
- **Third**: VM 105 (pfi-postgres) — many apps wait on this
- **Fourth**: ana-docker, then other VMs
4. After all VMs/CTs up, re-apply the Tier-0 ana-nas recovery steps for any consumers with stale NFS mounts.
**Prevention:** PBS-ANA → PBS-NH3 nightly sync means VM image backups are recoverable at NH3 if pfi-pve is unrecoverable; can restore key VMs to nh3-pve and cut over (cold DR).
---
### nh3-pve (10.100.250.60)
**Blast radius:**
- nh3-docker (VM) — NH3 site's Docker stacks
- PBS-NH3 (VM) — DR mirror target
- Other NH3 VMs
**Recovery:** same pattern as pfi-pve. PBS-NH3 being down breaks DR replication but doesn't affect primary backups (ANA side is authoritative).
---
### esh-pve + esh-pve-nas (10.0.250.35 + 10.0.50.55)
**Blast radius:**
- All ESH home-lab workloads: esh-docker-vm, vm-esh-nas, media services, home automation stacks
**Recovery:** restore VMs from PBS-ANA (cross-site pull). Not production-critical; no SLA.
---
### sfsrv-ana (Dell R630, 10.250.250.115, iDRAC 10.250.250.110)
**Blast radius:** SureFire tenant workloads.
- PBS-ANA namespace `sfsrv-pve` has daily snapshots
- Coordinate with tenant before any recovery action per hosting agreement
**Recovery:** restore from PBS-ANA sfsrv-pve namespace. Tenant may have app-level recovery procedures that take precedence.
---
## Tier 2 — Docker hosts (per-stack blast radius)
### ana-docker (VM 10.250.50.70)
**Blast radius — services hosted on it:**
- rest-server-ana (fleet restic target; depends on ana-nas NFS)
- Backrest UI, Beszel hub, Dozzle hub, Traefik (ANA), CrowdSec, Gitea, Vaultwarden, Seafile, Synapse, OpenWebUI, SearXNG, IT Tools, others
**Recovery:**
1. VM restart on pfi-pve: `qm stop <vmid>; qm start <vmid>`
2. After boot, stacks auto-start via `docker compose up -d` (compose files at `/opt/docker/compose/*/`)
3. Check each stack with `docker ps` / dockge
4. If `/mnt/backup` (NFS) is stale after pfi-pve reboot, force remount
### nh3-docker (VM 10.100.50.40)
**Blast radius:**
- rest-server-nh3 (if migrated there — currently on nh3-nas Synology directly)
- NH3-only stacks
**Recovery:** VM restart on nh3-pve. Cross-site Beszel/Dozzle agents on ana-docker keep reporting independently.
### esh-docker-vm + vm-esh-nas
**Blast radius:** ESH home-lab Docker workloads. Paperless, HA, CWA, pgadmin (esh-docker-vm); dockge, filezilla, agents (vm-esh-nas).
**Recovery:** VM restart on respective ESH hypervisors. Not production-critical.
---
## Tier 3 — Workload VMs/LXCs
### pfi-postgres (VM 105 on pfi-pve)
**Blast radius:** vaultwarden, gitea, paperless-ng databases.
**Recovery:** usually it's actually an ana-nas issue (see Tier 0). If postgres itself is the problem:
1. Check logs: `sudo tail -200 /var/lib/postgresql/*/main/log/postgresql-*.log`
2. If clean: `sudo systemctl restart postgresql`
3. If WAL corruption: restore from PBS vzdump (most recent snapshot)
### ana-wg (CT 113 on pfi-pve) — WireGuard VPN
**Blast radius:** remote-access VPN down; remote admin sessions drop but on-prem ops continue.
**Recovery:** `pct start 113` on pfi-pve. Restart wg-quick service if needed.
### pfi-pteradactyl (VM 107 on pfi-pve) — game panel
**Blast radius:** hosted game servers offline.
**Recovery:** VM restart. Not operationally critical.
### pfi-tacticalrmm (VM 111 on pfi-pve) — RMM tooling
**Blast radius:** RMM agent dashboards; agents continue running on endpoints but can't phone home.
**Recovery:** VM restart; TacticalRMM services auto-start.
### ana-filebot (CT 112 on pfi-pve) — file automation
**Blast radius:** scheduled file operations (rename, organize); low-criticality.
**Recovery:** `pct start 112`.
### pfi-ana-webhost (VM on pfi-pve) — web workload
**Blast radius:** hosted website(s) offline.
**Recovery:** VM restart; web server auto-start.
---
## Tier 4 — Specialty workloads
### ana-ml2 (bare metal Supermicro, 10.250.50.54, BMC 10.250.250.50)
**Blast radius:** AI inference services (llama-swap, vllm-qwen3). Consumer-facing chat/embedding endpoints fail.
**Recovery:**
1. Check OS via SSH. If unresponsive, BMC console at <https://10.250.250.50>.
2. If hardware issue: BMC logs, power cycle via IPMI, check GPU health (`nvidia-smi`).
3. Docker stacks auto-start via compose `restart: unless-stopped`.
**Notable:** no Proxmox hypervisor involved — recovery is pure bare-metal. File-level restic runs from inside the host. No vzdump coverage (it's not a VM). Backups via restic to rest-server-ana.
### SureFire hosts (sfsrv-ana, sf-ana-container, sf-r630)
**Blast radius:** tenant workloads (scoped to SureFire client).
**Recovery:** coordinate with tenant. PFI responsibility is hardware + OS layer; application recovery may need tenant input. See `servers/sfsrv-ana/README.md` for hosting agreement scope.
---
## Tier 5 — Networking
### FortiGate 60F (ESH gateway, 10.0.250.1)
**Blast radius:** ESH site WAN + inter-site VPN to ANA/NH3.
**Recovery:** physical console access; restore config from FortiManager if needed.
### UniFi controllers (ESH 10.0.0.1, PFI 10.100.0.1)
**Blast radius:** UniFi AP/switch management (existing config persists on devices; only changes need the controller).
**Recovery:** device power cycle; controller reboot. Not critical for ongoing operations.
---
## Cross-cutting: what to check FIRST for ambiguous outages
When symptoms are vague ("lots of things are down"), run this triage in order:
```bash
# 1. Is the Anaheim NAS alive?
ping -c 2 10.250.50.50
# 2. Are the hypervisors alive?
for h in 10.250.250.31 10.100.250.60 10.0.250.35 10.0.50.55 10.250.250.115; do
ping -c 1 -W 2 $h >/dev/null && echo "$h OK" || echo "$h DOWN"
done
# 3. WAN reachability between sites?
# (from nh3-dev to ana-side IP; from ana-side to nh3-side IP)
ping -c 2 10.250.50.70 # ana-docker from NH3
```
The first `DOWN` in the hypervisor list narrows blast radius to that site / that hypervisor's guests.
---
## What this runbook does NOT cover
- Application-level restores from PBS / restic snapshots — that's per-service and lives in respective stack READMEs
- Full site-failover (e.g., "move PFI fleet to NH3 entirely") — requires Phase-9-style DR plan, not built
- Data-level integrity recovery (DB corruption, ZFS pool damage) — beyond this runbook's scope
**Maintenance note:** update this file when:
- New host gets registered (add to applicable tier)
- New SPOF discovered (add blast-radius note)
- A recovery procedure changes in practice
+195 -18
View File
@@ -197,20 +197,77 @@ both have VM 100 for instance). Without namespaces, their backups
collide under the same `/vm/100/` path in the datastore. Create a
namespace per hypervisor up front:
`proxmox-backup-manager` doesn't manage namespaces — they're created
through the API/web UI or the `proxmox-backup-client` tool.
**Easiest: web UI.** Datastore → `backups` → Content → top of pane
there's a namespace selector with an **Add NS** button. Add one per
hypervisor.
**Scripted via `proxmox-backup-client`** (run on PBS-ANA):
```bash
for h in pfi-pve nh3-pve esh-pve esh-pve-nas sfsrv-ana; do
proxmox-backup-manager namespace create backups $h
export PBS_REPOSITORY='root@pam@localhost:backups'
export PBS_PASSWORD='<root-pam-password>'
for ns in pfi-pve nh3-pve esh-pve esh-pve-nas sfsrv-ana; do
proxmox-backup-client namespace create "$ns"
done
proxmox-backup-manager namespace list backups
proxmox-backup-client namespace list
```
(If the CLI errors — subcommand names shift between PBS versions —
use the web UI: Datastore → backups → Content → **Add NS**. Works
reliably regardless of version.)
**Scripted via API** (when ssh access is more convenient than shell
on PBS-ANA):
```bash
TOKEN='<fleet-vzdump-secret>'
for ns in pfi-pve nh3-pve esh-pve esh-pve-nas sfsrv-ana; do
curl -sk \
-H "Authorization: PBSAPIToken=root@pam!fleet-vzdump:$TOKEN" \
-X POST \
https://10.250.50.90:8007/api2/json/admin/datastore/backups/namespace \
-d "{\"ns\":\"$ns\"}"
done
```
Each PVE client later (Phase 2.1, 3.1, 4) sets its Namespace field to
its own hostname when configuring the PBS storage.
### 1.4c. Create a verify job
New snapshots land unverified — PBS treats "backup completed" and
"backup verified" as separate states. A verify job hash-checks chunk
data (not just the manifest), flips snapshots to verified, and catches
bitrot on aging data.
Web UI: **Datastore → backups → Verify Jobs → Add**:
| Field | Value |
|---|---|
| Schedule | `sat 23:00` |
| Ignore verified snapshots | ✓ |
| Re-verify after (days) | `30` |
| Max depth | blank (unlimited) |
| Namespace | blank (root — recurses into all namespaces) |
| Comment | `fleet verify — new + 30d re-check` |
Rationale for these defaults:
- **`sat 23:00`** — sits between the daily 03:00 backup window and the
Sunday 06:00 GC run, so verify never fights GC for I/O, and every
week's new backups get verified before pruning decisions happen.
- **`ignore-verified: true`** + **`outdated-after: 30d`** — efficient
steady state. First run after a backup night does the new snapshots
only; a 30-day rolling re-verify catches silent chunk corruption.
- **unlimited depth, root namespace** — one job covers all 5
hypervisor namespaces. Split into per-namespace jobs only if you
want per-hypervisor visibility into verify failures (not necessary
for a fleet this size).
Verify runs are I/O-heavy on the datastore — on the NFS-backed
PBS-ANA, expect a full-datastore verify to take hours once the
datastore grows. The `ignore-verified` flag keeps incremental verify
cheap; only the 30-day-aged portion is re-read each run.
### 1.5. Create an API token for hypervisors to use
PBS web UI: **Configuration → Access Control → API Token → Add**:
@@ -405,32 +462,152 @@ Same pattern as Phase 1.1:
Same as Phase 1.2.
### 5.3. Local datastore
### 5.3. Datastore backing — Synology NFS (chosen 2026-04-22)
PBS-NH3 uses local storage rather than NFS (keeps the two PBSes on
different failure domains). Two options:
PBS-NH3 mounts a Synology NFS share rather than using a local virtual
disk. Simpler storage admin; tradeoff is that PBS-NH3 now shares its
failure domain with the other NH3 backup paths (nh3-docker restic,
nh3-dev restic, Backrest repos). Acceptable for DR purposes because
PBS-ANA remains primary.
**Option A:** attach a second virtual disk to the VM (e.g. 2 TB on
nh3-pve's `local-lvm` or whatever storage is available), format ext4,
mount at `/mnt/pbs-datastore`.
**Synology-side export setup (DSM):**
**Option B:** mount Synology Btrfs share via NFS or CIFS. Simpler
storage admin but couples PBS-NH3 to the Synology's availability.
1. Create a dedicated share, e.g. `pbs-nh3`, on the target volume.
2. **Control Panel → Shared Folder → [share] → Edit → NFS Permissions → Add/Edit:**
| Field | Value |
|---|---|
| Hostname/IP | PBS-NH3 VM IP |
| Privilege | Read/Write |
| Squash | **No mapping** (Synology's label for `no_root_squash`) |
| Security | sys |
| Enable asynchronous | on |
| Allow connections from non-privileged ports | on |
Recommended: **Option A**. Keeps it self-contained.
3. Verify no **Advanced Permissions** ACL denies `root` write — those
override NFS perms and cause silent write failures.
**Client-side mount (on PBS-NH3 VM):**
Use **NFSv3**, not NFSv4 — see the gotcha below. Synology's "Advanced
Permissions" layer an NFSv4 ACL on top of POSIX mode that is invisible
to `ls` but denies writes to unprivileged users (including the
`backup` uid-34 that PBS runs as), even when the directory mode is
777. Root bypasses this via `no_root_squash`, which is why a root
`touch` succeeds but the datastore init fails.
```bash
# Inside the PBS-NH3 VM, after attaching a disk /dev/sdb:
mkfs.ext4 /dev/sdb
mkdir -p /mnt/pbs-datastore
echo '/dev/sdb /mnt/pbs-datastore ext4 defaults 0 2' >> /etc/fstab
cat >> /etc/fstab <<'EOF'
10.100.50.50:/volume1/pbs-nh3 /mnt/pbs-datastore nfs defaults,_netdev,bg,hard,timeo=600,retrans=2,vers=3 0 0
EOF
mount -a
df -h /mnt/pbs-datastore
# Two-layer smoke test: root AND the backup user that PBS runs as.
# The backup-user check is the one that actually matters — if root
# works but backup doesn't, you're hitting the ACL override.
touch /mnt/pbs-datastore/root-write-test && rm /mnt/pbs-datastore/root-write-test
sudo -u backup touch /mnt/pbs-datastore/backup-user-test && rm /mnt/pbs-datastore/backup-user-test
```
Ensure NFSv3 is reachable on the Synology side: **Control Panel →
File Services → NFS → Advanced** — confirm "Minimum NFS Protocol" is
3 (not 4.0+). Default is usually 3; only an issue if someone
hardened it previously.
**Fallback if NFSv3 still fails the backup-user check:** the deny is
a Synology-local syno_acl (`ls -la` on the Synology shows POSIX mode
`d---------+` with a `+` for extended ACL). DSM "Enable Advanced
Permissions" unchecked does NOT reset this once the archive flag
`has_ACL,is_support_ACL` is set on the share.
Diagnose from the Synology shell:
```
sudo synoacltool -get /volume1/<share>
sudo ls -la /volume1/<share>/
```
A `+` after the permission string + a `group:administrators:allow:...`
ACL entry + POSIX `d---------` is the smoking gun: only members of
the `administrators` group have access, which is why root (with
`no_root_squash`) writes but uid-34 (backup) doesn't.
**Fix (confirmed working 2026-04-22):** keep **Squash: `No mapping`**
(i.e. `no_root_squash`) AND flatten the share to Linux/POSIX mode
with `chmod 777`. This drops the syno_acl entirely — verifiable by
`synoacltool -get` returning "It's Linux mode" and `ls -la`
showing `drwxrwxrwx` with NO trailing `+`.
On the Synology shell:
```bash
sudo chmod 777 /volume1/<share>
# Verify pure POSIX, no ACL
sudo synoacltool -get /volume1/<share> # should say "It's Linux mode"
ls -la /volume1/<share>/ # should show drwxrwxrwx (no '+')
```
Pure POSIX 777 is a cleaner long-term config than the ACL-grant
approach — fewer permission-translation layers between NFSv3 and
Btrfs, and no chance of ACL inheritance surprises on PBS-created
subdirectories.
**Why `all_squash` alone doesn't work:** it lets backup-user writes
through (because the ACL-granted admin gets the mapped uid), but
breaks PBS's `chown()` during init. Squashed-admin doesn't have
CAP_CHOWN on the Synology side → EPERM. Only real root (via
`no_root_squash`) can chown to uid 34.
**Why squash alone fails:** even with `all_squash + anonuid=1024`
letting backup-user writes succeed (because admin is in
`administrators` ACL), PBS's datastore-init calls `chown` on
newly-created paths. Squashed-admin doesn't have CAP_CHOWN on the
Synology side → EPERM. Only a real root (via `no_root_squash`) can
chown to uid 34.
**Smoke-test all three paths after the fix:**
```bash
sudo touch /mnt/pbs-datastore/t && rm /mnt/pbs-datastore/t && echo ROOT_OK
sudo -u backup touch /mnt/pbs-datastore/t && rm /mnt/pbs-datastore/t && echo BACKUP_OK
touch /mnt/pbs-datastore/t && chown 34:34 /mnt/pbs-datastore/t && rm /mnt/pbs-datastore/t && echo CHOWN_OK
```
All three must pass before PBS will init cleanly.
**Historical note:** both PBS instances ended up on NFSv3 for
unrelated reasons — the ANA side because of ZFS case-insensitivity,
the NH3 side because of Synology ACL override. NFSv3 is the safer
default for PBS-on-NFS regardless of backend.
### 5.4. Create datastore
Web UI: Datastore → Add, name `backups-mirror`, path `/mnt/pbs-datastore`.
### 5.5. Create a verify job on the mirror
Same pattern as Phase 1.4c on PBS-ANA — verification doesn't replicate
across PBS instances, so the mirror needs its own job to catch bitrot
on the local datastore disk.
Web UI: **Datastore → backups-mirror → Verify Jobs → Add**:
| Field | Value |
|---|---|
| Schedule | `sun 12:00` |
| Ignore verified snapshots | ✓ |
| Re-verify after (days) | `30` |
| Max depth | blank (unlimited) |
| Namespace | blank |
| Comment | `mirror verify — catches DR-side bitrot` |
Schedule sits after the 06:00 sync job finishes, so newly-synced
snapshots get verified same day.
### Phase 5 done-state
- PBS-NH3 reachable, datastore ready