restic(esh-vm-db): a backup that exited 0 nightly while keeping an April dump

Work by a parallel session on 2026-09-12; committed here with the rest of the
day's changes. Rationale in persistent-memory.d/2026-09-12-esh-vm-db-restic-repair.md.

The only visible symptom was a systemd-failed unit from a Sep 6 repository
network timeout after boot. The real fault was quieter and much worse: the
pre-backup hook logged failures as WARN and returned zero, so pg_dumpall could
fail every single night -- it used TCP localhost and wanted a password nobody
supplied -- while restic dutifully backed up the stale April 23 dump still
sitting in the staging directory and reported success. Mongo was fine, which
is part of why it went unnoticed.

Postgres now dumps over the /var/run/postgresql socket with peer auth and -w,
and both database failures now fail the backup rather than masking it, while
still preserving any prior per-DB dump rather than truncating to nothing. An
ERRORS counter replaces the warn-and-continue path, and the staging directory
is overridable via RESTIC_STAGE_DIR so the new test can exercise it.

Adds backup.contract.md, retry.conf and test_pre_backup.py -- three red-green
regression tests covering the failure modes above. systemd drop-ins on both
jobs add network-online ordering plus Restart=on-failure with a 5m delay and a
3-per-hour limit, which addresses the original boot-timeout symptom.

Verified against a real run: snapshot bc5eeaff at 07:01 PDT with a fresh 3.46MB
PG dump, retrieved from the repository with decompression and completion marker
checked (not a full restore). Repository check passed, 99 snapshots. The old
hook and stale dump are preserved root-only at /var/lib/restic/repair-20260912.
This commit is contained in:
vh
2026-09-12 22:12:24 -07:00
parent 3e727dbeb5
commit a5691ce796
6 changed files with 118 additions and 13 deletions
+29 -4
View File
@@ -31,10 +31,35 @@ profile adds DB-level granularity via pre-backup dumps.
2. Checks mongo ping via `mongosh` — if OK, runs `mongodump` into
`$STAGE/mongodump/`
Both dumps are atomic (write to `.tmp`, then rename). If either DB is
unreachable, the script logs a WARN and continues — a failed DB dump
doesn't abort the whole restic run, and restic falls back to whatever
stage content is left over from the prior successful dump.
PostgreSQL uses the local Unix socket and peer authentication as postgres,
not TCP localhost. Dumps are staged before replacement. If either DB dump
fails, the hook returns nonzero and aborts the backup, preserving that DB's
previous dump. The two databases are not a single transactional snapshot.
`RESTIC_STAGE_DIR` supports isolated regression tests.
## Repair verified 2026-09-12
Weekly check failed September 6 on a repository connection timeout after boot.
Nightly backup returned success despite PostgreSQL TCP authentication failures,
reusing a dump last modified April 23. Fixed socket authentication, propagated
both DB failures, and added network-online ordering plus bounded retries to
both services (5-minute delay, maximum 3 starts per hour).
Deploy with `playbooks/esh-vm-db-restic-repair.yaml`. Service drop-ins survive
regeneration of resticprofile's main units. Previous hook and PG dump retained
under `/var/lib/restic/repair-20260912/` (root-only).
Fresh snapshot `bc5eeaff` at 07:01 PDT contains today's 3,460,215-byte compressed
PG dump. Retrieved FROM repository, decompressed successfully, and verified its
cluster-dump completion marker. This is not a full database restore test.
Mongo dump also completed. Weekly check rerun at 07:02 passed: 99 snapshots,
configured 10% data sample (19 packs), no errors. Both jobs Result=success,
no failed systemd units, timers active, PostgreSQL/MongoDB remain active.
Inactive/dead between scheduled runs is normal for these finite jobs.
Tests: `python3 configs/restic/esh-vm-db/test_pre_backup.py` — three passing
regressions for successful peer-auth dump and preservation/failure propagation
for each database. No DB authentication policy or service restarts changed.
## Deploy (one-time)