Work by a parallel session on 2026-09-12; committed here with the rest of the day's changes. Rationale in persistent-memory.d/2026-09-12-esh-vm-db-restic-repair.md. The only visible symptom was a systemd-failed unit from a Sep 6 repository network timeout after boot. The real fault was quieter and much worse: the pre-backup hook logged failures as WARN and returned zero, so pg_dumpall could fail every single night -- it used TCP localhost and wanted a password nobody supplied -- while restic dutifully backed up the stale April 23 dump still sitting in the staging directory and reported success. Mongo was fine, which is part of why it went unnoticed. Postgres now dumps over the /var/run/postgresql socket with peer auth and -w, and both database failures now fail the backup rather than masking it, while still preserving any prior per-DB dump rather than truncating to nothing. An ERRORS counter replaces the warn-and-continue path, and the staging directory is overridable via RESTIC_STAGE_DIR so the new test can exercise it. Adds backup.contract.md, retry.conf and test_pre_backup.py -- three red-green regression tests covering the failure modes above. systemd drop-ins on both jobs add network-online ordering plus Restart=on-failure with a 5m delay and a 3-per-hour limit, which addresses the original boot-timeout symptom. Verified against a real run: snapshot bc5eeaff at 07:01 PDT with a fresh 3.46MB PG dump, retrieved from the repository with decompression and completion marker checked (not a full restore). Repository check passed, 99 snapshots. The old hook and stale dump are preserved root-only at /var/lib/restic/repair-20260912.
7.5 KiB
restic / esh-vm-db
Two-database host at the ESH site (PostgreSQL 15 + MongoDB). Covered at the VM-image layer by PBS-ANA via esh-pve (or whichever ESH hypervisor owns this VM — confirm on next inventory pass). This restic profile adds DB-level granularity via pre-backup dumps.
What's backed up
| Path | Purpose |
|---|---|
/etc |
Host config — systemd, ssh, chrony, apt, pg_hba.conf, mongod.conf |
/root |
Root's ad-hoc scripts, shell history, ssh keys |
/home |
User homes (lkraven + any DB-admin locals) |
/var/lib/restic/stage |
pg_dumpall.sql.gz + mongodump/ produced by pre-backup.sh |
What's not backed up (by design)
/var/lib/postgresql— raw PGDATA. Live-capture risk;pg_dumpallin pre-backup covers it consistently./var/lib/mongodb— raw mongo dbPath. Same reasoning;mongodumpcovers it.- NFS mount
/mnt/backup(from esh-nas — not ours to mirror).
Pre-backup hook
pre-backup.sh runs as root before restic. It:
- Checks
pg_isreadyon :5432 — if OK, runspg_dumpallpiped through gzip to$STAGE/pg_dumpall.sql.gz - Checks mongo ping via
mongosh— if OK, runsmongodumpinto$STAGE/mongodump/
PostgreSQL uses the local Unix socket and peer authentication as postgres,
not TCP localhost. Dumps are staged before replacement. If either DB dump
fails, the hook returns nonzero and aborts the backup, preserving that DB's
previous dump. The two databases are not a single transactional snapshot.
RESTIC_STAGE_DIR supports isolated regression tests.
Repair verified 2026-09-12
Weekly check failed September 6 on a repository connection timeout after boot. Nightly backup returned success despite PostgreSQL TCP authentication failures, reusing a dump last modified April 23. Fixed socket authentication, propagated both DB failures, and added network-online ordering plus bounded retries to both services (5-minute delay, maximum 3 starts per hour).
Deploy with playbooks/esh-vm-db-restic-repair.yaml. Service drop-ins survive
regeneration of resticprofile's main units. Previous hook and PG dump retained
under /var/lib/restic/repair-20260912/ (root-only).
Fresh snapshot bc5eeaff at 07:01 PDT contains today's 3,460,215-byte compressed
PG dump. Retrieved FROM repository, decompressed successfully, and verified its
cluster-dump completion marker. This is not a full database restore test.
Mongo dump also completed. Weekly check rerun at 07:02 passed: 99 snapshots,
configured 10% data sample (19 packs), no errors. Both jobs Result=success,
no failed systemd units, timers active, PostgreSQL/MongoDB remain active.
Inactive/dead between scheduled runs is normal for these finite jobs.
Tests: python3 configs/restic/esh-vm-db/test_pre_backup.py — three passing
regressions for successful peer-auth dump and preservation/failure propagation
for each database. No DB authentication policy or service restarts changed.
Deploy (one-time)
1. Create rest-server-ana htpasswd entry
Do NOT use sudo for .htpasswd writes on ana-docker. The file is
NFS-mounted from ana-nas and owned by uid 1000 (the rest-server user,
which equals lkraven). Sudo-root on the client gets squashed to
nobody on the NFS server and can't read/write the file. lkraven
writes it natively, using the docker group for the bcrypt helper.
# Pick password in password manager first
HTPW='<new-pw-saved-to-pw-manager>'
ssh -t ana-docker "docker run --rm httpd:2.4-alpine htpasswd -nbB esh-vm-db '$HTPW' | \
tee /tmp/htline.txt > /dev/null && \
sed -i '/^esh-vm-db:/d' /mnt/backup/restic/repo/ana/.htpasswd && \
cat /tmp/htline.txt >> /mnt/backup/restic/repo/ana/.htpasswd && \
rm /tmp/htline.txt && \
grep ^esh-vm-db: /mnt/backup/restic/repo/ana/.htpasswd && \
docker restart rest-server"
unset HTPW
2. Install secrets on esh-vm-db
ssh -t esh-vm-db 'sudo install -d -o root -g root -m 0700 /etc/restic /var/lib/restic /var/lib/restic/stage'
# restic.env — URL-encode the password if it has special chars
ssh -t esh-vm-db "sudo bash -c '
read -sp \"htpasswd pw for rest-server-ana: \" HTPW; echo
cat > /etc/restic/restic.env <<EOF
RESTIC_REPOSITORY=rest:http://esh-vm-db:\$HTPW@10.250.50.70:8000/esh-vm-db/
EOF
chmod 600 /etc/restic/restic.env
'"
# Repo passphrase (prints once — save to password manager)
ssh -t esh-vm-db 'sudo bash -c "
openssl rand -base64 48 | tr -d \"\\n\" > /etc/restic/password
chmod 600 /etc/restic/password
echo === SAVE THIS TO PASSWORD MANAGER NOW ===
cat /etc/restic/password
echo
"'
3. Initialize the repo
ssh -t esh-vm-db 'sudo bash -c "
set -a; . /etc/restic/restic.env; set +a
RESTIC_PASSWORD_FILE=/etc/restic/password restic init
"'
4. Install prerequisites (restic, resticprofile, mongosh client)
ssh -t esh-vm-db 'which restic || sudo apt-get install -y restic; \
which mongosh || echo "NOTE: mongosh not found; pre-backup mongo ping will fail safely — install via MongoDB APT repo if needed"; \
curl -sfL https://raw.githubusercontent.com/creativeprojects/resticprofile/master/install.sh | sudo sh -s -- -b /usr/local/bin; \
/usr/local/bin/resticprofile --version'
5. Deploy profile + hook
scp configs/restic/esh-vm-db/profiles.yaml esh-vm-db:/tmp/
scp configs/restic/esh-vm-db/pre-backup.sh esh-vm-db:/tmp/
ssh -t esh-vm-db 'sudo install -o root -g root -m 0644 /tmp/profiles.yaml /etc/restic/profiles.yaml && \
sudo install -o root -g root -m 0755 /tmp/pre-backup.sh /etc/restic/pre-backup.sh && \
rm /tmp/profiles.yaml /tmp/pre-backup.sh'
6. Schedule + verify
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml schedule --all && \
systemctl list-timers "resticprofile*" --no-pager'
# First manual run
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml backup --verbose'
Expect first run to land ~50-200 MB (mostly the mongodump directory + pg_dumpall). Cross-check from Backrest UI on ana-docker.
Restore
Full host config
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /etc'
Just the PG dump
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /var/lib/restic/stage/pg_dumpall.sql.gz'
# Then: gunzip + psql < pg_dumpall.sql
Just a mongo DB
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /var/lib/restic/stage/mongodump'
# Then: mongorestore /tmp/restore/var/lib/restic/stage/mongodump/
Gotchas
- mongosh must be installed or the mongo pre-backup step silently
skips (logged as WARN). Install from the MongoDB APT repo if not
already present — the stock Debian
mongodb-clientspackage is out of date and doesn't includemongosh. - Mongo authentication — if mongod ever gets auth enabled (it's
currently open to 0.0.0.0 with no auth, which is its own concern),
mongodumpwill need--username/--passwordflags. Reference in pre-backup.sh when that change happens. - pg_hba.conf —
pg_dumpallrequires local postgres superuser access. Currently works viasudo -u postgres+ peer auth on the local socket. Ifpg_hba.confever changes peer → md5 for local, the hook needs a~postgres/.pgpassentry.