Files
esh-pfi-infrastructure/configs/restic/esh-vm-db
vh 6e203dcb99 fix(restic): stop publishing rest-server passwords in systemd units
resticprofile schedule copies env-file values into the generated units, which
are 0644, so RESTIC_REPOSITORY (the rest-server basic-auth password included)
was readable by every local user on every restic host.

New playbooks/restic-repository-file.yaml:
- derives /etc/restic/repository (root 0400) from restic.env;
- uploads the profile switched to repository-file, but only when the live
  profile's sha matches the repo copy it was edited from (drift guard);
- checks the repository is reachable through the new profile (cat config);
- regenerates the units and verifies they exist and contain no rest:http.

Applied to ana-docker, fv-ml1 (configs/restic/ana-ml2), esh-docker-vm,
esh-vm-db, irv-ml1, nh3-dev and nh3-docker. An independent check across all
eight restic hosts (these seven plus esh-ml1) found 0 leaking units. nh3-docker's
scheduled unit ran a real backup afterwards (snapshot a29b889d). Each host's
URL and passphrase are vaulted as <host>/etc/restic/{repository,password}.

restic.env is kept (root 0600) because the per-host READMEs and the freshness
probe source it. A rotation must update the vault, restic.env and repository.

vm-esh-nas has no infra-ops account. Its in-place migration script is staged
for Prime to run with sudo, and its repo profile is pre-edited to match.

Also mirrors augaman-dev's 401df2d (compose header only). The config hash on
esh-ml1 is unchanged.
2026-09-27 01:51:05 -07:00
..

restic / esh-vm-db

Two-database host at the ESH site (PostgreSQL 15 + MongoDB). Covered at the VM-image layer by PBS-ANA via esh-pve (or whichever ESH hypervisor owns this VM — confirm on next inventory pass). This restic profile adds DB-level granularity via pre-backup dumps.

What's backed up

Path Purpose
/etc Host config — systemd, ssh, chrony, apt, pg_hba.conf, mongod.conf
/root Root's ad-hoc scripts, shell history, ssh keys
/home User homes (lkraven + any DB-admin locals)
/var/lib/restic/stage pg_dumpall.sql.gz + mongodump/ produced by pre-backup.sh

What's not backed up (by design)

  • /var/lib/postgresql — raw PGDATA. Live-capture risk; pg_dumpall in pre-backup covers it consistently.
  • /var/lib/mongodb — raw mongo dbPath. Same reasoning; mongodump covers it.
  • NFS mount /mnt/backup (from esh-nas — not ours to mirror).

Pre-backup hook

pre-backup.sh runs as root before restic. It:

  1. Checks pg_isready on :5432 — if OK, runs pg_dumpall piped through gzip to $STAGE/pg_dumpall.sql.gz
  2. Checks mongo ping via mongosh — if OK, runs mongodump into $STAGE/mongodump/

PostgreSQL uses the local Unix socket and peer authentication as postgres, not TCP localhost. Dumps are staged before replacement. If either DB dump fails, the hook returns nonzero and aborts the backup, preserving that DB's previous dump. The two databases are not a single transactional snapshot. RESTIC_STAGE_DIR supports isolated regression tests.

Repair verified 2026-09-12

Weekly check failed September 6 on a repository connection timeout after boot. Nightly backup returned success despite PostgreSQL TCP authentication failures, reusing a dump last modified April 23. Fixed socket authentication, propagated both DB failures, and added network-online ordering plus bounded retries to both services (5-minute delay, maximum 3 starts per hour).

Deploy with playbooks/esh-vm-db-restic-repair.yaml. Service drop-ins survive regeneration of resticprofile's main units. Previous hook and PG dump retained under /var/lib/restic/repair-20260912/ (root-only).

Fresh snapshot bc5eeaff at 07:01 PDT contains today's 3,460,215-byte compressed PG dump. Retrieved FROM repository, decompressed successfully, and verified its cluster-dump completion marker. This is not a full database restore test. Mongo dump also completed. Weekly check rerun at 07:02 passed: 99 snapshots, configured 10% data sample (19 packs), no errors. Both jobs Result=success, no failed systemd units, timers active, PostgreSQL/MongoDB remain active. Inactive/dead between scheduled runs is normal for these finite jobs.

Tests: python3 configs/restic/esh-vm-db/test_pre_backup.py — three passing regressions for successful peer-auth dump and preservation/failure propagation for each database. No DB authentication policy or service restarts changed.

Deploy (one-time)

1. Create rest-server-ana htpasswd entry

Do NOT use sudo for .htpasswd writes on ana-docker. The file is NFS-mounted from ana-nas and owned by uid 1000 (the rest-server user, which equals lkraven). Sudo-root on the client gets squashed to nobody on the NFS server and can't read/write the file. lkraven writes it natively, using the docker group for the bcrypt helper.

# Pick password in password manager first
HTPW='<new-pw-saved-to-pw-manager>'

ssh -t ana-docker "docker run --rm httpd:2.4-alpine htpasswd -nbB esh-vm-db '$HTPW' | \
                   tee /tmp/htline.txt > /dev/null && \
                   sed -i '/^esh-vm-db:/d' /mnt/backup/restic/repo/ana/.htpasswd && \
                   cat /tmp/htline.txt >> /mnt/backup/restic/repo/ana/.htpasswd && \
                   rm /tmp/htline.txt && \
                   grep ^esh-vm-db: /mnt/backup/restic/repo/ana/.htpasswd && \
                   docker restart rest-server"
unset HTPW

2. Install secrets on esh-vm-db

ssh -t esh-vm-db 'sudo install -d -o root -g root -m 0700 /etc/restic /var/lib/restic /var/lib/restic/stage'

# restic.env — URL-encode the password if it has special chars
ssh -t esh-vm-db "sudo bash -c '
  read -sp \"htpasswd pw for rest-server-ana: \" HTPW; echo
  cat > /etc/restic/restic.env <<EOF
RESTIC_REPOSITORY=rest:http://esh-vm-db:\$HTPW@10.250.50.70:8000/esh-vm-db/
EOF
  chmod 600 /etc/restic/restic.env
'"

# Repo passphrase (prints once — save to password manager)
ssh -t esh-vm-db 'sudo bash -c "
  openssl rand -base64 48 | tr -d \"\\n\" > /etc/restic/password
  chmod 600 /etc/restic/password
  echo === SAVE THIS TO PASSWORD MANAGER NOW ===
  cat /etc/restic/password
  echo
"'

3. Initialize the repo

ssh -t esh-vm-db 'sudo bash -c "
  set -a; . /etc/restic/restic.env; set +a
  RESTIC_PASSWORD_FILE=/etc/restic/password restic init
"'

4. Install prerequisites (restic, resticprofile, mongosh client)

ssh -t esh-vm-db 'which restic || sudo apt-get install -y restic; \
                  which mongosh || echo "NOTE: mongosh not found; pre-backup mongo ping will fail safely — install via MongoDB APT repo if needed"; \
                  curl -sfL https://raw.githubusercontent.com/creativeprojects/resticprofile/master/install.sh | sudo sh -s -- -b /usr/local/bin; \
                  /usr/local/bin/resticprofile --version'

5. Deploy profile + hook

scp configs/restic/esh-vm-db/profiles.yaml esh-vm-db:/tmp/
scp configs/restic/esh-vm-db/pre-backup.sh esh-vm-db:/tmp/

ssh -t esh-vm-db 'sudo install -o root -g root -m 0644 /tmp/profiles.yaml /etc/restic/profiles.yaml && \
                  sudo install -o root -g root -m 0755 /tmp/pre-backup.sh /etc/restic/pre-backup.sh && \
                  rm /tmp/profiles.yaml /tmp/pre-backup.sh'

6. Schedule + verify

ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml schedule --all && \
                  systemctl list-timers "resticprofile*" --no-pager'

# First manual run
ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml backup --verbose'

Expect first run to land ~50-200 MB (mostly the mongodump directory + pg_dumpall). Cross-check from Backrest UI on ana-docker.

Restore

Full host config

ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /etc'

Just the PG dump

ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /var/lib/restic/stage/pg_dumpall.sql.gz'
# Then: gunzip + psql < pg_dumpall.sql

Just a mongo DB

ssh -t esh-vm-db 'sudo resticprofile --config /etc/restic/profiles.yaml restore latest --target /tmp/restore --path /var/lib/restic/stage/mongodump'
# Then: mongorestore /tmp/restore/var/lib/restic/stage/mongodump/

Gotchas

  • mongosh must be installed or the mongo pre-backup step silently skips (logged as WARN). Install from the MongoDB APT repo if not already present — the stock Debian mongodb-clients package is out of date and doesn't include mongosh.
  • Mongo authentication — if mongod ever gets auth enabled (it's currently open to 0.0.0.0 with no auth, which is its own concern), mongodump will need --username/--password flags. Reference in pre-backup.sh when that change happens.
  • pg_hba.conf — pg_dumpall requires local postgres superuser access. Currently works via sudo -u postgres + peer auth on the local socket. If pg_hba.conf ever changes peer → md5 for local, the hook needs a ~postgres/.pgpass entry.