Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md
T

1.4 KiB

[2026-09-19] The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking.

⭐⭐⭐ The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking. (1) backup-freshness-alert.sh called althing-cli, DELETED by the 2026-08-28 v3 cutover — the check detected stale backups every morning and told nobody. Swapped to postbox, added the missing ALTHING_POST_OFFICE to the unit and its installer (the binary swap alone would have failed differently), recipient infra-ops→infra-hermes (it was mailing itself), exit 2 now means the ALARM is broken vs exit 1 the backups, and --test-alert is a positive control because nothing had ever exercised the healthy path. (2) The checker is now coverage-aware — a guest is "not backed up by policy" only when NO enabled vzdump job covers it, a UNION across jobs: reading one exclude list would have silently stopped alarming on ana CT 109, which is excluded from the 03:00 job AND has its own 22:00 job. (3) It now reads vzdump TASK STATUS, because snapshot age is structurally blind to a job that runs and errors nightly — that cost 6 days on both CT 107 and VM 102, and it immediately found a third case on esh-nas-pve whose guests all read 0-1h FRESH. ⚠ ops-log earned its keep on day one: it captured all three failed playbook attempts automatically. Commits e979ccb, 5be25be, ba26852.