Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md
T

4 lines
1.4 KiB
Markdown

# `[2026-09-19]` The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking.
⭐⭐⭐ **The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking.** (1) `backup-freshness-alert.sh` called `althing-cli`, DELETED by the 2026-08-28 v3 cutover — the check detected stale backups every morning and told nobody. Swapped to `postbox`, added the missing `ALTHING_POST_OFFICE` to the unit **and its installer** (the binary swap alone would have failed differently), recipient `infra-ops`→`infra-hermes` (it was mailing itself), exit 2 now means the ALARM is broken vs exit 1 the backups, and `--test-alert` is a positive control because nothing had ever exercised the healthy path. (2) The checker is now **coverage-aware** — a guest is "not backed up by policy" only when NO enabled vzdump job covers it, a UNION across jobs: reading one `exclude` list would have silently stopped alarming on ana CT 109, which is excluded from the 03:00 job AND has its own 22:00 job. (3) It now reads **vzdump TASK STATUS**, because snapshot age is structurally blind to a job that runs and errors nightly — that cost 6 days on both CT 107 and VM 102, and it immediately found a third case on esh-nas-pve whose guests all read 0-1h FRESH. ⚠ **`ops-log` earned its keep on day one**: it captured all three failed playbook attempts automatically. Commits e979ccb, 5be25be, ba26852.