fix(backups): stop saying STALE over a fleet whose every backup is fresh
The check collapsed two different findings into one verdict. On 2026-09-20 it printed "RESULT: STALE" while reporting 37 FRESH layers and zero stale ones -- every backup body provably current, the three ❌ rows all yesterday's pre-fix runs aging out of the 36h window. infra-hermes caught it in triage: a reader, or a forwarder, could page someone over a state where nothing is stale. STALE is a claim about backup AGE. A job that ran and errored is a different claim with different urgency. They now have different words and different exit codes: 0 all backups fresh 1 STALE -- a body past the threshold, or an endpoint down 3 ERRORED-JOBS -- every body fresh, a vzdump job errored recently The alert wrapper mirrors the code and matches its own wording to the finding: 🟡 "Backup jobs errored — all bodies fresh" instead of 🔴 "Backup freshness ALERT", and it now exits with the check's code rather than flattening everything to 1, so `systemctl status` distinguishes the states too. This is the same defect class the rest of this script was built to fix, one level up: not an instrument that fails to look, but one that looks correctly and then reports the wrong word for what it saw. An alarm that cries outage over a healthy fleet earns being ignored exactly as fast as one that stays silent over a broken one. Verified all three states by forcing each: BACKUP_JOB_WINDOW_HOURS=1 -> exit 0, default -> exit 3, BACKUP_MAX_AGE_HOURS=1 -> exit 1.
This commit is contained in:
@@ -26,10 +26,16 @@
|
||||
# infra-ops. ⚠ It must NOT be infra-ops — this runs AS infra-ops, so that
|
||||
# would mail the alarm to itself, which is the mirror trap named in CLAUDE.md.
|
||||
#
|
||||
# Exit codes:
|
||||
# Exit codes — mirrored from the check, so `systemctl status` distinguishes
|
||||
# the three states that matter:
|
||||
# 0 backups fresh
|
||||
# 1 backups STALE/DOWN, and the alert was delivered
|
||||
# 2 the ALERT PATH ITSELF FAILED (investigate this before the backups)
|
||||
# 3 ERRORED-JOBS — every body fresh, a job errored recently; alert delivered
|
||||
#
|
||||
# 3 is deliberately not 1. A nightly job that errored is worth a look; a stale
|
||||
# backup body is worth a page. Reporting both as "STALE" over a fleet whose
|
||||
# every backup is current is how an alarm teaches you to ignore it.
|
||||
set -uo pipefail
|
||||
|
||||
REPO=/home/lkraven/development/eshpfi-management
|
||||
@@ -72,14 +78,24 @@ out=$("$REPO/scripts/check-backup-freshness.sh" 2>&1); rc=$?
|
||||
printf '%s\n' "$out"
|
||||
[ "$rc" -eq 0 ] && exit 0
|
||||
|
||||
# Match the words and the emoji to the finding. A job-errors-only state is not
|
||||
# an outage and must not dress like one.
|
||||
if [ "$rc" -eq 3 ]; then
|
||||
subject="🟡 Backup jobs errored ($(date '+%Y-%m-%d')) — all bodies fresh"
|
||||
lead="A vzdump job errored recently. Every backup body on the fleet is FRESH; nothing is stale."
|
||||
else
|
||||
subject="🔴 Backup freshness ALERT ($(date '+%Y-%m-%d'))"
|
||||
lead="Automated daily backup-freshness check found STALE or DOWN backup layer(s) on the PFI fleet."
|
||||
fi
|
||||
|
||||
printf '%s\n' \
|
||||
"Automated daily backup-freshness check found STALE or DOWN backup layer(s) on the PFI fleet." \
|
||||
"$lead" \
|
||||
"Runbook: docs/runbooks/backups.md (topology, 2-min check, rest-server-ana recovery)." \
|
||||
"" \
|
||||
"$out" \
|
||||
| send_alert "🔴 Backup freshness ALERT ($(date '+%Y-%m-%d'))" || {
|
||||
echo "FATAL: backups are STALE *and* the alert post FAILED — nobody was told." >&2
|
||||
| send_alert "$subject" || {
|
||||
echo "FATAL: the check found problems *and* the alert post FAILED — nobody was told." >&2
|
||||
echo " Fix the alert path first; that is the fault that hides the others." >&2
|
||||
exit 2; }
|
||||
|
||||
exit 1
|
||||
exit "$rc"
|
||||
|
||||
Reference in New Issue
Block a user