ops(nh3-dev): seat healthz watchdog — the 5-day pull-only outage must page, not age

althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so
Restart=on-failure never revived it; hermes-gateway sat pull-only for five
days with a yellow verdict nobody was watching. Fixed the death mode with a
Restart=always drop-in (also on the jekyll twin, same latent bug) and added
this watchdog for every other death mode: 5-min timer, alert on the second
consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery
mail on green, exit 2 distinguishes a broken alarm wire from a down seat
(backup-freshness precedent). Lifecycle exercised end-to-end against a dead
port before install.
This commit is contained in:
vh
2026-10-02 07:47:17 -07:00
parent 902e16630f
commit 59fb6b77a8
2 changed files with 131 additions and 0 deletions
+38
View File
@@ -0,0 +1,38 @@
#!/usr/bin/env bash
# install-seat-healthz-timer.sh — install/refresh the seat-healthz watchdog as
# a systemd USER timer on nh3-dev (5-min ticks; alert after 2 consecutive
# failing ticks ~= 10 min sustained). Idempotent; re-run after editing the
# wrapper. Pattern follows install-backup-freshness-timer.sh.
set -euo pipefail
UNIT_DIR="$HOME/.config/systemd/user"
REPO=/home/lkraven/development/eshpfi-management
mkdir -p "$UNIT_DIR"
cat > "$UNIT_DIR/seat-healthz.service" <<EOF
[Unit]
Description=Seat status-page healthz watchdog + althing alert
After=network-online.target
[Service]
Type=oneshot
Environment=PATH=/home/lkraven/.local/bin:/usr/local/bin:/usr/bin:/bin
Environment=ALTHING_POST_OFFICE=http://10.100.50.40:8390
ExecStart=$REPO/scripts/seat-healthz-alert.sh
EOF
cat > "$UNIT_DIR/seat-healthz.timer" <<EOF
[Unit]
Description=5-min tick for the seat healthz watchdog
[Timer]
OnBootSec=2min
OnUnitActiveSec=5min
RandomizedDelaySec=30
[Install]
WantedBy=timers.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now seat-healthz.timer
systemctl --user list-timers seat-healthz.timer --no-pager | tail -2