althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so Restart=on-failure never revived it; hermes-gateway sat pull-only for five days with a yellow verdict nobody was watching. Fixed the death mode with a Restart=always drop-in (also on the jekyll twin, same latent bug) and added this watchdog for every other death mode: 5-min timer, alert on the second consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery mail on green, exit 2 distinguishes a broken alarm wire from a down seat (backup-freshness precedent). Lifecycle exercised end-to-end against a dead port before install.
39 lines
1.1 KiB
Bash
Executable File
39 lines
1.1 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# install-seat-healthz-timer.sh — install/refresh the seat-healthz watchdog as
|
|
# a systemd USER timer on nh3-dev (5-min ticks; alert after 2 consecutive
|
|
# failing ticks ~= 10 min sustained). Idempotent; re-run after editing the
|
|
# wrapper. Pattern follows install-backup-freshness-timer.sh.
|
|
set -euo pipefail
|
|
UNIT_DIR="$HOME/.config/systemd/user"
|
|
REPO=/home/lkraven/development/eshpfi-management
|
|
mkdir -p "$UNIT_DIR"
|
|
|
|
cat > "$UNIT_DIR/seat-healthz.service" <<EOF
|
|
[Unit]
|
|
Description=Seat status-page healthz watchdog + althing alert
|
|
After=network-online.target
|
|
|
|
[Service]
|
|
Type=oneshot
|
|
Environment=PATH=/home/lkraven/.local/bin:/usr/local/bin:/usr/bin:/bin
|
|
Environment=ALTHING_POST_OFFICE=http://10.100.50.40:8390
|
|
ExecStart=$REPO/scripts/seat-healthz-alert.sh
|
|
EOF
|
|
|
|
cat > "$UNIT_DIR/seat-healthz.timer" <<EOF
|
|
[Unit]
|
|
Description=5-min tick for the seat healthz watchdog
|
|
|
|
[Timer]
|
|
OnBootSec=2min
|
|
OnUnitActiveSec=5min
|
|
RandomizedDelaySec=30
|
|
|
|
[Install]
|
|
WantedBy=timers.target
|
|
EOF
|
|
|
|
systemctl --user daemon-reload
|
|
systemctl --user enable --now seat-healthz.timer
|
|
systemctl --user list-timers seat-healthz.timer --no-pager | tail -2
|