docs(notify-failure): a stop that exits non-zero pages as a failure
A deliberate `systemctl --user restart hermes-gateway` paged infra-ops as a FAILED unit (msg 3716) while the unit was already back up. Cause is a Hermes v0.21.1 race: the planned-stop marker watcher runs the shutdown handler before systemd's SIGTERM, consumes the marker, and the SIGTERM re-runs the handler, which then classifies the stop as unexpected and exits 1. Corrects the README claim that OnFailure never fires on a deliberate restart: that holds only when the main process exits with a success status. Measured on a throwaway unit (3/3 paged without SuccessExitStatus, 0/3 with it, and crash restarts are unaffected), and a survey of every stop on nh3-dev since 09-15 found hermes-gateway to be the only unit that does this. The host-side fix is drop-in hermes-gateway.service.d/ 20-planned-stop-exit1-is-clean.conf (SuccessExitStatus=1). No alarm coverage is lost: Restart=always ignores the classification, and StartLimitIntervalSec=0 means the unit can never reach `failed` from a start failure.
This commit is contained in:
@@ -30,6 +30,18 @@ delivery, a blind spot in the notification path every other alarm depends on.
|
||||
`OnFailure` does not fire on a clean restart or a deliberate stop. svos-dev's
|
||||
four restarts and three deploys in one day would have produced **zero** alerts.
|
||||
|
||||
⚠ **"Clean" means the main process exits with a success status when it gets
|
||||
SIGTERM.** A daemon that exits non-zero on a *deliberate* `systemctl restart`
|
||||
or `stop` lands the unit in `failed` for an instant, and `OnFailure` fires even
|
||||
though the restart then carries on and succeeds. The page it sends reads
|
||||
`state active (running)` — that combination is the tell. Measured 2026-09-23 on
|
||||
a throwaway unit that exits 1 on SIGTERM: 3/3 restarts paged. Adding
|
||||
`SuccessExitStatus=<code>` to the unit brought it to 0/3, and an exit outside a
|
||||
stop job still auto-restarted under `Restart=always`. A survey of every stop on
|
||||
nh3-dev since 09-15 (booth 49, peedlar 30, svos 10, …) found **one** unit that
|
||||
does this — `hermes-gateway`, below — so the fix is per-unit, not in the
|
||||
notifier.
|
||||
|
||||
### Duplicate suppression, and a claim I had to correct
|
||||
|
||||
This README originally asserted that a crash-loop "yields one message per
|
||||
@@ -77,7 +89,7 @@ alarm would have fired.
|
||||
| **Covered** | the 7 timer-driven oneshots (dev-backup, ha-backup, fleet-tls-cert-check, headscale-ddns, seat-inventory-drift, brokkr-landscape-scan, soong-ci-relay) | `Restart=no` — any failure lands in `failed` immediately |
|
||||
| **Covered** | `svos.service` | burst 3 per **5min** with `RestartSec=5s` — 3 restarts fit in 15s, so it genuinely gives up |
|
||||
| **NOT covered** | booth, althing-po-herald, althing-seat-page, peedlar, wherethef, ttyd-caddy, ttyd-ro, ttyd-rw, lrpg-demo, zellij-web | `RestartSec` ≥ 2s against burst 5 per 10s — they flap forever instead |
|
||||
| **NOT covered, definitively** | `hermes-gateway` | `StartLimitIntervalUSec=0` — start limiting disabled, it never gives up at all |
|
||||
| **NOT covered, definitively** | `hermes-gateway` | `StartLimitIntervalUSec=0` — start limiting disabled, it never gives up at all. Its hook could only ever fire false positives (a stop-time exit 1, from a Hermes race; see `servers/nh3-dev/README.md`), so drop-in `20-planned-stop-exit1-is-clean.conf` sets `SuccessExitStatus=1`, 2026-09-23 |
|
||||
|
||||
`svos.service`'s divergent 5-minute window is **deliberate and load-bearing**
|
||||
(operator ruling 2026-09-11, "fatal both ways"). Do **not** harmonise it to 10s:
|
||||
|
||||
Reference in New Issue
Block a user