docs(notify-failure): a stop that exits non-zero pages as a failure

A deliberate `systemctl --user restart hermes-gateway` paged infra-ops as a
FAILED unit (msg 3716) while the unit was already back up. Cause is a Hermes
v0.21.1 race: the planned-stop marker watcher runs the shutdown handler before
systemd's SIGTERM, consumes the marker, and the SIGTERM re-runs the handler,
which then classifies the stop as unexpected and exits 1.

Corrects the README claim that OnFailure never fires on a deliberate restart:
that holds only when the main process exits with a success status. Measured on a
throwaway unit (3/3 paged without SuccessExitStatus, 0/3 with it, and crash
restarts are unaffected), and a survey of every stop on nh3-dev since 09-15
found hermes-gateway to be the only unit that does this.

The host-side fix is drop-in hermes-gateway.service.d/
20-planned-stop-exit1-is-clean.conf (SuccessExitStatus=1). No alarm coverage
is lost: Restart=always ignores the classification, and StartLimitIntervalSec=0
means the unit can never reach `failed` from a start failure.
This commit is contained in:
vh
2026-09-23 02:15:13 -07:00
parent 6d616391fd
commit c2b0a05754
2 changed files with 28 additions and 1 deletions
+13 -1
View File
@@ -30,6 +30,18 @@ delivery, a blind spot in the notification path every other alarm depends on.
`OnFailure` does not fire on a clean restart or a deliberate stop. svos-dev's
four restarts and three deploys in one day would have produced **zero** alerts.
⚠ **"Clean" means the main process exits with a success status when it gets
SIGTERM.** A daemon that exits non-zero on a *deliberate* `systemctl restart`
or `stop` lands the unit in `failed` for an instant, and `OnFailure` fires even
though the restart then carries on and succeeds. The page it sends reads
`state active (running)` — that combination is the tell. Measured 2026-09-23 on
a throwaway unit that exits 1 on SIGTERM: 3/3 restarts paged. Adding
`SuccessExitStatus=<code>` to the unit brought it to 0/3, and an exit outside a
stop job still auto-restarted under `Restart=always`. A survey of every stop on
nh3-dev since 09-15 (booth 49, peedlar 30, svos 10, …) found **one** unit that
does this — `hermes-gateway`, below — so the fix is per-unit, not in the
notifier.
### Duplicate suppression, and a claim I had to correct
This README originally asserted that a crash-loop "yields one message per
@@ -77,7 +89,7 @@ alarm would have fired.
| **Covered** | the 7 timer-driven oneshots (dev-backup, ha-backup, fleet-tls-cert-check, headscale-ddns, seat-inventory-drift, brokkr-landscape-scan, soong-ci-relay) | `Restart=no` — any failure lands in `failed` immediately |
| **Covered** | `svos.service` | burst 3 per **5min** with `RestartSec=5s` — 3 restarts fit in 15s, so it genuinely gives up |
| **NOT covered** | booth, althing-po-herald, althing-seat-page, peedlar, wherethef, ttyd-caddy, ttyd-ro, ttyd-rw, lrpg-demo, zellij-web | `RestartSec` ≥ 2s against burst 5 per 10s — they flap forever instead |
| **NOT covered, definitively** | `hermes-gateway` | `StartLimitIntervalUSec=0` — start limiting disabled, it never gives up at all |
| **NOT covered, definitively** | `hermes-gateway` | `StartLimitIntervalUSec=0` — start limiting disabled, it never gives up at all. Its hook could only ever fire false positives (a stop-time exit 1, from a Hermes race; see `servers/nh3-dev/README.md`), so drop-in `20-planned-stop-exit1-is-clean.conf` sets `SuccessExitStatus=1`, 2026-09-23 |
`svos.service`'s divergent 5-minute window is **deliberate and load-bearing**
(operator ruling 2026-09-11, "fatal both ways"). Do **not** harmonise it to 10s: