docs(notify-failure): a stop that exits non-zero pages as a failure

A deliberate `systemctl --user restart hermes-gateway` paged infra-ops as a
FAILED unit (msg 3716) while the unit was already back up. Cause is a Hermes
v0.21.1 race: the planned-stop marker watcher runs the shutdown handler before
systemd's SIGTERM, consumes the marker, and the SIGTERM re-runs the handler,
which then classifies the stop as unexpected and exits 1.

Corrects the README claim that OnFailure never fires on a deliberate restart:
that holds only when the main process exits with a success status. Measured on a
throwaway unit (3/3 paged without SuccessExitStatus, 0/3 with it, and crash
restarts are unaffected), and a survey of every stop on nh3-dev since 09-15
found hermes-gateway to be the only unit that does this.

The host-side fix is drop-in hermes-gateway.service.d/
20-planned-stop-exit1-is-clean.conf (SuccessExitStatus=1). No alarm coverage
is lost: Restart=always ignores the classification, and StartLimitIntervalSec=0
means the unit can never reach `failed` from a start failure.
This commit is contained in:
vh
2026-09-23 02:15:13 -07:00
parent 6d616391fd
commit c2b0a05754
2 changed files with 28 additions and 1 deletions
+15
View File
@@ -96,6 +96,21 @@ local Bash already executes here — no SSH-to-self needed for non-privileged wo
is what SVOS's `_hermes_roster` derives its required-config line from —
narrowing does not blind it. Becomes `[svos_miranda]` once SVOS's plugin lands
in `$HERMES_HOME/plugins/`.
⚠ **A deliberate restart can exit 1, and that is a Hermes race, not a crash**
(2026-09-23). `ExecStop=gateway.systemd_stop_mark` writes a planned-stop
marker; the gateway's marker-watcher thread sometimes sees it before systemd's
SIGTERM lands, runs the shutdown handler ("Received UNKNOWN as a planned
gateway stop"), and consumes the marker. The SIGTERM then runs the handler a
second time, finds no marker, and exits 1 ("signal-initiated shutdown without
restart request"). `~/.hermes/logs/gateway.log` shows both lines 26 ms apart.
Before the 09-21 Hermes update every stop exited 1; since then it is
intermittent. The exit 1 used to page infra-ops through the `OnFailure` hook,
so drop-in `hermes-gateway.service.d/20-planned-stop-exit1-is-clean.conf`
sets `SuccessExitStatus=1`. That costs nothing: `Restart=always` restarts
whatever the exit status, and `StartLimitIntervalSec=0` means the unit never
reaches `failed` on a start failure, so the hook can't catch anything real on
this unit anyway. Delete the drop-in once upstream makes the handler
idempotent.
- **Hermes model backend → `gen-large` on the fleet LiteLLM gateway** (free local
compute), set 2026-09-14 per operator ruling. Until then `model.default` said
`anthropic/claude-opus-4.6` with `model.base_url` at openrouter, but