docs(alerts): the alarm covers far less than I claimed — measured

OnFailure fires when a unit enters failed. A unit that keeps re-earning its
retries never enters failed, so the alarm never fires for it. svos-dev warned
this was the consequence of the interval divergence; measuring it showed the
gap is most of the box.

MEASURED: a unit shaped like booth/althing-po-herald (Restart=always,
RestartSec=5s, burst 5, interval 10s) ran 9 restarts over 45 seconds cycling
active/activating and NEVER reached failed. No alarm would have fired.

Covered: the 7 timer-driven oneshots (Restart=no, so any failure lands in
failed immediately) and svos.service (burst 3 per 5min -- it genuinely gives
up). NOT covered: ten daemons that flap instead, and hermes-gateway, whose
start limiting is disabled outright.

svos.service's divergent 5min window is deliberate and load-bearing (operator
ruling 2026-09-11, 'fatal both ways'). NOT to be harmonised: that would restore
the flapping the ruling forbids AND silence the alarm on the one daemon it
currently works for. Whether the OTHERS move to 5min is a behaviour change for
ten services and an operator call.

Also recorded, svos-dev's framing of the two installer bugs: a tool that
enumerates 'things that are fine' and acts on them has selected against its own
subject. Neither bug failed randomly -- both were anti-correlated with purpose,
working better the healthier the fleet looked.
This commit is contained in:
vh
2026-09-22 08:55:07 -07:00
parent ef70b2ffa1
commit 163bb97dea
+35 -35
View File
@@ -57,47 +57,47 @@ proven here. Two properties it must have, both tested:
Verified: same cause fired 3× → 1 sent, 2 suppressed and logged; a different
cause inside the same window → sent.
### Restart policy decides the exposure
### ⚠ COVERAGE — read this before trusting the alarm
| policy | units |
|---|---|
| `Restart=always` | althing-po-herald, hermes-gateway, lrpg-demo, ttyd-caddy, ttyd-ro, ttyd-rw |
| `Restart=on-failure` | althing-seat-page, booth, peedlar, svos, wherethef, zellij-web |
**`OnFailure` fires when a unit enters `failed`. A unit that keeps re-earning
its retries never enters `failed`, so the alarm never fires for it.** That is
not a subtlety; it silently excludes most daemons on this box.
`hermes-gateway` has `StartLimitIntervalUSec=0` — start limiting disabled, so it
retries forever and may never reach `failed` at all. Worth knowing before
trusting this alarm to cover it.
A unit reaches `failed` only by exhausting `StartLimitBurst` starts inside
`StartLimitIntervalSec`. With `RestartSec=5s` against burst 5 per 10s, roughly
two restarts fit in a window — the burst is **unexhaustible**.
## Two traps, both hit during the build
**Measured 2026-09-22**, a unit shaped like `booth` / `althing-po-herald`
(`Restart=always`, `RestartSec=5s`, burst 5, interval 10s): **9 restarts over
45 seconds, state cycling `active`/`activating`, never once `failed`.** No
alarm would have fired.
⚠ **`%i`, never `%I`.** `%I` *unescapes* the instance name and systemd escaping
maps `-` to `/`. The acceptance test fired with `%I` and delivered a message for
`onfailure/selftest.service` — a unit that does not exist — with a spool path
that tried to create directories. `althing-po-herald.service` would have arrived
as `althing/po/herald.service`. The referring unit passes `%n` **raw**, so the
literal `%i` is correct.
| coverage | units | why |
|---|---|---|
| **Covered** | the 7 timer-driven oneshots (dev-backup, ha-backup, fleet-tls-cert-check, headscale-ddns, seat-inventory-drift, brokkr-landscape-scan, soong-ci-relay) | `Restart=no` — any failure lands in `failed` immediately |
| **Covered** | `svos.service` | burst 3 per **5min** with `RestartSec=5s` — 3 restarts fit in 15s, so it genuinely gives up |
| **NOT covered** | booth, althing-po-herald, althing-seat-page, peedlar, wherethef, ttyd-caddy, ttyd-ro, ttyd-rw, lrpg-demo, zellij-web | `RestartSec` ≥ 2s against burst 5 per 10s — they flap forever instead |
| **NOT covered, definitively** | `hermes-gateway` | `StartLimitIntervalUSec=0` — start limiting disabled, it never gives up at all |
⚠ **The notifier must never report itself.** Guarded twice on purpose: the
template carries no `OnFailure` of its own, and the script bails on an instance
name matching itself. A notification loop is the one bug that pages you forever.
`svos.service`'s divergent 5-minute window is **deliberate and load-bearing**
(operator ruling 2026-09-11, "fatal both ways"). Do **not** harmonise it to 10s:
that would restore the flapping the ruling forbids *and* silence the alarm on
the one unit that currently works. The open question is whether the others
should move **to** 5min — a behaviour change for ten services, so an operator
call, not a tidy-up.
## Delivery, honestly
## The shape both of this tool's bugs shared
The script writes a **durable local record first**
(`~/.local/state/althing-notify-failure/`), then sends. `postbox` has no outbox —
a send that cannot reach the post office is dropped — and this alarm exists
precisely for moments nobody is watching, so the signal must survive the
post office being one of the things that is down.
svos-dev's framing, worth keeping: **a tool that enumerates "things that are
fine" and acts on them has selected against its own subject.**
If `althing-po-herald` is the unit that failed, the message still **reaches**
the post office (postbox talks to it directly; the herald only delivers inbound
pokes). It will not be pushed into a live session, but it is stored and the next
`postbox read` finds it — a memo to the successor, which is the point.
- `install.sh` selected `--state=running`, so a unit already down at install
time was never hooked — the unit most needing the alarm is by definition the
one that is not healthy.
- The parser took `awk '{print $1}'`, which returns systemd's `●` decoration
for a **failed** unit, so those rows were reduced to empty and dropped.
## Verified
Acceptance-tested by firing a real unit that exits 42, twice: once to catch the
`%I` bug, once to confirm the fix. Delivered subject
`[systemd] onfailure-selftest.service FAILED on nh3-dev`, with state, result,
exit status, timestamp and the last 25 journal lines; spool file written. The
test unit was removed afterwards.
Neither failed randomly. Both failure modes were *perfectly anti-correlated
with the purpose*: the healthier the fleet, the better they appeared to work.
That is worse than a broken instrument, which at least fails visibly on the
easy cases too.