Files
esh-pfi-infrastructure/services/althing-notify-failure
vh 00a03921ff docs(alerts): reset-failed sits on the remediation path, which a flap sampler must design for
svos-dev found an interaction between the two detectors by doing the recovery
rather than reasoning about it: clearing a failed unit is REQUIRED before
systemd will start it again, and reset-failed zeroes NRestarts in the same
motion. Verified here -- a start-limited unit refused to start until
reset-failed, which took NRestarts 3 -> 0.

So the standard recovery from a failed-START alarm erases the history a flap
sampler would depend on, routinely rather than as an edge case. A unit that
fails, is cleared, then flaps starts its history from zero, delaying the first
flap alarm by the whole threshold -- exactly the window where a just-fixed unit
is most likely still sick. Mitigation recorded: store a last-seen timestamp so
a counter going BACKWARDS is itself detectable as a reset rather than read as
quiet.

Also recorded as its own shape: 'nothing to count is not the same as not
counting'. The first manual-restart check read NRestarts flat and would have
concluded a sampler goes blind whenever a human touches a unit. Artifact of the
test -- that unit had already hit its start limit and stopped auto-restarting,
so the instrument was reading correctly from a subject that had stopped
producing. Verify, then ask what the verification could not have seen.
2026-09-22 08:59:21 -07:00
..

althing-notify-failure — failed-START alarms for fleet user units

A systemd OnFailure template that reports any unit entering failed state to the althing infra-ops inbox. Installed on nh3-dev; hooks 12 user units.

services/althing-notify-failure/install.sh --dry-run   # see what it would hook
services/althing-notify-failure/install.sh             # idempotent; re-run to pick up new units

The failure it exists to catch

An uptime check cannot see the dangerous one. Measured by svos-dev, 2026-09-22: a config change on 09-19 made svos.service refuse to boot, the running process predated the change and kept serving, and the service sat one restart from dark for three days. Every uptime probe was green and correct the whole time — the thing was up. Nothing would have fired until the next restart, and then it would have been an outage rather than a warning.

The signal that catches it is failed-START, not down. A count taken after that conversation:

13 running user units on nh3-dev. Zero with an OnFailure hook.

Including althing-po-herald — whose silent failure cuts infra-ops's own mail delivery, a blind spot in the notification path every other alarm depends on.

Why this is not noise

OnFailure does not fire on a clean restart or a deliberate stop. svos-dev's four restarts and three deploys in one day would have produced zero alerts.

Duplicate suppression, and a claim I had to correct

This README originally asserted that a crash-loop "yields one message per episode". svos-dev challenged it with a measurement: one 2026-09-19 boot-gate refusal on svos.service produced five transitions into failed, and the operator got five messages.

Measured here before accepting or rejecting it: a unit with Restart=on-failure, burst 3, interval 30s produced 7 journal failure lines and exactly 1 notifier invocation. So the multiplier is not universal — it needs retries spanning start-limit windows, or an external restarter. svos.service has StartLimitIntervalSec=5min, which is how it accumulated five.

Both conditions exist on this box, so there is now a cooldown (900s default, NOTIFY_COOLDOWN_SECONDS) — cheap insurance rather than a fix for something proven here. Two properties it must have, both tested:

  • Keyed on the CAUSE, not the unit. A genuinely different failure inside the window is a new fact and still pages. Suppressing by unit name alone would hide a second, worse failure behind the first.
  • Suppression is logged (suppressed.log), never silent. An alarm that quietly declines to fire is indistinguishable from one that is broken, and this whole mechanism exists because something that looked fine was not.

Verified: same cause fired 3× → 1 sent, 2 suppressed and logged; a different cause inside the same window → sent.

⚠ COVERAGE — read this before trusting the alarm

OnFailure fires when a unit enters failed. A unit that keeps re-earning its retries never enters failed, so the alarm never fires for it. That is not a subtlety; it silently excludes most daemons on this box.

A unit reaches failed only by exhausting StartLimitBurst starts inside StartLimitIntervalSec. With RestartSec=5s against burst 5 per 10s, roughly two restarts fit in a window — the burst is unexhaustible.

Measured 2026-09-22, a unit shaped like booth / althing-po-herald (Restart=always, RestartSec=5s, burst 5, interval 10s): 9 restarts over 45 seconds, state cycling active/activating, never once failed. No alarm would have fired.

coverage units why
Covered the 7 timer-driven oneshots (dev-backup, ha-backup, fleet-tls-cert-check, headscale-ddns, seat-inventory-drift, brokkr-landscape-scan, soong-ci-relay) Restart=no — any failure lands in failed immediately
Covered svos.service burst 3 per 5min with RestartSec=5s — 3 restarts fit in 15s, so it genuinely gives up
NOT covered booth, althing-po-herald, althing-seat-page, peedlar, wherethef, ttyd-caddy, ttyd-ro, ttyd-rw, lrpg-demo, zellij-web RestartSec ≥ 2s against burst 5 per 10s — they flap forever instead
NOT covered, definitively hermes-gateway StartLimitIntervalUSec=0 — start limiting disabled, it never gives up at all

svos.service's divergent 5-minute window is deliberate and load-bearing (operator ruling 2026-09-11, "fatal both ways"). Do not harmonise it to 10s: that would restore the flapping the ruling forbids and silence the alarm on the one unit that currently works. The open question is whether the others should move to 5min — a behaviour change for ten services, so an operator call, not a tidy-up.

The shape both of this tool's bugs shared

svos-dev's framing, worth keeping: a tool that enumerates "things that are fine" and acts on them has selected against its own subject.

  • install.sh selected --state=running, so a unit already down at install time was never hooked — the unit most needing the alarm is by definition the one that is not healthy.
  • The parser took awk '{print $1}', which returns systemd's ● decoration for a failed unit, so those rows were reduced to empty and dropped.

Neither failed randomly. Both failure modes were perfectly anti-correlated with the purpose: the healthier the fleet, the better they appeared to work. That is worse than a broken instrument, which at least fails visibly on the easy cases too.

If a flap sampler is ever built, design for these first

The uncovered class above is detectable but not eventable: a flapping unit increments NRestarts on every automatic restart while never entering failed. That needs a sampler (timer reads the counter, alarms on delta), not a hook. Design notes from svos-dev, with what was measured here rather than assumed:

property measured 2026-09-22
survives daemon-reload yes (4 → 4)
survives a manual restart yes, does not reset
keeps counting after one yes (3 → 4 → 7 → 10)
reset-failed zeroes it (4 → 0)
stop → start resets

⚠ reset-failed is on the REMEDIATION PATH of the other alarm, so the reset is routine rather than an edge case. A unit that exhausts its start limit lands in failed and systemd refuses further starts until it is cleared — verified: a start-limited unit could only be restarted after reset-failed, which zeroed NRestarts from 3 to 0 in the same motion. So the standard recovery from a failed-START alarm erases the history a flap sampler depends on, for that unit, every time.

Consequence to design for rather than discover: a unit that fails, is cleared, then begins flapping starts its flap history from zero, so the first alarm after a reset is delayed by the whole threshold — precisely the window in which a just-"fixed" unit is most likely to still be sick. Mitigation (svos-dev): store a last-seen timestamp beside the count, so a counter going BACKWARDS is itself a detectable event and can be recorded as a reset rather than read as quiet.

⚠ "Nothing to count is not the same as not counting." The first run of the manual-restart check showed NRestarts flat and would have concluded a sampler goes blind whenever a human touches a unit. It was an artifact of the test: that unit had burst 5 per 10s, had already hit its limit, and had stopped auto-restarting — so there was nothing to increment. The instrument was reading correctly from a subject that had stopped producing. Its own shape, alongside the ● bug: verify, then ask what the verification could not have seen.