Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-22-instrument-errors.md
T
vh a2ade84470 memory: snapshot — Booth handed to booth-dev, Kuma rebuilt, alerts generalized, five instrument errors
In-flight rewritten for 2026-09-22. The Booth is booth-dev's now (extracted to
vh/booth, v0.2.0 shipped by them); infra-ops retains only its Uptime Kuma
monitor.

Eight decisions recorded, and three OPEN OPERATOR DECISIONS filed under Recent
decisions with commit tracking pointers rather than under in-flight, per the
deferred-work rule: the NRestarts flap sampler (163bb97), safe-rm, and
confirming the restic content-assertion ruling that reached me relayed by
svos-dev rather than Miranda (ba60fda).

Detail file for the day's throughline: five instrument errors, one shape -- a
tool that enumerates 'things that are fine' has selected against its own
subject. Four were mine, one a peer's, and mine included making theirs, on the
same day we were both cataloguing the habit. Written down because a rule both
parties broke while discussing the rule needs to live where a fresh session
reads it.

Two loose threads carried: talk.service failed while still serving on :8092,
and restic/ana/esh-docker-vm 36h behind the fleet.
2026-09-22 15:03:55 -07:00

72 lines
3.6 KiB
Markdown

# Five instrument errors in one day — and they are one shape
`[2026-09-22]` · infra-ops, in a running exchange with svos-dev
## The shape
> **A tool that enumerates "things that are fine" and acts on them has selected
> against its own subject.** (svos-dev's framing, taken close to verbatim.)
None of these failed randomly. Several had failure modes *anti-correlated with
their purpose* — the healthier the fleet looked, the better they appeared to
work. That is worse than an ordinary broken instrument, which at least fails
visibly on the easy cases too.
## The five
1. **`--state=running` in `install.sh`** (`ef70b2f`). Selected only healthy
units to hook, so a unit already down at install time was never hooked. The
unit most needing the alarm is by definition the one that is not healthy.
Found because `talk.service` was sitting in `failed` and uncovered.
2. **`awk '{print $1}'` over `systemctl list-units`** (`ef70b2f`). systemd
decorates a FAILED unit with a leading `●`, so field 1 is the bullet, not
the name; the sanitiser then reduced it to empty and dropped the row. The
parser silently lost *exactly the rows that matter*. Fix: `--plain`.
3. **`grep -ic restic backup-freshness-alert.sh` → 0** (`ba60fda`). That is the
WRAPPER. The checking lives in `check-backup-freshness.sh`, which had always
checked seven hosts. I grepped one file of a two-file tool and generalised to
the whole system — **the identical error svos-dev had made with `dev-backup`
four hours earlier, made by me while correcting them for it.**
4. **`restic ls` always prints a header line** (`ba60fda`). Absent path returns
1 line, real path returned 6; counting with `grep -c .` read the header as a
hit, so the ABSENT branch fell through and reported *"index intact, blobs
gone"*. An operator would have been sent to a repo-corruption runbook by a
mistyped path — wrong in the most expensive available direction. Found ONLY
by running the negative control. Fix: `grep -c '^/'`.
5. **A cooldown test that invoked the script before failing the unit**
(`02a6fc0`). The two calls legitimately saw different states, computed
different fingerprints, and the suppression correctly did not fire — so the
test "proved" a bug that did not exist. The code was right; the test was
wrong.
Adjacent, same family: **"nothing to count is not the same as not counting."**
An `NRestarts` probe read flat after a manual restart and would have concluded a
sampler goes blind whenever a human touches a unit. Artifact of the test — that
unit had already hit its start limit and stopped auto-restarting. The instrument
was reading correctly from a subject that had stopped producing.
## The rule this cost
**Not "verify". *Verify, then ask what the verification could not have seen.***
svos-dev sharpened it one notch and the sharpening is the useful half:
> The rung where you are CHECKING an instrument is the rung where you are least
> likely to check, because you are already in verification posture and it feels
> like the checking is happening.
## Why it is written down
Four of the five were mine, one was svos-dev's, and mine included making theirs.
Neither of us is careless — we each checked one instrument and generalised from
it, on the same day, while cataloguing the habit. A rule that both parties broke
while discussing the rule is a rule that needs to be written where a fresh
session will read it, not left as a lesson someone remembers.
Related: [[2026-09-21-booth-two-dead-controls]] — same family, found by the
operator rather than by me, because I read templates instead of rendering them.