Files
esh-pfi-infrastructure/stacks/uptimekuma
vh 6f0a9b9fae feat(alerts): generalize the althing bridge, wire Kuma to it, retire chamber
THE BRIDGE. `beszel-althing` hardcoded a "[Beszel] " subject prefix and a
Beszel hub footer from when Beszel was its only caller. Routing Uptime Kuma
through it unchanged would have delivered Kuma outages labelled [Beszel],
pointing the reader at the wrong dashboard -- an alert that lies about its own
source is worse than no alert.

Now a route registry: /beszel and /kuma, each with its own prefix, footer and
payload parser, because the tools do not agree on a shape (Beszel sends
{title, message}; Kuma sends {heartbeat, monitor, msg}). Generalising cost a
dict; a sibling service would have cost a second unit, a second port and a
second thing to notice had died.

Renamed beszel-althing -> althing-alert-bridge with it. A service named after
one consumer that carries two is the invisible coupling that sends a future
session looking in the wrong place.

⚠ /beszel IS FROZEN and this refactor proves it rather than claiming it. The
three original tests were kept BYTE-UNCHANGED -- including the one asserting
the exact postbox argv -- and deliver() still defaults to the Beszel route so
they exercise it. A new test asserts the Kuma footer never leaks into a Beszel
body or vice versa. Verified live after the rename: a real POST to /beszel
landed as "[Beszel] BRIDGE RENAME CHECK" with the correct hub footer, read back
from the thread rather than trusted from the receipt.

Payload shapes are parsed HERE, not via Kuma's custom-webhook-body feature,
because Kuma's notification config lives in its own database -- and that
database was destroyed and rebuilt from scratch hours ago. Anything living only
in a tool's DB is lost on the next rebuild; format knowledge belongs in git,
next to a test.

parse_kuma also handles the monitorless case. testNotification and cert-expiry
alerts carry no monitor and no heartbeat, and the first cut fabricated "unknown
monitor is ?" from them -- caught by sending a real one and reading the subject,
not by the suite. Fixed, pinned, and the earlier test asserting the bad
behaviour was corrected rather than worked around.

KUMA IS NOW WIRED. scripts/kuma gained notification support and the channel is
in monitors.yaml, seeded BEFORE the monitors and with applyExisting, so a
rebuild restores alerting and not just detection. Ground truth from the DB:
13 of 13 monitors carry the channel.

⚠ A THIRD instance of the same class of bug, worth naming: notifications() is
pushed as `notificationList` at LOGIN ONLY -- there is no event to ask with. The
first cut cleared the captured value before waiting, discarding the only copy it
would ever be sent, then blocked for the full timeout and reported an empty
list. That reads exactly like "no channels configured" and is a lie. Same family
as the getMonitorList ack-vs-push trap, different shape.

End-to-end, both shapes, read back from the inbox:
  [Uptime Kuma] Homepage is DOWN          + target + board link
  [Uptime Kuma] althing (infra-ops) Testing   (no fabricated subject)
  [Beszel] BRIDGE RENAME CHECK            + hub footer, unchanged

ALTHING CHAMBER RETIRED (operator). Three of its four containers had never
started -- created 2026-09-19, StartedAt epoch-zero, 0 restarts -- so :7881
refused, and nothing was watching it. Only its valkey was running, on the
project's own network with no external consumer. Stack, compose/build/conf dirs
and the local image removed; the Homepage card went with the label.
2026-09-21 23:11:54 -07:00
..

uptimekuma — the fleet's SERVICE-layer monitor (ana-docker)

louislam/uptime-kuma:2.5.5 on ana-docker (10.250.50.70:3001), beside the Beszel and Dozzle hubs.

Why it exists — the lane split

Measured 2026-09-21, after the fleet dashboard sat dead for three days while every monitoring tool reported correctly:

tool what it asks scope
Beszel is the box alive, and is it out of CPU / memory / disk / heat 18 hosts × {Status, CPU, Memory, Disk, Temperature}
Uptime Kuma (this) is the service actually serving the fleet's user-facing endpoints
Homepage — DISPLAY ONLY. Polls 38 URLs, alerts nobody.

Beszel and Kuma are not redundant — they are disjoint, and the seam between them is where things die quietly. Beszel's alerts bind to a system with a threshold (alerts table: system, name, value, min); there is no URL column, so it is structurally incapable of "this endpoint should return 200". That is not a configuration gap, it is the data model.

On 2026-09-18 homepage wedged in an unkillable D-state. Beszel reported esh-vm-docker: up — correctly; the host was up. Kuma was not watching it. The dead dashboard fell between two working instruments and stayed dark for three days until the operator hit a 404.

Rebuilt from scratch, 2026-09-21

Operator-authorised: "uptime-kuma was never really used… you can even dump the existing container and config and start over from scratch." The prior instance had four monitors — three firewall pings and one HTTPS check — and nothing was migrated. This also skipped the one-way v1→v2 database migration entirely.

Two things changed with the rebuild:

Pinned to 2.5.5, and :latest is now a documented trap. Upstream keeps latest on the 1.x line: an August 2026 pull produced an image built 2024-12-20 running 1.23.16. Verified by digest — latest and 1 resolve to the same image, while 2/next carry 2.5.5. 2.x has been stable since 2.2.0 (2026-03-05), twelve releases, zero prereleases. Pinned exactly rather than floating on 2, for the same reason latest burned us.

Moved esh-docker-vm → ana-docker. House placement rule (CLAUDE.md): cross- site services live on ana-docker and pull from agents on the other hosts — a fleet-wide service monitor is exactly that. And esh-docker-vm has wedged unkillably twice in four months (2026-06-03, 2026-09-18), so it is the worst box in the fleet to host the thing that would tell us. A monitor also cannot report the failure of the host it runs on; Beszel covers that layer, and the two should not share a failure domain.

Normalised at the same time: restart: always → unless-stopped, which the 2026-08-18 README flagged as "worth normalising on the next deliberate touch". always revives a container that was stopped on purpose.

Scripting it

There is no REST CRUD API — in either major version. Checked against the 2.5.5 source tree, not the docs: server/routers/ contains exactly two files, api-router.js (Prometheus /metrics, badges, entry page) and status-page-router.js. API keys unlock /metrics and badges only.

Automation goes over Socket.IO, the same channel the web UI uses; server/socket-handlers/general-socket-handler.js carries add, editMonitor, deleteMonitor, getMonitorList.

⚠️ Do not reach for uptime-kuma-api (the Python wrapper). It is abandoned — last release 2023-09-26 — and its stated ceiling is "support for uptime kuma 1.23.0 and 1.23.1". It does not support 2.x and never will. Use a thin first-party Socket.IO client instead.

The Homepage widget is deliberately absent

The compose carries no homepage.widget.* labels. The widget needs a published status-page slug, and a from-scratch install has none — it would poll a 404 forever. That matters more here than elsewhere: homepage widgets pointed at dead targets are the suspected mechanism behind both of that dashboard's unkillable wedges (see incident_esh_docker_nfs_boot_race, 2026-06-03 entry).

Re-add homepage.widget.type / .url / .slug only once the status page actually exists. Label changes need docker compose up -d, not restart.

Data

Named volume uptimekuma_uptime-kuma → /app/data holds every monitor definition and all history. Recreating the container is safe; deleting that volume is not.