Files
esh-pfi-infrastructure/stacks/uptimekuma
vh e6da607767 chore(task-board): mothball it; superseded by the High Seat and ledger
Operator ruling 2026-09-24. On ana-docker the stack is `docker compose
down`: the container is removed and port 7878 is closed. Kept for revival:
- the data dir /opt/docker/conf/task-board/data (tasks.db, last written
  2026-09-11)
- the task-board:local image
- stacks/task-board/ and the host's compose + .env

The Uptime Kuma monitor (id 3) was deleted before the stop so it could
not page, and its row is removed from monitors.yaml. Homepage drops the
card on its own, since it reads the container's labels.

Hooks: the container log showed no hook POSTs in 30 days. The only
traffic was open browser tabs holding /events, and the plugin was already
uninstalled on nh3-dev. Removed the paragraph that told sessions to call
task_* tools (CLAUDE.md, and the fork-fleet.sh template that seeds new
repos), the vestigial TASK_BOARD_SESSION env in .claude/settings.json, and
the listings in README and FLEETTOOLS.
2026-09-24 09:22:50 -07:00
..

uptimekuma — the fleet's SERVICE-layer monitor (ana-docker)

louislam/uptime-kuma:2.5.5 on ana-docker (10.250.50.70:3001), beside the Beszel and Dozzle hubs.

Why it exists — the lane split

Measured 2026-09-21, after the fleet dashboard sat dead for three days while every monitoring tool reported correctly:

tool what it asks scope
Beszel is the box alive, and is it out of CPU / memory / disk / heat 18 hosts × {Status, CPU, Memory, Disk, Temperature}
Uptime Kuma (this) is the service actually serving the fleet's user-facing endpoints
Homepage — DISPLAY ONLY. Polls 38 URLs, alerts nobody.

Beszel and Kuma are not redundant — they are disjoint, and the seam between them is where things die quietly. Beszel's alerts bind to a system with a threshold (alerts table: system, name, value, min); there is no URL column, so it is structurally incapable of "this endpoint should return 200". That is not a configuration gap, it is the data model.

On 2026-09-18 homepage wedged in an unkillable D-state. Beszel reported esh-vm-docker: up — correctly; the host was up. Kuma was not watching it. The dead dashboard fell between two working instruments and stayed dark for three days until the operator hit a 404.

Rebuilt from scratch, 2026-09-21

Operator-authorised: "uptime-kuma was never really used… you can even dump the existing container and config and start over from scratch." The prior instance had four monitors — three firewall pings and one HTTPS check — and nothing was migrated. This also skipped the one-way v1→v2 database migration entirely.

Two things changed with the rebuild:

Pinned to 2.5.5, and :latest is now a documented trap. Upstream keeps latest on the 1.x line: an August 2026 pull produced an image built 2024-12-20 running 1.23.16. Verified by digest — latest and 1 resolve to the same image, while 2/next carry 2.5.5. 2.x has been stable since 2.2.0 (2026-03-05), twelve releases, zero prereleases. Pinned exactly rather than floating on 2, for the same reason latest burned us.

Moved esh-docker-vm → ana-docker. House placement rule (CLAUDE.md): cross- site services live on ana-docker and pull from agents on the other hosts — a fleet-wide service monitor is exactly that. And esh-docker-vm has wedged unkillably twice in four months (2026-06-03, 2026-09-18), so it is the worst box in the fleet to host the thing that would tell us. A monitor also cannot report the failure of the host it runs on; Beszel covers that layer, and the two should not share a failure domain.

Normalised at the same time: restart: always → unless-stopped, which the 2026-08-18 README flagged as "worth normalising on the next deliberate touch". always revives a container that was stopped on purpose.

Scripting it

There is no REST CRUD API — in either major version. Checked against the 2.5.5 source tree, not the docs: server/routers/ contains exactly two files, api-router.js (Prometheus /metrics, badges, entry page) and status-page-router.js. API keys unlock /metrics and badges only.

Automation goes over Socket.IO, the same channel the web UI uses; server/socket-handlers/general-socket-handler.js carries add, editMonitor, deleteMonitor, getMonitorList.

⚠️ Do not reach for uptime-kuma-api (the Python wrapper). It is abandoned — last release 2023-09-26 — and its stated ceiling is "support for uptime kuma 1.23.0 and 1.23.1". It does not support 2.x and never will. Use a thin first-party Socket.IO client instead.

The Homepage widget is deliberately absent

The compose carries no homepage.widget.* labels. The widget needs a published status-page slug, and a from-scratch install has none — it would poll a 404 forever. That matters more here than elsewhere: homepage widgets pointed at dead targets are the suspected mechanism behind both of that dashboard's unkillable wedges (see incident_esh_docker_nfs_boot_race, 2026-06-03 entry).

Re-add homepage.widget.type / .url / .slug only once the status page actually exists. Label changes need docker compose up -d, not restart.

Data

Named volume uptimekuma_uptime-kuma → /app/data holds every monitor definition and all history. Recreating the container is safe; deleting that volume is not.