Files
esh-pfi-infrastructure/stacks/uptimekuma/monitors.yaml
T
vh 94899d6fa3 feat(uptimekuma): normalize names off Homepage, publish the status page, restore the widget
NAMES. Homepage already answers "what is this service called", so the monitor
name is now that name verbatim -- a second naming authority is how drift starts,
and an alert reading "[Uptime Kuma] Beszel hub is DOWN" sends you hunting for a
card that does not exist. Only two rows moved (Beszel hub -> Beszel, Dozzle hub
-> Dozzle); the " hub" suffixes were mine, not the services'.

The remaining mixed case is deliberate and is now documented as such. talk, vor
and task-board are lowercase on Homepage and in their own repos; title-casing
them here would make this board disagree with both. What actually looked messy
was scripts/kuma's own ASCII-ordinal sort, which buried every lowercase name
below every capitalised one. Fixed to case-insensitive.

⚠ RENAME SAFETY, which this pass needed and did not have. The seed keys on NAME,
so editing a name would have read as a brand-new monitor: added fresh, with the
old row orphaned, still checking, still alerting, and holding all the history.
`rename_from:` names the old row for one run. Verified: both renamed monitors
kept their IDs and all 67 heartbeats.

Added with it, an orphan warning for any row on the board the spec no longer
names -- because a forgotten monitor keeps paging. Its first cut diffed against
the PRE-EDIT snapshot and so cried wolf on its own successful renames; it
re-reads the board now. A warning that fires on its own correct work is worse
than no warning.

STATUS PAGE + WIDGET. The Homepage uptimekuma widget reads a PUBLISHED status
page (/api/status-page/<slug>), not the admin API -- which is why the widget
labels were deliberately absent from the rebuild: a dashboard widget pointed at
a 404 is the suspected mechanism behind both of Homepage's unkillable D-state
wedges, so shipping one on purpose would have been daft.

The page now exists at slug `nethealth` (the pre-rebuild slug, so old references
still resolve) and is DECLARED IN monitors.yaml, applied by `kuma seed`. Same
principle as the notification channel: a from-scratch rebuild restores the page,
the channel and the monitors together, and nothing the widget depends on lives
only in Kuma's database. Verified in a browser: "13 SITES UP / 0 SITES DOWN /
100% UPTIME" on the dashboard.

⚠ saveStatusPage calls imgDataUrl.startsWith() unconditionally, so passing null
throws and leaves the page CREATED BUT EMPTY -- which reads as success from
/api/status-page (200, correct title) while the group list is silently blank.
Pass "" instead. Commented at the call site.
2026-09-22 00:55:29 -07:00

113 lines
5.0 KiB
YAML

# Uptime Kuma — the fleet's service-layer monitors, as code.
#
# scripts/kuma seed stacks/uptimekuma/monitors.yaml
#
# Idempotent: keyed on NAME, so re-running edits rather than duplicating. Name
# and not URL, because a URL legitimately repeats (a root and a /healthz on one
# service) and a re-run must never fork the board.
#
# WHAT BELONGS HERE — the lane rule, so this file does not sprawl to 115 rows:
# IN — services that should ALWAYS be up, whose silent death costs us work.
# OUT — inference seats (they come and go BY DESIGN; a dormant seat is normal
# and alerting on it is pure noise), hosts and hypervisors and firewalls
# (Beszel's lane: 18 hosts x Status/CPU/Memory/Disk/Temp), and Uptime
# Kuma itself (it cannot report its own death — that is Beszel's job).
#
# NAMES COME FROM HOMEPAGE, VERBATIM. Homepage already answers "what is this
# service called", and a second naming authority is how drift starts: an alert
# reading "[Uptime Kuma] Beszel hub is DOWN" sends you looking for a card called
# "Beszel hub" that does not exist. So the monitor name IS the dashboard name --
# which is why this normalisation pass only moved two rows. The remaining mixed
# case (talk, vor, task-board against Gitea, Backrest) is NOT an inconsistency to
# fix: those are the products' own names, lowercase on Homepage and lowercase in
# their own repos. Title-casing them here would make this board disagree with
# both. `rename_from` exists for exactly this operation -- see scripts/kuma.
#
# Every URL below was probed before being written here: all returned 200 on
# 2026-09-21. A seed that ships red on day one teaches everyone to ignore the
# board, which is how you end up with a monitor nobody reads.
#
# "althing chamber" was a candidate and was RETIRED instead (operator,
# 2026-09-21): three of its four containers had never started since being
# created on 2026-09-19, so :7881 refused. Stack, host dirs and image removed.
# ---- where alerts GO ---------------------------------------------------------
# Seeded BEFORE the monitors, and `applyExisting` attaches the channel to rows
# that already exist. A board that detects and notifies nobody is precisely the
# failure this service layer was built to close -- Homepage's own healthcheck
# caught its 2026-09-18 death correctly and nothing was subscribed.
#
# Route: Kuma -> althing-alert-bridge (/kuma) -> postbox -> infra-ops inbox.
# Same path Beszel uses, different route, so the subject says which tool spoke:
# "[Uptime Kuma] Homepage is DOWN" rather than a Beszel-labelled lie.
# ---- the published status page -----------------------------------------------
# Exists so Homepage's `uptimekuma` widget has something to read: it calls
# /api/status-page/<slug> and /api/status-page/heartbeat/<slug>, and without a
# published page at that slug it polls a 404 forever. The widget labels on
# stacks/uptimekuma/compose.yaml were deliberately held back until this existed.
# Slug kept as `nethealth` -- the same one the pre-rebuild instance used, so any
# bookmark or older reference still resolves.
status_page:
slug: nethealth
title: PFI fleet services
group: Services
notifications:
- name: althing (infra-ops)
webhookURL: http://10.100.10.50:8096/kuma
monitors:
# ---- fleet toolchain: dead = agents and the operator are blocked ----
- name: althing post office
url: http://10.100.50.40:8390/
description: the bus every agent session reads mail from
- name: The Booth
url: http://10.100.10.50:8090/healthz
description: operator-review surface; real healthz, not a root page
- name: task-board
url: http://10.250.50.70:7878
- name: vor
url: http://10.250.50.70:7879
# ---- the dashboard that started all this ----
# Dead for three days in Sept 2026 while Beszel correctly reported its host
# UP. This row is the entire reason the service layer exists.
- name: Homepage
url: http://10.0.50.45:5100/
description: fleet dashboard (esh-docker-vm) — the 2026-09-18 silent death
# ---- credentials + code: dead = nothing ships ----
- name: Gitea
url: https://gitea.phasefinal.com
- name: Vaultwarden
url: https://vaultwarden.phasefinal.com
description: the vault every agent reads credentials from
# ---- AI control plane (the gateways, NOT the seats behind them) ----
- name: LiteLLM Gateway
url: http://10.250.50.70:4000/health/liveliness
description: liveliness endpoint, so a model outage does not read as a gateway outage
- name: Asset Engine
url: http://10.250.50.70:8200/api/v1/services
- name: talk
url: https://talk.nh3.phasefinal.com:8092/
description: fleet voice bench
# ---- monitoring + backup: a blind monitor is worse than none ----
- name: Beszel
url: http://10.250.50.70:8090
description: the host layer; if this is down we are blind to 18 hosts
- name: Backrest
url: http://10.250.50.70:9898
description: restic orchestration — a silent backup failure is the expensive kind
- name: Dozzle
url: http://10.250.50.70:8088