# Uptime Kuma — the fleet's service-layer monitors, as code. # # scripts/kuma seed stacks/uptimekuma/monitors.yaml # # Idempotent: keyed on NAME, so re-running edits rather than duplicating. Name # and not URL, because a URL legitimately repeats (a root and a /healthz on one # service) and a re-run must never fork the board. # # WHAT BELONGS HERE — the lane rule, so this file does not sprawl to 115 rows: # IN — services that should ALWAYS be up, whose silent death costs us work. # OUT — inference seats (they come and go BY DESIGN; a dormant seat is normal # and alerting on it is pure noise), hosts and hypervisors and firewalls # (Beszel's lane: 18 hosts x Status/CPU/Memory/Disk/Temp), and Uptime # Kuma itself (it cannot report its own death — that is Beszel's job). # # NAMES COME FROM HOMEPAGE, VERBATIM. Homepage already answers "what is this # service called", and a second naming authority is how drift starts: an alert # reading "[Uptime Kuma] Beszel hub is DOWN" sends you looking for a card called # "Beszel hub" that does not exist. So the monitor name IS the dashboard name -- # which is why this normalisation pass only moved two rows. The remaining mixed # case (talk, vor, task-board against Gitea, Backrest) is NOT an inconsistency to # fix: those are the products' own names, lowercase on Homepage and lowercase in # their own repos. Title-casing them here would make this board disagree with # both. `rename_from` exists for exactly this operation -- see scripts/kuma. # # Every URL below was probed before being written here: all returned 200 on # 2026-09-21. A seed that ships red on day one teaches everyone to ignore the # board, which is how you end up with a monitor nobody reads. # # "althing chamber" was a candidate and was RETIRED instead (operator, # 2026-09-21): three of its four containers had never started since being # created on 2026-09-19, so :7881 refused. Stack, host dirs and image removed. # ---- where alerts GO --------------------------------------------------------- # Seeded BEFORE the monitors, and `applyExisting` attaches the channel to rows # that already exist. A board that detects and notifies nobody is precisely the # failure this service layer was built to close -- Homepage's own healthcheck # caught its 2026-09-18 death correctly and nothing was subscribed. # # Route: Kuma -> althing-alert-bridge (/kuma) -> postbox -> infra-ops inbox. # Same path Beszel uses, different route, so the subject says which tool spoke: # "[Uptime Kuma] Homepage is DOWN" rather than a Beszel-labelled lie. # ---- the published status page ----------------------------------------------- # Exists so Homepage's `uptimekuma` widget has something to read: it calls # /api/status-page/ and /api/status-page/heartbeat/, and without a # published page at that slug it polls a 404 forever. The widget labels on # stacks/uptimekuma/compose.yaml were deliberately held back until this existed. # Slug kept as `nethealth` -- the same one the pre-rebuild instance used, so any # bookmark or older reference still resolves. status_page: slug: nethealth title: PFI fleet services group: Services notifications: - name: althing (infra-ops) webhookURL: http://10.100.10.50:8096/kuma monitors: # ---- fleet toolchain: dead = agents and the operator are blocked ---- - name: althing post office url: http://10.100.50.40:8390/ description: the bus every agent session reads mail from - name: The Booth url: http://10.100.10.50:8090/healthz description: operator-review surface; real healthz, not a root page - name: task-board url: http://10.250.50.70:7878 - name: vor url: http://10.250.50.70:7879 # ---- the dashboard that started all this ---- # Dead for three days in Sept 2026 while Beszel correctly reported its host # UP. This row is the entire reason the service layer exists. - name: Homepage url: http://10.0.50.45:5100/ description: fleet dashboard (esh-docker-vm) — the 2026-09-18 silent death # ---- credentials + code: dead = nothing ships ---- - name: Gitea url: https://gitea.phasefinal.com - name: Vaultwarden url: https://vaultwarden.phasefinal.com description: the vault every agent reads credentials from # ---- AI control plane (the gateways, NOT the seats behind them) ---- - name: LiteLLM Gateway url: http://10.250.50.70:4000/health/liveliness description: liveliness endpoint, so a model outage does not read as a gateway outage - name: Asset Engine url: http://10.250.50.70:8200/api/v1/services - name: talk url: https://talk.nh3.phasefinal.com:8092/ description: fleet voice bench # ---- monitoring + backup: a blind monitor is worse than none ---- - name: Beszel url: http://10.250.50.70:8090 description: the host layer; if this is down we are blind to 18 hosts - name: Backrest url: http://10.250.50.70:9898 description: restic orchestration — a silent backup failure is the expensive kind - name: Dozzle url: http://10.250.50.70:8088