# Uptime Kuma — the fleet's service-layer monitors, as code. # # scripts/kuma seed stacks/uptimekuma/monitors.yaml # # Idempotent: keyed on NAME, so re-running edits rather than duplicating. Name # and not URL, because a URL legitimately repeats (a root and a /healthz on one # service) and a re-run must never fork the board. # # WHAT BELONGS HERE — the lane rule, so this file does not sprawl to 115 rows: # IN — services that should ALWAYS be up, whose silent death costs us work. # OUT — inference seats (they come and go BY DESIGN; a dormant seat is normal # and alerting on it is pure noise), hosts and hypervisors and firewalls # (Beszel's lane: 18 hosts x Status/CPU/Memory/Disk/Temp), and Uptime # Kuma itself (it cannot report its own death — that is Beszel's job). # # Every URL below was probed before being written here: all returned 200 on # 2026-09-21. A seed that ships red on day one teaches everyone to ignore the # board, which is how you end up with a monitor nobody reads. # # "althing chamber" was a candidate and was RETIRED instead (operator, # 2026-09-21): three of its four containers had never started since being # created on 2026-09-19, so :7881 refused. Stack, host dirs and image removed. # ---- where alerts GO --------------------------------------------------------- # Seeded BEFORE the monitors, and `applyExisting` attaches the channel to rows # that already exist. A board that detects and notifies nobody is precisely the # failure this service layer was built to close -- Homepage's own healthcheck # caught its 2026-09-18 death correctly and nothing was subscribed. # # Route: Kuma -> althing-alert-bridge (/kuma) -> postbox -> infra-ops inbox. # Same path Beszel uses, different route, so the subject says which tool spoke: # "[Uptime Kuma] Homepage is DOWN" rather than a Beszel-labelled lie. notifications: - name: althing (infra-ops) webhookURL: http://10.100.10.50:8096/kuma monitors: # ---- fleet toolchain: dead = agents and the operator are blocked ---- - name: althing post office url: http://10.100.50.40:8390/ description: the bus every agent session reads mail from - name: The Booth url: http://10.100.10.50:8090/healthz description: operator-review surface; real healthz, not a root page - name: task-board url: http://10.250.50.70:7878 - name: vor url: http://10.250.50.70:7879 # ---- the dashboard that started all this ---- # Dead for three days in Sept 2026 while Beszel correctly reported its host # UP. This row is the entire reason the service layer exists. - name: Homepage url: http://10.0.50.45:5100/ description: fleet dashboard (esh-docker-vm) — the 2026-09-18 silent death # ---- credentials + code: dead = nothing ships ---- - name: Gitea url: https://gitea.phasefinal.com - name: Vaultwarden url: https://vaultwarden.phasefinal.com description: the vault every agent reads credentials from # ---- AI control plane (the gateways, NOT the seats behind them) ---- - name: LiteLLM Gateway url: http://10.250.50.70:4000/health/liveliness description: liveliness endpoint, so a model outage does not read as a gateway outage - name: Asset Engine url: http://10.250.50.70:8200/api/v1/services - name: talk url: https://talk.nh3.phasefinal.com:8092/ description: fleet voice bench # ---- monitoring + backup: a blind monitor is worse than none ---- - name: Beszel hub url: http://10.250.50.70:8090 description: the host layer; if this is down we are blind to 18 hosts - name: Backrest url: http://10.250.50.70:9898 description: restic orchestration — a silent backup failure is the expensive kind - name: Dozzle hub url: http://10.250.50.70:8088