Files
esh-pfi-infrastructure/stacks/uptimekuma/monitors.yaml
T
vh e6da607767 chore(task-board): mothball it; superseded by the High Seat and ledger
Operator ruling 2026-09-24. On ana-docker the stack is `docker compose
down`: the container is removed and port 7878 is closed. Kept for revival:
- the data dir /opt/docker/conf/task-board/data (tasks.db, last written
  2026-09-11)
- the task-board:local image
- stacks/task-board/ and the host's compose + .env

The Uptime Kuma monitor (id 3) was deleted before the stop so it could
not page, and its row is removed from monitors.yaml. Homepage drops the
card on its own, since it reads the container's labels.

Hooks: the container log showed no hook POSTs in 30 days. The only
traffic was open browser tabs holding /events, and the plugin was already
uninstalled on nh3-dev. Removed the paragraph that told sessions to call
task_* tools (CLAUDE.md, and the fork-fleet.sh template that seeds new
repos), the vestigial TASK_BOARD_SESSION env in .claude/settings.json, and
the listings in README and FLEETTOOLS.
2026-09-24 09:22:50 -07:00

110 lines
4.9 KiB
YAML

# Uptime Kuma — the fleet's service-layer monitors, as code.
#
# scripts/kuma seed stacks/uptimekuma/monitors.yaml
#
# Idempotent: keyed on NAME, so re-running edits rather than duplicating. Name
# and not URL, because a URL legitimately repeats (a root and a /healthz on one
# service) and a re-run must never fork the board.
#
# WHAT BELONGS HERE — the lane rule, so this file does not sprawl to 115 rows:
# IN — services that should ALWAYS be up, whose silent death costs us work.
# OUT — inference seats (they come and go BY DESIGN; a dormant seat is normal
# and alerting on it is pure noise), hosts and hypervisors and firewalls
# (Beszel's lane: 18 hosts x Status/CPU/Memory/Disk/Temp), and Uptime
# Kuma itself (it cannot report its own death — that is Beszel's job).
#
# NAMES COME FROM HOMEPAGE, VERBATIM. Homepage already answers "what is this
# service called", and a second naming authority is how drift starts: an alert
# reading "[Uptime Kuma] Beszel hub is DOWN" sends you looking for a card called
# "Beszel hub" that does not exist. So the monitor name IS the dashboard name --
# which is why this normalisation pass only moved two rows. The remaining mixed
# case (talk, vor, task-board against Gitea, Backrest) is NOT an inconsistency to
# fix: those are the products' own names, lowercase on Homepage and lowercase in
# their own repos. Title-casing them here would make this board disagree with
# both. `rename_from` exists for exactly this operation -- see scripts/kuma.
#
# Every URL below was probed before being written here: all returned 200 on
# 2026-09-21. A seed that ships red on day one teaches everyone to ignore the
# board, which is how you end up with a monitor nobody reads.
#
# "althing chamber" was a candidate and was RETIRED instead (operator,
# 2026-09-21): three of its four containers had never started since being
# created on 2026-09-19, so :7881 refused. Stack, host dirs and image removed.
# ---- where alerts GO ---------------------------------------------------------
# Seeded BEFORE the monitors, and `applyExisting` attaches the channel to rows
# that already exist. A board that detects and notifies nobody is precisely the
# failure this service layer was built to close -- Homepage's own healthcheck
# caught its 2026-09-18 death correctly and nothing was subscribed.
#
# Route: Kuma -> althing-alert-bridge (/kuma) -> postbox -> infra-ops inbox.
# Same path Beszel uses, different route, so the subject says which tool spoke:
# "[Uptime Kuma] Homepage is DOWN" rather than a Beszel-labelled lie.
# ---- the published status page -----------------------------------------------
# Exists so Homepage's `uptimekuma` widget has something to read: it calls
# /api/status-page/<slug> and /api/status-page/heartbeat/<slug>, and without a
# published page at that slug it polls a 404 forever. The widget labels on
# stacks/uptimekuma/compose.yaml were deliberately held back until this existed.
# Slug kept as `nethealth` -- the same one the pre-rebuild instance used, so any
# bookmark or older reference still resolves.
status_page:
slug: nethealth
title: PFI fleet services
group: Services
notifications:
- name: althing (infra-ops)
webhookURL: http://10.100.10.50:8096/kuma
monitors:
# ---- fleet toolchain: dead = agents and the operator are blocked ----
- name: althing post office
url: http://10.100.50.40:8390/
description: the bus every agent session reads mail from
- name: The Booth
url: http://10.100.10.50:8090/healthz
description: operator-review surface; real healthz, not a root page
- name: vor
url: http://10.250.50.70:7879
# ---- the dashboard that started all this ----
# Dead for three days in Sept 2026 while Beszel correctly reported its host
# UP. This row is the entire reason the service layer exists.
- name: Homepage
url: http://10.0.50.45:5100/
description: fleet dashboard (esh-docker-vm) — the 2026-09-18 silent death
# ---- credentials + code: dead = nothing ships ----
- name: Gitea
url: https://gitea.phasefinal.com
- name: Vaultwarden
url: https://vaultwarden.phasefinal.com
description: the vault every agent reads credentials from
# ---- AI control plane (the gateways, NOT the seats behind them) ----
- name: LiteLLM Gateway
url: http://10.250.50.70:4000/health/liveliness
description: liveliness endpoint, so a model outage does not read as a gateway outage
- name: Asset Engine
url: http://10.250.50.70:8200/api/v1/services
- name: talk
url: https://talk.nh3.phasefinal.com:8092/
description: fleet voice bench
# ---- monitoring + backup: a blind monitor is worse than none ----
- name: Beszel
url: http://10.250.50.70:8090
description: the host layer; if this is down we are blind to 18 hosts
- name: Backrest
url: http://10.250.50.70:9898
description: restic orchestration — a silent backup failure is the expensive kind
- name: Dozzle
url: http://10.250.50.70:8088