Operator ruling 2026-09-24. On ana-docker the stack is `docker compose down`: the container is removed and port 7878 is closed. Kept for revival: - the data dir /opt/docker/conf/task-board/data (tasks.db, last written 2026-09-11) - the task-board:local image - stacks/task-board/ and the host's compose + .env The Uptime Kuma monitor (id 3) was deleted before the stop so it could not page, and its row is removed from monitors.yaml. Homepage drops the card on its own, since it reads the container's labels. Hooks: the container log showed no hook POSTs in 30 days. The only traffic was open browser tabs holding /events, and the plugin was already uninstalled on nh3-dev. Removed the paragraph that told sessions to call task_* tools (CLAUDE.md, and the fork-fleet.sh template that seeds new repos), the vestigial TASK_BOARD_SESSION env in .claude/settings.json, and the listings in README and FLEETTOOLS.
110 lines
4.9 KiB
YAML
110 lines
4.9 KiB
YAML
# Uptime Kuma — the fleet's service-layer monitors, as code.
|
|
#
|
|
# scripts/kuma seed stacks/uptimekuma/monitors.yaml
|
|
#
|
|
# Idempotent: keyed on NAME, so re-running edits rather than duplicating. Name
|
|
# and not URL, because a URL legitimately repeats (a root and a /healthz on one
|
|
# service) and a re-run must never fork the board.
|
|
#
|
|
# WHAT BELONGS HERE — the lane rule, so this file does not sprawl to 115 rows:
|
|
# IN — services that should ALWAYS be up, whose silent death costs us work.
|
|
# OUT — inference seats (they come and go BY DESIGN; a dormant seat is normal
|
|
# and alerting on it is pure noise), hosts and hypervisors and firewalls
|
|
# (Beszel's lane: 18 hosts x Status/CPU/Memory/Disk/Temp), and Uptime
|
|
# Kuma itself (it cannot report its own death — that is Beszel's job).
|
|
#
|
|
# NAMES COME FROM HOMEPAGE, VERBATIM. Homepage already answers "what is this
|
|
# service called", and a second naming authority is how drift starts: an alert
|
|
# reading "[Uptime Kuma] Beszel hub is DOWN" sends you looking for a card called
|
|
# "Beszel hub" that does not exist. So the monitor name IS the dashboard name --
|
|
# which is why this normalisation pass only moved two rows. The remaining mixed
|
|
# case (talk, vor, task-board against Gitea, Backrest) is NOT an inconsistency to
|
|
# fix: those are the products' own names, lowercase on Homepage and lowercase in
|
|
# their own repos. Title-casing them here would make this board disagree with
|
|
# both. `rename_from` exists for exactly this operation -- see scripts/kuma.
|
|
#
|
|
# Every URL below was probed before being written here: all returned 200 on
|
|
# 2026-09-21. A seed that ships red on day one teaches everyone to ignore the
|
|
# board, which is how you end up with a monitor nobody reads.
|
|
#
|
|
# "althing chamber" was a candidate and was RETIRED instead (operator,
|
|
# 2026-09-21): three of its four containers had never started since being
|
|
# created on 2026-09-19, so :7881 refused. Stack, host dirs and image removed.
|
|
|
|
# ---- where alerts GO ---------------------------------------------------------
|
|
# Seeded BEFORE the monitors, and `applyExisting` attaches the channel to rows
|
|
# that already exist. A board that detects and notifies nobody is precisely the
|
|
# failure this service layer was built to close -- Homepage's own healthcheck
|
|
# caught its 2026-09-18 death correctly and nothing was subscribed.
|
|
#
|
|
# Route: Kuma -> althing-alert-bridge (/kuma) -> postbox -> infra-ops inbox.
|
|
# Same path Beszel uses, different route, so the subject says which tool spoke:
|
|
# "[Uptime Kuma] Homepage is DOWN" rather than a Beszel-labelled lie.
|
|
# ---- the published status page -----------------------------------------------
|
|
# Exists so Homepage's `uptimekuma` widget has something to read: it calls
|
|
# /api/status-page/<slug> and /api/status-page/heartbeat/<slug>, and without a
|
|
# published page at that slug it polls a 404 forever. The widget labels on
|
|
# stacks/uptimekuma/compose.yaml were deliberately held back until this existed.
|
|
# Slug kept as `nethealth` -- the same one the pre-rebuild instance used, so any
|
|
# bookmark or older reference still resolves.
|
|
status_page:
|
|
slug: nethealth
|
|
title: PFI fleet services
|
|
group: Services
|
|
|
|
notifications:
|
|
- name: althing (infra-ops)
|
|
webhookURL: http://10.100.10.50:8096/kuma
|
|
|
|
monitors:
|
|
# ---- fleet toolchain: dead = agents and the operator are blocked ----
|
|
- name: althing post office
|
|
url: http://10.100.50.40:8390/
|
|
description: the bus every agent session reads mail from
|
|
|
|
- name: The Booth
|
|
url: http://10.100.10.50:8090/healthz
|
|
description: operator-review surface; real healthz, not a root page
|
|
|
|
- name: vor
|
|
url: http://10.250.50.70:7879
|
|
|
|
# ---- the dashboard that started all this ----
|
|
# Dead for three days in Sept 2026 while Beszel correctly reported its host
|
|
# UP. This row is the entire reason the service layer exists.
|
|
- name: Homepage
|
|
url: http://10.0.50.45:5100/
|
|
description: fleet dashboard (esh-docker-vm) — the 2026-09-18 silent death
|
|
|
|
# ---- credentials + code: dead = nothing ships ----
|
|
- name: Gitea
|
|
url: https://gitea.phasefinal.com
|
|
|
|
- name: Vaultwarden
|
|
url: https://vaultwarden.phasefinal.com
|
|
description: the vault every agent reads credentials from
|
|
|
|
# ---- AI control plane (the gateways, NOT the seats behind them) ----
|
|
- name: LiteLLM Gateway
|
|
url: http://10.250.50.70:4000/health/liveliness
|
|
description: liveliness endpoint, so a model outage does not read as a gateway outage
|
|
|
|
- name: Asset Engine
|
|
url: http://10.250.50.70:8200/api/v1/services
|
|
|
|
- name: talk
|
|
url: https://talk.nh3.phasefinal.com:8092/
|
|
description: fleet voice bench
|
|
|
|
# ---- monitoring + backup: a blind monitor is worse than none ----
|
|
- name: Beszel
|
|
url: http://10.250.50.70:8090
|
|
description: the host layer; if this is down we are blind to 18 hosts
|
|
|
|
- name: Backrest
|
|
url: http://10.250.50.70:9898
|
|
description: restic orchestration — a silent backup failure is the expensive kind
|
|
|
|
- name: Dozzle
|
|
url: http://10.250.50.70:8088
|