Files
esh-pfi-infrastructure/docs/fleettools/observability.md
T
vh e6da607767 chore(task-board): mothball it; superseded by the High Seat and ledger
Operator ruling 2026-09-24. On ana-docker the stack is `docker compose
down`: the container is removed and port 7878 is closed. Kept for revival:
- the data dir /opt/docker/conf/task-board/data (tasks.db, last written
  2026-09-11)
- the task-board:local image
- stacks/task-board/ and the host's compose + .env

The Uptime Kuma monitor (id 3) was deleted before the stop so it could
not page, and its row is removed from monitors.yaml. Homepage drops the
card on its own, since it reads the container's labels.

Hooks: the container log showed no hook POSTs in 30 days. The only
traffic was open browser tabs holding /events, and the plugin was already
uninstalled on nh3-dev. Removed the paragraph that told sessions to call
task_* tools (CLAUDE.md, and the fork-fleet.sh template that seeds new
repos), the vestigial TASK_BOARD_SESSION env in .claude/settings.json, and
the listings in README and FLEETTOOLS.
2026-09-24 09:22:50 -07:00

2.0 KiB

Observability — where to look when something is wrong

surface URL what it answers
Homepage http://10.0.50.45:5100 the curated fleet service map; /api/services is machine-readable
Beszel http://10.250.50.70:8090 host + container up/down, resource history; alerts land in althing
Dozzle http://10.250.50.70:8088 container logs, ana-docker + ana-ml2
Uptime Kuma http://10.0.50.45:3001 ESH-side service monitoring
Backrest http://10.250.50.70:9898 restic snapshot browse / restore
task-board — mothballed 2026-09-24; what a session is working on now lives on the High Seat (http://10.100.10.50:8770)
Dockge (per host) <host>:5001 compose UI on ana-docker, ana-ml2, esh-docker-vm, irv-ml1, nh3-docker
PBS 10.250.50.90:8007 (primary) · 10.100.50.90:8007 (DR) VM/LXC image backups

Homepage's /api/services is the single best machine-readable inventory of the fleet and it is maintained as a side effect of normal work. Query it rather than keeping a list.

Diagnostic reflexes

  • All hosts at one site flapping together = a WAN/mesh event, not host failure. Check uptimes first; they will show weeks. Then read the site's subnet-router tailscaled log.
  • ana-nas is a SPOF for postgres, rest-server-ana, PBS-ANA and cross-site restic. Check it first on any multi-failure.
  • After a power loss, crowdsec fails closed: it exits 255 and the Traefik bouncer 403s ALL HTTP. Restart crowdsec THEN traefik. This is a recurring recovery step.
  • A filter can silently narrow your window. docker logs --since HH:MM returning 0 while the line demonstrably exists, a find -maxdepth undercount, a character class that omits digits — re-run unfiltered before concluding absence. Absence of a signal is not a safe reading of it.
  • Timestamp every state report. An untimestamped claim is a claim about "now" and has manufactured phantom events in peer artifacts before.