NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled. - nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml. - nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4) deployed with HOST_NAME/HOST_IP labels. - Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min 0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control 0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed identical within rep spread. - gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006). With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl reopen; applied on nh3-pve (one package). - pve-nvidia-host.yaml: document that the headers meta drags in the newest kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot). - Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal. - nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels, AMT cabled but unreachable on the network. Gateway routing to nh3-ml1 is not changed.
uptimekuma — the fleet's SERVICE-layer monitor (ana-docker)
louislam/uptime-kuma:2.5.5 on ana-docker (10.250.50.70:3001), beside the
Beszel and Dozzle hubs.
Why it exists — the lane split
Measured 2026-09-21, after the fleet dashboard sat dead for three days while every monitoring tool reported correctly:
| tool | what it asks | scope |
|---|---|---|
| Beszel | is the box alive, and is it out of CPU / memory / disk / heat | 18 hosts × {Status, CPU, Memory, Disk, Temperature} |
| Uptime Kuma (this) | is the service actually serving | the fleet's user-facing endpoints |
| Homepage | — | DISPLAY ONLY. Polls 38 URLs, alerts nobody. |
Beszel and Kuma are not redundant — they are disjoint, and the seam between
them is where things die quietly. Beszel's alerts bind to a system with a
threshold (alerts table: system, name, value, min); there is no URL
column, so it is structurally incapable of "this endpoint should return 200".
That is not a configuration gap, it is the data model.
On 2026-09-18 homepage wedged in an unkillable D-state. Beszel reported
esh-vm-docker: up — correctly; the host was up. Kuma was not watching it.
The dead dashboard fell between two working instruments and stayed dark for
three days until the operator hit a 404.
Rebuilt from scratch, 2026-09-21
Operator-authorised: "uptime-kuma was never really used… you can even dump the existing container and config and start over from scratch." The prior instance had four monitors — three firewall pings and one HTTPS check — and nothing was migrated. This also skipped the one-way v1→v2 database migration entirely.
Two things changed with the rebuild:
Pinned to 2.5.5, and :latest is now a documented trap. Upstream keeps
latest on the 1.x line: an August 2026 pull produced an image built
2024-12-20 running 1.23.16. Verified by digest — latest and 1 resolve to
the same image, while 2/next carry 2.5.5. 2.x has been stable since 2.2.0
(2026-03-05), twelve releases, zero prereleases. Pinned exactly rather than
floating on 2, for the same reason latest burned us.
Moved esh-docker-vm → ana-docker. House placement rule (CLAUDE.md): cross-
site services live on ana-docker and pull from agents on the other hosts — a
fleet-wide service monitor is exactly that. And esh-docker-vm has wedged
unkillably twice in four months (2026-06-03, 2026-09-18), so it is the worst
box in the fleet to host the thing that would tell us. A monitor also cannot
report the failure of the host it runs on; Beszel covers that layer, and the
two should not share a failure domain.
Normalised at the same time: restart: always → unless-stopped, which the
2026-08-18 README flagged as "worth normalising on the next deliberate touch".
always revives a container that was stopped on purpose.
Scripting it
There is no REST CRUD API — in either major version. Checked against the
2.5.5 source tree, not the docs: server/routers/ contains exactly two files,
api-router.js (Prometheus /metrics, badges, entry page) and
status-page-router.js. API keys unlock /metrics and badges only.
Automation goes over Socket.IO, the same channel the web UI uses;
server/socket-handlers/general-socket-handler.js carries add,
editMonitor, deleteMonitor, getMonitorList.
⚠️ Do not reach for uptime-kuma-api (the Python wrapper). It is abandoned
— last release 2023-09-26 — and its stated ceiling is "support for uptime kuma
1.23.0 and 1.23.1". It does not support 2.x and never will. Use a thin
first-party Socket.IO client instead.
The Homepage widget is deliberately absent
The compose carries no homepage.widget.* labels. The widget needs a
published status-page slug, and a from-scratch install has none — it would poll
a 404 forever. That matters more here than elsewhere: homepage widgets pointed
at dead targets are the suspected mechanism behind both of that dashboard's
unkillable wedges (see incident_esh_docker_nfs_boot_race, 2026-06-03 entry).
Re-add homepage.widget.type / .url / .slug only once the status page
actually exists. Label changes need docker compose up -d, not restart.
Data
Named volume uptimekuma_uptime-kuma → /app/data holds every monitor
definition and all history. Recreating the container is safe; deleting that
volume is not.