feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list

- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system
  registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples
  verified (RTX 2000E Ada util, VRAM, power, temp).
- Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013)
  services — the sole backends behind the gateway, so the one exception to
  "seats are out of Kuma's lane". Reward seat excluded (no consumer).
- Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step);
  esh-ml1-docker added to docker.yaml.
- Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still
  named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed
  to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing).
- Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an
  automounted NAS share); revived by hand, fix still open.
This commit is contained in:
vh
2026-09-25 09:20:36 -07:00
parent 52612cbe96
commit 65dc586497
9 changed files with 165 additions and 20 deletions
+10
View File
@@ -23,6 +23,16 @@ filesystem samples verified; fleet 13/14 up with known fv-ml1 outage.
| irv-ml1 | beszel-agent-irv | /worktank, /storetank, /mnt/smithy |
| vm-esh-nas | beszel-agent-esh-nas | /mnt/books, /mnt/share, /mnt/music, /mnt/media |
| nh3-dev | beszel | /mnt/backup, /mnt/smithy |
| esh-ml1 (added 2026-09-25) | beszel | none — NVIDIA image (`hosts/esh-ml1.yaml`) for the RTX 2000E Ada |
⚠ **nh3-dev's agent was DOWN from the 2026-09-24 NH3 power recovery until
2026-09-25.** Docker could not bind `/mnt/smithy` at boot ("no such device"): since
`1cbde50` the NAS shares are automounted, and nh3-nas was not up yet. Docker does not
retry a container that fails to *create*, so `unless-stopped` never brought it back.
Started by hand. **It will recur on the next NH3 cold start** until the agent stops
depending on NAS mounts at boot. The NAS capacity is already reported by
nh3-nas's own agent, so dropping the two extra filesystems here is the simplest fix.
Open item.
Use `infra-ops@<ip>` with passwordless sudo, except vm-esh-nas:
`lkraven@10.0.50.154` has Docker access. Irvine's hub address is
+14
View File
@@ -0,0 +1,14 @@
# esh-ml1 (CT 110 on esh-pve) — Beszel agent with NVIDIA GPU telemetry for the
# RTX 2000E Ada (utilization, VRAM, temperature, power). Docker-in-LXC: the
# NVIDIA container toolkit runs with no-cgroups=true (playbooks/esh-ml1-lxc.yaml).
# No extra filesystems: the root filesystem holds everything, models included.
services:
beszel-agent:
image: henrygd/beszel-agent-nvidia:0.18.7
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [utility]
+12 -2
View File
@@ -3,9 +3,19 @@
Container log viewer. One UI on **ana-docker** aggregates logs from every Docker host via remote agents.
**Deploys to:**
- **ana-docker** (hub) — UI at `http://10.250.50.70:8088`
- **ana-docker** (hub) — UI at `http://10.250.50.70:8088`; live compose dir `/opt/docker/compose/dozzle-hub`
- **fv-ml1** (agent) — listens on `10.251.50.54:7007`
- **nh3-docker** (agent, cross-site) — listens on `10.100.50.40:7007`
- **esh-docker-vm** / **vm-esh-nas** (agents) — `10.0.50.45:7007`, `10.0.50.154:7007`
- **irv-ml1** (agent) — `10.6.110.50:7007` (mesh address)
- **esh-ml1** (agent, added 2026-09-25) — `10.0.50.80:7007`, pinned `v10.4.1` = the hub's version; compose dir `dozzle-agent`. The host needed an (empty) `traefik-net` network because the compose declares it external.
- **nh3-docker** (agent, cross-site) — `10.100.50.40:7007`. ⚠ **Stopped by hand ~2026-04 (Exited 0) and left that way**; the hub logs a refused connection for it. Revive it or drop it from the list deliberately.
⚠ **The hub's agent list silently rots when a host moves.** Until 2026-09-25 it
still named ana-ml2's `10.250.50.54` and irv-ml1's retired `10.100.79.3`, so fv-ml1
and irv-ml1 logs had been missing from Dozzle since their moves (2026-09-12 and
2026-09-06). Nothing reported it. After any re-IP, check the hub's
`DOZZLE_REMOTE_AGENT`, and check that `docker logs dozzle` shows `"clients":N` equal to
reachable agents + 1.
- **corviduo-dev** (agent) — listens on `10.250.50.152:7007`. Compose at `/home/vh/docker/compose/dozzle-agent/` (not `/opt/docker/compose/` — see `servers/corviduo-dev/README.md` for why)
One compose.yaml lives on each host. The per-host `.env` sets `COMPOSE_PROFILES=hub` or `COMPOSE_PROFILES=agent` so `docker compose up -d` brings up the right service. On the hub, add every agent to `DOZZLE_REMOTE_AGENT` as a comma-separated list (e.g. `10.251.50.54:7007,10.100.50.40:7007`).
+6
View File
@@ -31,6 +31,12 @@ irv-ml1-docker:
host: 10.6.110.50
port: 2375
# esh-ml1 — CT 110 on esh-pve: the fleet embed/rerank (TEI) + reward seats.
# dockerd listens only on its own address (playbooks/esh-ml1-lxc.yaml), 2026-09-25.
esh-ml1-docker:
host: 10.0.50.80
port: 2375
# Example TLS socket (if/when a host moves off plaintext 2375):
# ana-pfi-docker:
# host: 10.250.50.70
+15
View File
@@ -92,6 +92,21 @@ monitors:
- name: Asset Engine
url: http://10.250.50.70:8200/api/v1/services
# ---- fleet embed/rerank: the one EXCEPTION to "seats are OUT" ----
# Since 2026-09-25 these are the SOLE backends behind the gateway's
# `qwen3-embedding` and `reranker` (TEI on esh-ml1, no failover until the second
# RTX 2000 arrives). They are not come-and-go seats: dead = Worldtree recall,
# nevermore clustering and Open WebUI RAG all fail. TEI's /health runs the
# backend, so a loaded-but-broken model reads DOWN, not UP. The reward seat on the
# same box stays OUT: it has no working consumer (stacks/reward-seat/README.md).
- name: Embed — Qwen3 0.6B (TEI, esh-ml1)
url: http://10.0.50.80:8001/health
description: sole backend for gateway `qwen3-embedding`
- name: Rerank — bge-v2-m3 (TEI, esh-ml1)
url: http://10.0.50.80:8013/health
description: sole backend for gateway `reranker`
- name: talk
url: https://talk.nh3.phasefinal.com:8092/
description: fleet voice bench