feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list

- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system
  registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples
  verified (RTX 2000E Ada util, VRAM, power, temp).
- Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013)
  services — the sole backends behind the gateway, so the one exception to
  "seats are out of Kuma's lane". Reward seat excluded (no consumer).
- Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step);
  esh-ml1-docker added to docker.yaml.
- Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still
  named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed
  to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing).
- Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an
  automounted NAS share); revived by hand, fix still open.
This commit is contained in:
vh
2026-09-25 09:20:36 -07:00
parent 52612cbe96
commit 65dc586497
9 changed files with 165 additions and 20 deletions
+15 -12
View File
@@ -136,17 +136,20 @@ for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in)
needs that series' headers meta-package installed first, or the module will not
build and this CT will fail to start at the next boot.
## Not yet wired
## Monitoring and telemetry (wired 2026-09-25)
⚠ **esh-ml1 became load-bearing on 2026-09-25, so these are no longer
optional.** Nothing alerts if it dies today. The gateway will just start
returning errors for `qwen3-embedding` and `reranker`.
| layer | what | where |
|---|---|---|
| **Beszel** (host + GPU telemetry) | NVIDIA agent `henrygd/beszel-agent-nvidia:0.18.7`, `stacks/beszel` + `hosts/esh-ml1.yaml`. Hub system `ridkfdpwfq3f730`. GPU util, VRAM, power and temperature are sampled. | alerts → infra-ops: Status down 2 m, Disk >85% 5 m, CPU >95% 15 m, Memory >90% 10 m, Temperature >85 °C 5 m (GPU included) |
| **Uptime Kuma** (service) | `Embed — Qwen3 0.6B (TEI, esh-ml1)` → `:8001/health` (#27); `Rerank — bge-v2-m3 (TEI, esh-ml1)` → `:8013/health` (#28). TEI's health runs the backend. | alerts → infra-ops via althing-alert-bridge; `stacks/uptimekuma/monitors.yaml` |
| **Homepage** | three cards under *AI - Eval & Retrieval*. dockerd exposes tcp/2375 on `10.0.50.80` only (fleet norm, `playbooks/esh-ml1-lxc.yaml`). | `stacks/homepage/conf/docker.yaml` → `esh-ml1-docker` |
| **Dozzle** (logs) | agent `v10.4.1` on `10.0.50.80:7007`, compose dir `dozzle-agent` | hub on ana-docker :8088 |
- **Beszel**: no agent yet.
- **Uptime Kuma**: no check on `:8001/health` / `:8013/health` yet.
- **Homepage**: the compose carries labels, but esh-ml1 is not in
`stacks/homepage/conf/docker.yaml` (it would need dockerd on tcp/2375 like
the other hosts).
- **Path:** every consumer reaches esh-ml1 through the gateway at ana-docker, so
ESH callers hairpin ESH → Anaheim → ESH over the mesh. A mesh outage cuts
every consumer off, the ones at ESH included.
The reward seat has **no Kuma check on purpose**: seats are outside Kuma's lane,
and this one has no working consumer. Beszel and Homepage cover it.
⚠ Disk: rootfs is ~66% used (vLLM image ~30 GB, TEI ~8 GB). The alert fires at 85%.
**Path:** every consumer reaches esh-ml1 through the gateway at ana-docker, so
ESH callers hairpin ESH → Anaheim → ESH over the mesh. A mesh outage cuts every
consumer off, the ones at ESH included.