feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list
- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples verified (RTX 2000E Ada util, VRAM, power, temp). - Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013) services — the sole backends behind the gateway, so the one exception to "seats are out of Kuma's lane". Reward seat excluded (no consumer). - Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step); esh-ml1-docker added to docker.yaml. - Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing). - Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an automounted NAS share); revived by hand, fix still open.
This commit is contained in:
+15
-12
@@ -136,17 +136,20 @@ for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in)
|
||||
needs that series' headers meta-package installed first, or the module will not
|
||||
build and this CT will fail to start at the next boot.
|
||||
|
||||
## Not yet wired
|
||||
## Monitoring and telemetry (wired 2026-09-25)
|
||||
|
||||
⚠ **esh-ml1 became load-bearing on 2026-09-25, so these are no longer
|
||||
optional.** Nothing alerts if it dies today. The gateway will just start
|
||||
returning errors for `qwen3-embedding` and `reranker`.
|
||||
| layer | what | where |
|
||||
|---|---|---|
|
||||
| **Beszel** (host + GPU telemetry) | NVIDIA agent `henrygd/beszel-agent-nvidia:0.18.7`, `stacks/beszel` + `hosts/esh-ml1.yaml`. Hub system `ridkfdpwfq3f730`. GPU util, VRAM, power and temperature are sampled. | alerts → infra-ops: Status down 2 m, Disk >85% 5 m, CPU >95% 15 m, Memory >90% 10 m, Temperature >85 °C 5 m (GPU included) |
|
||||
| **Uptime Kuma** (service) | `Embed — Qwen3 0.6B (TEI, esh-ml1)` → `:8001/health` (#27); `Rerank — bge-v2-m3 (TEI, esh-ml1)` → `:8013/health` (#28). TEI's health runs the backend. | alerts → infra-ops via althing-alert-bridge; `stacks/uptimekuma/monitors.yaml` |
|
||||
| **Homepage** | three cards under *AI - Eval & Retrieval*. dockerd exposes tcp/2375 on `10.0.50.80` only (fleet norm, `playbooks/esh-ml1-lxc.yaml`). | `stacks/homepage/conf/docker.yaml` → `esh-ml1-docker` |
|
||||
| **Dozzle** (logs) | agent `v10.4.1` on `10.0.50.80:7007`, compose dir `dozzle-agent` | hub on ana-docker :8088 |
|
||||
|
||||
- **Beszel**: no agent yet.
|
||||
- **Uptime Kuma**: no check on `:8001/health` / `:8013/health` yet.
|
||||
- **Homepage**: the compose carries labels, but esh-ml1 is not in
|
||||
`stacks/homepage/conf/docker.yaml` (it would need dockerd on tcp/2375 like
|
||||
the other hosts).
|
||||
- **Path:** every consumer reaches esh-ml1 through the gateway at ana-docker, so
|
||||
ESH callers hairpin ESH → Anaheim → ESH over the mesh. A mesh outage cuts
|
||||
every consumer off, the ones at ESH included.
|
||||
The reward seat has **no Kuma check on purpose**: seats are outside Kuma's lane,
|
||||
and this one has no working consumer. Beszel and Homepage cover it.
|
||||
|
||||
⚠ Disk: rootfs is ~66% used (vLLM image ~30 GB, TEI ~8 GB). The alert fires at 85%.
|
||||
|
||||
**Path:** every consumer reaches esh-ml1 through the gateway at ana-docker, so
|
||||
ESH callers hairpin ESH → Anaheim → ESH over the mesh. A mesh outage cuts every
|
||||
consumer off, the ones at ESH included.
|
||||
|
||||
Reference in New Issue
Block a user