feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list

- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system
  registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples
  verified (RTX 2000E Ada util, VRAM, power, temp).
- Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013)
  services — the sole backends behind the gateway, so the one exception to
  "seats are out of Kuma's lane". Reward seat excluded (no consumer).
- Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step);
  esh-ml1-docker added to docker.yaml.
- Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still
  named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed
  to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing).
- Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an
  automounted NAS share); revived by hand, fix still open.
This commit is contained in:
vh
2026-09-25 09:20:36 -07:00
parent 52612cbe96
commit 65dc586497
9 changed files with 165 additions and 20 deletions
+22
View File
@@ -32,6 +32,7 @@ vars:
ctid: 110
hostname: esh-ml1
ip_cidr: 10.0.50.80/24
ip_addr: 10.0.50.80
gateway: 10.0.50.1
vlan: 50
cores: 6
@@ -156,6 +157,27 @@ steps:
EOF
when: "! pct exec {{ ctid }} -- sh -c 'grep -q nvidia /etc/docker/daemon.json && grep -Eq \"^no-cgroups *= *true\" /etc/nvidia-container-runtime/config.toml' 2>/dev/null"
# Fleet norm: Homepage (on esh-docker-vm) discovers labelled containers by
# reading every host's Docker API on tcp/2375 (stacks/homepage/conf/docker.yaml).
# Same unauthenticated plaintext exposure as fv-ml1 and esh-docker-vm, bound to
# this CT's one address. ⚠ Restarting dockerd restarts every container here,
# the fleet embed/rerank service included (~5 s for TEI, ~60 s for the reward seat).
- name: Expose the Docker API on tcp/2375 for Homepage discovery
shell: |
pct exec {{ ctid }} -- bash -s <<'EOF'
set -euo pipefail
install -d /etc/systemd/system/docker.service.d
cat > /etc/systemd/system/docker.service.d/override.conf <<'EOC'
# Homepage discovery — see eshpfi playbooks/esh-ml1-lxc.yaml
[Service]
ExecStart=
ExecStart=/usr/bin/dockerd -H fd:// -H tcp://{{ ip_addr }}:2375 --containerd=/run/containerd/containerd.sock
EOC
systemctl daemon-reload
systemctl restart docker
EOF
when: "! pct exec {{ ctid }} -- grep -q 'tcp://' /etc/systemd/system/docker.service.d/override.conf 2>/dev/null"
verify:
- name: Container is running with onboot set, started after the core guests
shell: "pct status {{ ctid }} | grep -q running && pct config {{ ctid }} | grep -q '^onboot: 1' && pct config {{ ctid }} | grep -q '^startup: order=30'"