feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list
- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples verified (RTX 2000E Ada util, VRAM, power, temp). - Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013) services — the sole backends behind the gateway, so the one exception to "seats are out of Kuma's lane". Reward seat excluded (no consumer). - Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step); esh-ml1-docker added to docker.yaml. - Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing). - Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an automounted NAS share); revived by hand, fix still open.
This commit is contained in:
@@ -32,6 +32,7 @@ vars:
|
||||
ctid: 110
|
||||
hostname: esh-ml1
|
||||
ip_cidr: 10.0.50.80/24
|
||||
ip_addr: 10.0.50.80
|
||||
gateway: 10.0.50.1
|
||||
vlan: 50
|
||||
cores: 6
|
||||
@@ -156,6 +157,27 @@ steps:
|
||||
EOF
|
||||
when: "! pct exec {{ ctid }} -- sh -c 'grep -q nvidia /etc/docker/daemon.json && grep -Eq \"^no-cgroups *= *true\" /etc/nvidia-container-runtime/config.toml' 2>/dev/null"
|
||||
|
||||
# Fleet norm: Homepage (on esh-docker-vm) discovers labelled containers by
|
||||
# reading every host's Docker API on tcp/2375 (stacks/homepage/conf/docker.yaml).
|
||||
# Same unauthenticated plaintext exposure as fv-ml1 and esh-docker-vm, bound to
|
||||
# this CT's one address. ⚠ Restarting dockerd restarts every container here,
|
||||
# the fleet embed/rerank service included (~5 s for TEI, ~60 s for the reward seat).
|
||||
- name: Expose the Docker API on tcp/2375 for Homepage discovery
|
||||
shell: |
|
||||
pct exec {{ ctid }} -- bash -s <<'EOF'
|
||||
set -euo pipefail
|
||||
install -d /etc/systemd/system/docker.service.d
|
||||
cat > /etc/systemd/system/docker.service.d/override.conf <<'EOC'
|
||||
# Homepage discovery — see eshpfi playbooks/esh-ml1-lxc.yaml
|
||||
[Service]
|
||||
ExecStart=
|
||||
ExecStart=/usr/bin/dockerd -H fd:// -H tcp://{{ ip_addr }}:2375 --containerd=/run/containerd/containerd.sock
|
||||
EOC
|
||||
systemctl daemon-reload
|
||||
systemctl restart docker
|
||||
EOF
|
||||
when: "! pct exec {{ ctid }} -- grep -q 'tcp://' /etc/systemd/system/docker.service.d/override.conf 2>/dev/null"
|
||||
|
||||
verify:
|
||||
- name: Container is running with onboot set, started after the core guests
|
||||
shell: "pct status {{ ctid }} | grep -q running && pct config {{ ctid }} | grep -q '^onboot: 1' && pct config {{ ctid }} | grep -q '^startup: order=30'"
|
||||
|
||||
Reference in New Issue
Block a user