Wire Beszel fleet filesystems, GPU telemetry, dashboard and alerts

This commit is contained in:
vh
2026-09-10 08:35:28 -07:00
parent 20bbb95113
commit eb75713c1b
18 changed files with 438 additions and 84 deletions
@@ -0,0 +1,46 @@
# Beszel fleet wiring — 2026-09-10
Operator requested `/tmp/beszel.md` handoff execution, selected **infra-ops inbox**
for alerts (Miranda later), approved creation of a dedicated monitoring superuser,
and asked for GPU usage/power telemetry and the card's health detail.
Completed: all seven requested hosts up, alongside previously registered
corviduo-dev (8/8). nh3-docker revived; nh3-dev added. vm-esh-nas was already up,
contrary to the handoff; access is lkraven, not infra-ops. Irvine's agent was
healthy but its hub record still pointed at retired 10.100.79.3; fixed to
100.64.0.6. althing-post-office container remained up throughout.
Docker agents need bind mounts, not merely EXTRA_FILESYSTEMS=/tank. Host
overrides under stacks/beszel/hosts provide read-only mounts. Existing project
directories/volumes preserved using new deploy-stack options DEPLOY_DEST_STACK
and DEPLOY_SUDO=1. nh3-dev uses legacy docker-compose and needed the external
traefik-net network even with agent-only profile. No host Docker upgrade.
ana-ml2 tank: 4548.68 / 8791.46 GiB (~51.7%). ana-docker root: ~83.1%, close
to 85% disk warning. irv-ml1 storetank ~77.4%. NVIDIA agent 0.18.7 on both
GPU hosts reports all four cards' utilization, VRAM and watts. No GPU power
limits or serving workloads changed. GPU watts do not size a whole-host PSU.
Homepage uses existing discovery labels and version-2 widget; verified one
card and live authenticated data. This overview shows systems/up only;
reachability is not a degraded-health score. Per-system widget can expose
CPU/memory/root disk/network; hub charts contain the additional disks and GPUs.
Approved dedicated PocketBase superuser beszel-monitoring@phasefinal.com,
Vaultwarden ana-docker/beszel-monitoring; Homepage live .env contains its
credential, labels only placeholders. Existing operator login unchanged.
Thirty rules: disk >85% for 5m, CPU >95% for 15m, memory >90% for 10m,
offline 2m on all seven, temperature >85C for 5m on GPU hosts. Existing unused
email route replaced with verified webhook. nh3-dev system service
beszel-althing forwards JSON via supported postbox CLI, sender/recipient
infra-ops; configurable recipient for later Miranda move. See service README.
Real alert test: ana-ml2 Disk 1%/1m fired at 15:29:45Z into althing thread
01M25Z0WFDJM92GPTJQF769HJ7, receipt confirmed infra-ops reachable. Restored
85%/5m afterward. Fixed hub appURL from localhost to 10.250.50.70:8090 for
clickable alert deep links. Inbox verification did not mark mail read.
Still separate: ZFS degradation/SMART/scrubs and independent hub/bridge/post
office outage detection. Bridge deliberately has no hidden delivery queue;
downstream failure is logged and HTTP 502, not a claimed delivery.