Wire Beszel fleet filesystems, GPU telemetry, dashboard and alerts
This commit is contained in:
@@ -0,0 +1,46 @@
|
||||
# Beszel fleet wiring — 2026-09-10
|
||||
|
||||
Operator requested `/tmp/beszel.md` handoff execution, selected **infra-ops inbox**
|
||||
for alerts (Miranda later), approved creation of a dedicated monitoring superuser,
|
||||
and asked for GPU usage/power telemetry and the card's health detail.
|
||||
|
||||
Completed: all seven requested hosts up, alongside previously registered
|
||||
corviduo-dev (8/8). nh3-docker revived; nh3-dev added. vm-esh-nas was already up,
|
||||
contrary to the handoff; access is lkraven, not infra-ops. Irvine's agent was
|
||||
healthy but its hub record still pointed at retired 10.100.79.3; fixed to
|
||||
100.64.0.6. althing-post-office container remained up throughout.
|
||||
|
||||
Docker agents need bind mounts, not merely EXTRA_FILESYSTEMS=/tank. Host
|
||||
overrides under stacks/beszel/hosts provide read-only mounts. Existing project
|
||||
directories/volumes preserved using new deploy-stack options DEPLOY_DEST_STACK
|
||||
and DEPLOY_SUDO=1. nh3-dev uses legacy docker-compose and needed the external
|
||||
traefik-net network even with agent-only profile. No host Docker upgrade.
|
||||
|
||||
ana-ml2 tank: 4548.68 / 8791.46 GiB (~51.7%). ana-docker root: ~83.1%, close
|
||||
to 85% disk warning. irv-ml1 storetank ~77.4%. NVIDIA agent 0.18.7 on both
|
||||
GPU hosts reports all four cards' utilization, VRAM and watts. No GPU power
|
||||
limits or serving workloads changed. GPU watts do not size a whole-host PSU.
|
||||
|
||||
Homepage uses existing discovery labels and version-2 widget; verified one
|
||||
card and live authenticated data. This overview shows systems/up only;
|
||||
reachability is not a degraded-health score. Per-system widget can expose
|
||||
CPU/memory/root disk/network; hub charts contain the additional disks and GPUs.
|
||||
|
||||
Approved dedicated PocketBase superuser beszel-monitoring@phasefinal.com,
|
||||
Vaultwarden ana-docker/beszel-monitoring; Homepage live .env contains its
|
||||
credential, labels only placeholders. Existing operator login unchanged.
|
||||
|
||||
Thirty rules: disk >85% for 5m, CPU >95% for 15m, memory >90% for 10m,
|
||||
offline 2m on all seven, temperature >85C for 5m on GPU hosts. Existing unused
|
||||
email route replaced with verified webhook. nh3-dev system service
|
||||
beszel-althing forwards JSON via supported postbox CLI, sender/recipient
|
||||
infra-ops; configurable recipient for later Miranda move. See service README.
|
||||
|
||||
Real alert test: ana-ml2 Disk 1%/1m fired at 15:29:45Z into althing thread
|
||||
01M25Z0WFDJM92GPTJQF769HJ7, receipt confirmed infra-ops reachable. Restored
|
||||
85%/5m afterward. Fixed hub appURL from localhost to 10.250.50.70:8090 for
|
||||
clickable alert deep links. Inbox verification did not mark mail read.
|
||||
|
||||
Still separate: ZFS degradation/SMART/scrubs and independent hub/bridge/post
|
||||
office outage detection. Bridge deliberately has no hidden delivery queue;
|
||||
downstream failure is logged and HTTP 502, not a claimed delivery.
|
||||
Reference in New Issue
Block a user