Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-10-beszel-fleet-wiring.md
T

2.7 KiB

Beszel fleet wiring — 2026-09-10

Operator requested /tmp/beszel.md handoff execution, selected infra-ops inbox for alerts (Miranda later), approved creation of a dedicated monitoring superuser, and asked for GPU usage/power telemetry and the card's health detail.

Completed: all seven requested hosts up, alongside previously registered corviduo-dev (8/8). nh3-docker revived; nh3-dev added. vm-esh-nas was already up, contrary to the handoff; access is lkraven, not infra-ops. Irvine's agent was healthy but its hub record still pointed at retired 10.100.79.3; fixed to 100.64.0.6. althing-post-office container remained up throughout.

Docker agents need bind mounts, not merely EXTRA_FILESYSTEMS=/tank. Host overrides under stacks/beszel/hosts provide read-only mounts. Existing project directories/volumes preserved using new deploy-stack options DEPLOY_DEST_STACK and DEPLOY_SUDO=1. nh3-dev uses legacy docker-compose and needed the external traefik-net network even with agent-only profile. No host Docker upgrade.

ana-ml2 tank: 4548.68 / 8791.46 GiB (~51.7%). ana-docker root: ~83.1%, close to 85% disk warning. irv-ml1 storetank ~77.4%. NVIDIA agent 0.18.7 on both GPU hosts reports all four cards' utilization, VRAM and watts. No GPU power limits or serving workloads changed. GPU watts do not size a whole-host PSU.

Homepage uses existing discovery labels and version-2 widget; verified one card and live authenticated data. This overview shows systems/up only; reachability is not a degraded-health score. Per-system widget can expose CPU/memory/root disk/network; hub charts contain the additional disks and GPUs.

Approved dedicated PocketBase superuser beszel-monitoring@phasefinal.com, Vaultwarden ana-docker/beszel-monitoring; Homepage live .env contains its credential, labels only placeholders. Existing operator login unchanged.

Thirty rules: disk >85% for 5m, CPU >95% for 15m, memory >90% for 10m, offline 2m on all seven, temperature >85C for 5m on GPU hosts. Existing unused email route replaced with verified webhook. nh3-dev system service beszel-althing forwards JSON via supported postbox CLI, sender/recipient infra-ops; configurable recipient for later Miranda move. See service README.

Real alert test: ana-ml2 Disk 1%/1m fired at 15:29:45Z into althing thread 01M25Z0WFDJM92GPTJQF769HJ7, receipt confirmed infra-ops reachable. Restored 85%/5m afterward. Fixed hub appURL from localhost to 10.250.50.70:8090 for clickable alert deep links. Inbox verification did not mark mail read.

Still separate: ZFS degradation/SMART/scrubs and independent hub/bridge/post office outage detection. Bridge deliberately has no hidden delivery queue; downstream failure is logged and HTTP 502, not a claimed delivery.