Files
esh-pfi-infrastructure/servers/esh-docker-vm/README.md
T

5.8 KiB
Raw Blame History

esh-docker-vm

General-purpose Docker host at the ESH home-lab site (esteban.net / 10.0.0.0/8 space). VM, no GPU. Separate scope from the PFI colo work but tracked here because it's part of the same fleet.

Network

  • LAN IP: 10.0.50.45
  • FQDN: esh-vm-docker.esteban.net
  • Subnet: 10.0.50.0/24 (inferred from macvlan config)
  • SSH: standard port 22
  • Traefik entrypoints: 80/443 terminate here; DNS-01 challenge via AWS Route53; protected by CrowdSec Traefik plugin

Hardware

  • CPU: 16 vCPU (QEMU virtual — VM)
  • RAM: 15.6 GB
  • GPUs: none
  • Storage: 250 GB root (ext4) + NFS mounts from 10.0.50.50
  • OS: Debian 12 (bookworm), kernel 6.1.x
  • Docker: running (daemon socket at /var/run/docker.sock; also listens on :2375)

Key paths

Path Purpose
/opt/docker/compose/<stack>/ Compose files (also a git repo)
/opt/docker/conf/<stack>/ Config bind mounts
/opt/docker/docker-bu/ Legacy backup staging (ad-hoc)
/mnt/compose NFS — 10.0.50.50:/mnt/compose
/mnt/backup NFS — 10.0.50.50:/mnt/backup (volume-backup sidecars write here)
/mnt/documents NFS — 10.0.50.50:/mnt/documents (Paperless-ngx source)
/mnt/books NFS — 10.0.50.50:/mnt/books (Calibre library)

Running stacks

Stack Port(s) Notes
traefik 80 / 443 / 8380 Reverse proxy + TLS (AWS Route53 DNS-01); CrowdSec bouncer plugin
cloudflare-ddns-esh — Keeps the esteban.net record in sync with the dynamic WAN IP
adguard 53 / 853 / 8080 / 8443 / 3000 / 784 DNS + DoT for the ESH site
dockge 5001 Docker stack management UI
homepage 5100 Dashboard (eshhome / eshhome.esteban.net)
uptime-kuma 3001 Service uptime monitor
homeassistant macvlan 10.0.50.46:8123 Home automation (direct LAN IP via macvlan on ens18)
esphome host net / 6052 ESPHome firmware dashboard
mosquitto 1883 MQTT broker
calibre-web-automated 8083 All-in-one ebook library + web UI; replaced calibre + calibre-web pair on 2026-04-20
paperless-ngx (+ redis broker + volume-backup sidecar) 8200 Document archive; Postgres on 10.0.50.60:5432
pgadmin (+ volume-backup sidecar) 5050 Postgres admin UI
drawio 8087 / 8447 Diagram editor
dozzle-agent 7007 Log agent; feeds ana-docker's Dozzle hub
beszel-esh-vm-docker (agent) 45876 Metrics agent; feeds ana-docker's Beszel hub
portainer (+ portainer_agent) 9443 / 8000 / 9001 Container management UI (ad-hoc, no compose dir under /opt/docker/compose/)

Refresh state

scripts/refresh-server-info.sh esh-docker-vm

Latest snapshot: system-details.txt (regenerate as needed).

Cross-site monitoring

Like nh3-docker, this host runs Dozzle and Beszel agents that report back to the hubs on ana-docker, so container logs and metrics show up alongside PFI hosts in the shared dashboards.

Notes

  • Macvlan for Home Assistant — the HA container gets its own LAN IP (10.0.50.46) via a macvlan network on ens18, avoiding NAT so multicast/mDNS for HA discovery works cleanly.
  • External Postgres — Paperless-ngx connects to a DB running elsewhere (10.0.50.60:5432), not a sidecar. Paperless creds in that compose file currently look like defaults; rotate before exposing.
  • Volume backups already in place — paperless-ngx and pgadmin include offen/docker-volume-backup:latest sidecars that tar named volumes to /mnt/backup/docker/esh-vm-docker/<stack>/. When the fleet-wide restic plan lands, decide whether to subsume these or leave the per-stack sidecars alone.

Placement rule

Home-lab workloads for the ESH site go here. Not part of the PFI colo topology.

⚠ Recurring: homepage wedges unkillably (3× — 2026-06-03, 2026-09-18, 2026-09-24)

The Homepage container stops answering (Uptime Kuma: timeout of 16000ms exceeded; healthcheck Connecting to 127.0.0.1:3000 times out) and cannot be stopped. Signature captured 2026-09-24 ~2224 PT, guest kernel 6.1.0-41-amd64:

  • one node thread in D state in vm_mmap_pgoff — waiting for its own process's mmap_lock for write, with no visible holder (every other thread sat in futex_wait; a scan of every task's kernel stack found no reader in a fault, NFS or access_remote_vm path except the ps calls queued behind it).
  • docker restart → "tried to kill container, but did not receive an exit event"; the process then sits in exit_mmap (uninterruptible) with PID 1 of the container in zap_pid_ns_processes. The June entry in archival-memory records the same exit_mmap end state.
  • The rest of the VM is fine: DNS (AdGuard) answered, 17 other containers up.

⚠ Diagnosing it can hang your shell. ps, pgrep and docker top read /proc/<pid>/cmdline|environ, which takes the same lock, so they block in D state too. Read /proc/<pid>/task/*/stat and sudo cat .../stack instead (ssh infra-ops@10.0.50.45, NOPASSWD).

Only a VM reboot clears it, and this VM is ESH's DNS resolver, so the reboot is a short ESH-wide DNS outage — schedule it.

⚠ After the reboot, homepage does NOT come back by itself if anyone tried docker restart/stop on it first: docker records it as stopped on purpose, so restart: unless-stopped skips it. Run docker start homepage. Measured 2026-09-25: qm reboot 100 from esh-pve took 3m18s (the graceful shutdown stalls on the wedged process, then proceeds); the VM was unreachable 0613:44–0616:58 PT; every other container came back unaided. Root cause not established; a kernel-side mmap_lock problem is the leading guess, not a finding. Moving Homepage to ana-docker (as was done for Uptime Kuma on 2026-09-21 for the same box's history) would take the dashboard out of this failure domain.