Files
esh-pfi-infrastructure/CLAUDE.md
T
vh 0be8de8ab0 fleet: register vm-esh-nas as 5th Docker host + canonical dockge stack
vm-esh-nas (10.0.50.154) is a NAS-adjacent Docker VM on the esh-pve-nas
hypervisor. Runs filezilla (port 5800), dockge, beszel-agent, dozzle-agent
with /mnt/{share,music,books,media} NFS-mounted from 10.0.50.50.
Use this host when a stack needs direct NFS mounts to the ESH NAS shares.

Canonicalize dockge as stacks/dockge/ — single compose used on all five
Docker hosts with per-host DOCKGE_HOST_LABEL/DOCKGE_HOST_IP in .env so
each card on the homepage points at its own instance. Labeled
homepage.group=Service Networking.

Beszel + Dozzle agent dirs also renamed to beszel-agent-<site> /
dozzle-agent-<site> pattern across the fleet for consistency.
2026-04-20 22:14:13 -07:00

157 lines
7.8 KiB
Markdown

# CLAUDE.md
This workspace is for managing PFI infrastructure — servers, Docker stacks, and related configs. Spawn a dedicated Claude Code session here when working on infra so it doesn't clutter AIPA-MCP development context.
## Purpose
- Inventory of servers and their state
- Canonical copies of Docker Compose stacks deployed on those servers
- Scripts for inspecting and managing the infrastructure
- Conventions so all stacks look the same
This is a **reference workspace** — the authoritative copies of compose files and configs live **on the servers** under `/opt/docker/compose/<stack>/` and `/opt/docker/conf/<stack>/`. This workspace mirrors them for version control, editing, and planning.
## Conventions (enforce for every new stack)
Observed and standardized across servers:
- **Compose location on server:** `/opt/docker/compose/<stack>/compose.yaml`
- **Config mounts on server:** `/opt/docker/conf/<stack>/...`
- **Networks:** external `traefik-net`, aliased as `tnet` in compose
```yaml
networks:
tnet:
name: traefik-net
external: true
```
- **GPU reservation:** prefer `deploy.resources.reservations.devices` with explicit `device_ids` for pinning
```yaml
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["1"]
capabilities: [gpu]
```
- **Tunables:** `.env` in the same directory as `compose.yaml` — keep the compose file constant, edit the `.env`
- **Named volumes** for service state (pattern: `<stack>_<name>`)
- **Bind mounts** only for: model files (`/tank/aimodels/...`), config files (`/opt/docker/conf/...`), docker socket where required
- **Restart policy:** `restart: unless-stopped` for daemons
- **Homepage labels** on user-facing services:
```yaml
labels:
- homepage.group=AI Systems
- homepage.name=<ServiceName>
- homepage.icon=mdi-<icon>
- homepage.description=<short>
- homepage.href=http://<host-ip>:<port>
```
- **Healthchecks** on services that expose HTTP
## Servers
| Name | IP | Site | Role | Details |
|------|-----|------|------|---------|
| ana-ml2 | 10.250.50.54 | Anaheim (`10.250.0.0/16`) | GPU / AI inference (bare metal) | `servers/ana-ml2/README.md` |
| ana-docker | 10.250.50.70 | Anaheim | General-purpose Docker host (non-GPU VM on pfi-pve) | `servers/ana-docker/README.md` |
| pfi-pve | 10.250.250.31 | Anaheim | Proxmox VE hypervisor | `servers/pfi-pve/README.md` |
| nh3-docker | 10.100.50.40 | NH3 (`10.100.0.0/16`) | General-purpose Docker host (non-GPU VM on nh3-pve) | `servers/nh3-docker/README.md` |
| nh3-pve | 10.100.250.60 | NH3 | Proxmox VE hypervisor | `servers/nh3-pve/README.md` |
| esh-docker-vm | 10.0.50.45 | ESH home lab (`esteban.net`, `10.0.50.0/24`) | Home-lab Docker host (VM on esh-pve) | `servers/esh-docker-vm/README.md` |
| vm-esh-nas | 10.0.50.154 | ESH home lab | NAS-adjacent Docker host (VM on esh-pve-nas) | `servers/vm-esh-nas/README.md` |
| esh-pve | 10.0.250.35 | ESH home lab | Proxmox VE hypervisor | `servers/esh-pve/README.md` |
| esh-pve-nas | 10.0.50.55 | ESH home lab | Proxmox VE hypervisor (storage / media) | `servers/esh-pve-nas/README.md` |
**Placement rules:**
- GPU-required stacks → `ana-ml2`.
- Anaheim non-GPU services → `ana-docker`.
- NH-site non-GPU services → `nh3-docker`.
- ESH home-lab workloads (`esteban.net`) → `esh-docker-vm` (general) or `vm-esh-nas` (needs direct NFS mounts from 10.0.50.50). Not part of the PFI colo topology, but shares monitoring/backup tooling.
- Cross-site services (e.g. Beszel hub, Dozzle hub) live on `ana-docker` and pull from agents on the other hosts.
- **Hypervisors** (`pfi-pve`, `nh3-pve`, `esh-pve`, `esh-pve-nas`) are tracked for inventory / capacity planning. Don't deploy Docker stacks directly on them; new workloads land as VMs. `server_inspect.sh` captures host-level detail only — VM/LXC/ZFS enumeration needs Proxmox-native tooling (`qm list`, `pvesh get …`, `zpool list`).
## How to refresh a server's state
```bash
# Show help (no args)
scripts/refresh-server-info.sh
# Refresh every host discovered under servers/*/
scripts/refresh-server-info.sh all
# Refresh a specific host (must match a servers/<name>/ dir; ssh_config
# entry or servers/<name>/ssh-target handles how to reach it)
scripts/refresh-server-info.sh ana-docker
```
Fleet-wide runs require the literal `all` keyword — no-args prints help so you can't accidentally hit every host by forgetting a name.
The script pipes `server_inspect.sh` over SSH via stdin (no scp, no remote cleanup) and writes each `servers/<host>/system-details.txt` atomically — a failed run never clobbers the previous snapshot. The inspect script itself is read-only.
Each server dir can hold an `ssh-target` file (one line, `<ip>` or `<user>@<ip>`) as a fallback for when the dir name doesn't resolve via DNS or `~/.ssh/config`. The script prefers whatever ssh would resolve normally and only consults the file when that fails.
To register a new server:
```bash
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name> # pull the first snapshot
```
To audit discovery without touching the network (checks permissions, unresolvable names with no fallback, missing README / system-details, malformed `ssh-target`):
```bash
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh --validate-only <host>
```
## Stack mirror (pull / push)
Compose and config trees are mirrored into `stacks-mirror/<host>/<stack>/` so they can be diffed and version-controlled. Pull is fleet-wide and safe; push is one stack at a time with a diff + prompt.
```bash
# Pull compose + conf from every host into stacks-mirror/
scripts/sync-stacks.sh
scripts/sync-stacks.sh --dry-run # see what would change
scripts/sync-stacks.sh ana-docker # one host
# Push a local stack back to the server (diffs each file, prompts y/N)
scripts/deploy-stack.sh <host> <stack>
scripts/deploy-stack.sh <host> <stack> --compose # skip conf
scripts/deploy-stack.sh <host> <stack> --conf # skip compose
```
**Opt-out per stack:** create `stacks-mirror/<host>/<stack>/.no-sync` (skip both sides) or `stacks-mirror/<host>/<stack>/conf/.no-sync` (skip conf only).
**Always excluded in both directions** (secrets / runtime state): `.env`, `.env.*`, `acme.json`, `client_secrets.json`, `*.pem`, `*.key`, `*.crt`, `*.pfx`, `*.sqlite`, `*.sqlite3`, `*.db`, `*.log`, `*.log.*`, `*.pid`, `hub/`, `logs/`.
Requires `rsync` installed on this workstation and every host you sync against (`apt install rsync`).
## Layout
```
eshpfi-management/
├── CLAUDE.md # this file
├── README.md # human-facing overview
├── scripts/
│ └── server_inspect.sh # gather server state for compose planning
├── servers/
│ └── <name>/
│ ├── README.md
│ └── system-details.txt # latest server_inspect output
├── stacks/
│ └── <stack>/
│ ├── compose.yaml # deployed to /opt/docker/compose/<stack>/
│ ├── .env.example # template; real .env lives on server
│ └── README.md # what this stack does, how to deploy
└── docs/
└── pfi/ # general PFI infrastructure reference
```
## Working rules
- **Copies, not symlinks.** Files here reflect what's on the server at the time of the last sync. When you edit here, the server doesn't change until you deploy.
- **Never commit secrets.** Use `.env.example` templates; real `.env` files (with tokens, passwords) live on the server and are gitignored if/when this becomes a git repo.
- **Surgical edits.** When fixing one stack, don't touch unrelated ones. Follow AIPA-MCP's CLAUDE.md rules about scope discipline.
- **Sanity-check before deploying.** Run `docker compose config` (dry parse) before `docker compose up -d` on the server.