# CLAUDE.md This workspace is for managing PFI infrastructure — servers, Docker stacks, and related configs. Spawn a dedicated Claude Code session here when working on infra so it doesn't clutter AIPA-MCP development context. ## Purpose - Inventory of servers and their state - Canonical copies of Docker Compose stacks deployed on those servers - Scripts for inspecting and managing the infrastructure - Conventions so all stacks look the same This is a **reference workspace** — the authoritative copies of compose files and configs live **on the servers** under `/opt/docker/compose//` and `/opt/docker/conf//`. This workspace mirrors them for version control, editing, and planning. ## Conventions (enforce for every new stack) Observed and standardized across servers: - **Compose location on server:** `/opt/docker/compose//compose.yaml` - **Config mounts on server:** `/opt/docker/conf//...` - **Networks:** external `traefik-net`, aliased as `tnet` in compose ```yaml networks: tnet: name: traefik-net external: true ``` - **GPU reservation:** prefer `deploy.resources.reservations.devices` with explicit `device_ids` for pinning ```yaml deploy: resources: reservations: devices: - driver: nvidia device_ids: ["1"] capabilities: [gpu] ``` - **Tunables:** `.env` in the same directory as `compose.yaml` — keep the compose file constant, edit the `.env` - **Named volumes** for service state (pattern: `_`) - **Bind mounts** only for: model files (`/tank/aimodels/...`), config files (`/opt/docker/conf/...`), docker socket where required - **Restart policy:** `restart: unless-stopped` for daemons - **Homepage labels** on user-facing services: ```yaml labels: - homepage.group=AI Systems - homepage.name= - homepage.icon=mdi- - homepage.description= - homepage.href=http://: ``` - **Healthchecks** on services that expose HTTP ## Servers | Name | IP | Site | Role | Details | |------|-----|------|------|---------| | ana-ml2 | 10.250.50.54 | Anaheim (`10.250.0.0/16`) | GPU / AI inference (bare metal) | `servers/ana-ml2/README.md` | | ana-docker | 10.250.50.70 | Anaheim | General-purpose Docker host (non-GPU VM on pfi-pve) | `servers/ana-docker/README.md` | | pfi-pve | 10.250.250.31 | Anaheim | Proxmox VE hypervisor | `servers/pfi-pve/README.md` | | nh3-docker | 10.100.50.40 | NH3 (`10.100.0.0/16`) | General-purpose Docker host (non-GPU VM on nh3-pve) | `servers/nh3-docker/README.md` | | nh3-pve | 10.100.250.60 | NH3 | Proxmox VE hypervisor | `servers/nh3-pve/README.md` | | esh-docker-vm | 10.0.50.45 | ESH home lab (`esteban.net`, `10.0.50.0/24`) | Home-lab Docker host (VM on esh-pve) | `servers/esh-docker-vm/README.md` | | vm-esh-nas | 10.0.50.154 | ESH home lab | NAS-adjacent Docker host (VM on esh-pve-nas) | `servers/vm-esh-nas/README.md` | | esh-pve | 10.0.250.35 | ESH home lab | Proxmox VE hypervisor | `servers/esh-pve/README.md` | | esh-pve-nas | 10.0.50.55 | ESH home lab | Proxmox VE hypervisor (storage / media) | `servers/esh-pve-nas/README.md` | **Placement rules:** - GPU-required stacks → `ana-ml2`. - Anaheim non-GPU services → `ana-docker`. - NH-site non-GPU services → `nh3-docker`. - ESH home-lab workloads (`esteban.net`) → `esh-docker-vm` (general) or `vm-esh-nas` (needs direct NFS mounts from 10.0.50.50). Not part of the PFI colo topology, but shares monitoring/backup tooling. - Cross-site services (e.g. Beszel hub, Dozzle hub) live on `ana-docker` and pull from agents on the other hosts. - **Hypervisors** (`pfi-pve`, `nh3-pve`, `esh-pve`, `esh-pve-nas`) are tracked for inventory / capacity planning. Don't deploy Docker stacks directly on them; new workloads land as VMs. `server_inspect.sh` captures host-level detail only — VM/LXC/ZFS enumeration needs Proxmox-native tooling (`qm list`, `pvesh get …`, `zpool list`). ## How to refresh a server's state ```bash # Show help (no args) scripts/refresh-server-info.sh # Refresh every host discovered under servers/*/ scripts/refresh-server-info.sh all # Refresh a specific host (must match a servers// dir; ssh_config # entry or servers//ssh-target handles how to reach it) scripts/refresh-server-info.sh ana-docker ``` Fleet-wide runs require the literal `all` keyword — no-args prints help so you can't accidentally hit every host by forgetting a name. The script pipes `server_inspect.sh` over SSH via stdin (no scp, no remote cleanup) and writes each `servers//system-details.txt` atomically — a failed run never clobbers the previous snapshot. The inspect script itself is read-only. Each server dir can hold an `ssh-target` file (one line, `` or `@`) as a fallback for when the dir name doesn't resolve via DNS or `~/.ssh/config`. The script prefers whatever ssh would resolve normally and only consults the file when that fails. To register a new server: ```bash scripts/add-host.sh scripts/refresh-server-info.sh # pull the first snapshot ``` To audit discovery without touching the network (checks permissions, unresolvable names with no fallback, missing README / system-details, malformed `ssh-target`): ```bash scripts/refresh-server-info.sh --validate-only all scripts/refresh-server-info.sh --validate-only ``` ## Stack mirror (pull / push) Compose and config trees are mirrored into `stacks-mirror///` so they can be diffed and version-controlled. Pull is fleet-wide and safe; push is one stack at a time with a diff + prompt. ```bash # Pull compose + conf from every host into stacks-mirror/ scripts/sync-stacks.sh scripts/sync-stacks.sh --dry-run # see what would change scripts/sync-stacks.sh ana-docker # one host # Push a local stack back to the server (diffs each file, prompts y/N) scripts/deploy-stack.sh scripts/deploy-stack.sh --compose # skip conf scripts/deploy-stack.sh --conf # skip compose ``` **Opt-out per stack:** create `stacks-mirror///.no-sync` (skip both sides) or `stacks-mirror///conf/.no-sync` (skip conf only). **Always excluded in both directions** (secrets / runtime state): `.env`, `.env.*`, `acme.json`, `client_secrets.json`, `*.pem`, `*.key`, `*.crt`, `*.pfx`, `*.sqlite`, `*.sqlite3`, `*.db`, `*.log`, `*.log.*`, `*.pid`, `hub/`, `logs/`. Requires `rsync` installed on this workstation and every host you sync against (`apt install rsync`). ## Layout ``` eshpfi-management/ ├── CLAUDE.md # this file ├── README.md # human-facing overview ├── scripts/ │ └── server_inspect.sh # gather server state for compose planning ├── servers/ │ └── / │ ├── README.md │ └── system-details.txt # latest server_inspect output ├── stacks/ │ └── / │ ├── compose.yaml # deployed to /opt/docker/compose// │ ├── .env.example # template; real .env lives on server │ └── README.md # what this stack does, how to deploy └── docs/ └── pfi/ # general PFI infrastructure reference ``` ## Working rules - **Copies, not symlinks.** Files here reflect what's on the server at the time of the last sync. When you edit here, the server doesn't change until you deploy. - **Never commit secrets.** Use `.env.example` templates; real `.env` files (with tokens, passwords) live on the server and are gitignored if/when this becomes a git repo. - **Surgical edits.** When fixing one stack, don't touch unrelated ones. Follow AIPA-MCP's CLAUDE.md rules about scope discipline. - **Sanity-check before deploying.** Run `docker compose config` (dry parse) before `docker compose up -d` on the server.