Names for fleet hosts so addresses stop needing to be memorised. Built because IPv6 makes that hopeless — and, more to the point, because v6 addresses are derived rather than assigned, so they cannot reliably be written down once and trusted either. dns/internal.yaml source of truth: 38 hosts + 4 service aliases scripts/dns-sync.py reconciles AdGuard resolvers against it stacks/adguard-ana/ the colo's resolver, which did not exist Naming is <host>.<site>.internal with sites ana/esh/nh3 (operator's call). .internal is ICANN-reserved for this; .local is reserved for mDNS, which is why searxng.pfi.local was a collision that merely happened to work. Same posture as deploy-stack.sh: file is intent, resolvers are derived state, you see a diff before anything changes. Every name is published to every resolver, so the site label says where a host IS, not who knows about it. Two properties that matter: - Authority is scoped to the ZONE, not the resolver. ESH carries hand-made esteban.net rewrites predating this; they are read, ignored and preserved. Resolver-wide authority would have silently deleted them. - Within .internal it IS authoritative, so UI-added names get removed. That is the point — one place to look. Colo gap closed: ana-docker had no resolver at all (hosts went straight to 1.1.1.1). Its AdGuard runs API on 8053 because 8080/3000 were taken, so the port is carried per-site in the yaml rather than assumed by the script. It ships with no blocklists — a false positive on a server network breaks service-to-service calls for no upside. Auth is a dedicated infra-ops AdGuard user, not the operator's account, password vaulted at nh3-dev/adguard-infra-ops-password. Pre-change configs backed up on each resolver. Both resolvers stayed answering across the restart. searxng.pfi.local -> searxng.ana.internal, with the old Host() kept alongside so nothing breaks mid-migration. matrix.pfi.local deliberately NOT migrated: a Matrix server_name is baked into every user id, room id and signing key, so renaming it rebuilds the homeserver's identity rather than changing a DNS name. The v6 column is empty and correct — no fleet host has a global v6 address yet. The file documents why addresses must be pinned statically before they go in, since a record that silently stops matching is worse than no record.
16 KiB
CLAUDE.md
This workspace is for managing PFI infrastructure — servers, Docker stacks, and related configs. Spawn a dedicated Claude Code session here when working on infra so it doesn't clutter AIPA-MCP development context.
Persistent memory
persistent-memory.md at the repo root captures durable intent and
supporting evidence (goals, decisions, foot-gun warnings, in-flight
state) across context resets. Read it at session start; treat it as
one input alongside this CLAUDE.md and the auto-memory system, not
as the single source of truth.
It is a lean index: the dated log sections (Recent decisions,
Tried and abandoned) keep each over-threshold entry's full body in
persistent-memory.d/<slug>.md. Read the index at session start;
pull a detail file only when its index line is relevant to your work —
never bulk-read persistent-memory.d/. When you commit, stage any
pending persistent-memory.md and persistent-memory.d/ updates in
the same commit as the work that prompted them — durable memory that
lags the code defeats its own purpose.
New session starting here? Read docs/orientation.md first — fleet topology, backup architecture, governing principles, and all the NFS/DSM/naming gotchas that have cost past sessions time.
For SSH-driven work: use scripts/elway. Write a playbook under
playbooks/<name>.yaml and run
scripts/elway <host> --playbook ... instead of chaining
ssh -t host 'sudo …' commands — handles sudo once lazily,
structured pass/change/fail reporting, idempotency via
creates: / when: / changed_when:. Template:
playbooks/elway-smoke.yaml.
Task visibility via task-board. If the Claude Code session has
the task-board plugin enabled (installed from
git@gitea.phasefinal.com:vh/task-board.git), a card at
http://10.250.50.70:7878/ tracks work in progress. Hooks flip the
card on turn boundaries automatically; call task_start /
task_update / task_wait / task_complete MCP tools to set the
activity subheader and post meaningful log entries.
When you launch a Bash tool with run_in_background: true (or any
long-running shell / monitor / poll loop), call task_set_shells
with one short description per active background shell — and call it
again with the updated list (or []) when one completes. The board
flips a waiting card to orange while the list is non-empty so the
user can tell at a glance the session is parked on background work,
not stalled on them. Hooks have no way to enumerate the bg-task list
externally, so this is on the assistant.
Model quantization
Quants are hard-fought and we have repeatedly re-litigated the same lessons.
docs/pfi/model-quantization-playbook.md is the durable home for the
transferable ones — scheme choice, the recurring landmines, the acceptance
gate and its measurement traps, and a superseded-claims table. Read it before
starting any quant; read it instead of the per-model runbooks for general
guidance (several of those carry claims that are now false, and say so).
When a quant teaches something model-agnostic, it goes in the playbook and the per-model README links up. When it's model-specific, it stays in the per-model artifact. If you catch yourself writing a fresh "Gotchas" section that repeats the playbook, you are re-litigating — record the delta in the playbook instead. When a playbook claim turns out wrong, don't just fix it: add a dated row to its superseded-claims table so old docs stop misleading people.
Purpose
- Inventory of servers and their state
- Canonical copies of Docker Compose stacks deployed on those servers
- Scripts for inspecting and managing the infrastructure
- Conventions so all stacks look the same
This is a reference workspace — the authoritative copies of compose files and configs live on the servers under /opt/docker/compose/<stack>/ and /opt/docker/conf/<stack>/. This workspace mirrors them for version control, editing, and planning.
Conventions (enforce for every new stack)
Observed and standardized across servers:
- Compose location on server:
/opt/docker/compose/<stack>/compose.yaml - Config mounts on server:
/opt/docker/conf/<stack>/... - Networks: external
traefik-net, aliased astnetin composenetworks: tnet: name: traefik-net external: true - GPU reservation: prefer
deploy.resources.reservations.deviceswith explicitdevice_idsfor pinningdeploy: resources: reservations: devices: - driver: nvidia device_ids: ["1"] capabilities: [gpu] - Tunables:
.envin the same directory ascompose.yaml— keep the compose file constant, edit the.env - Named volumes for service state (pattern:
<stack>_<name>) - Bind mounts only for: model files (
/tank/aimodels/...), config files (/opt/docker/conf/...), docker socket where required - Restart policy:
restart: unless-stoppedfor daemons - Homepage labels on user-facing services:
labels: - homepage.group=AI Systems - homepage.name=<ServiceName> - homepage.icon=mdi-<icon> - homepage.description=<short> - homepage.href=http://<host-ip>:<port> - Healthchecks on services that expose HTTP
Servers
| Name | IP | Site | Role | Details |
|---|---|---|---|---|
| ana-ml2 | 10.250.50.54 | Anaheim (10.250.0.0/16) |
GPU / AI inference (bare metal, dual RTX PRO 6000 Blackwell Max-Q, 96 GB each) | servers/ana-ml2/README.md |
| irv-ml1 | 10.100.79.3 (WG) | Irvine — reachable only via WireGuard tunnel from NH3 | GPU / AI inference (bare metal, RTX 3090 + RTX A6000, native stacks) | servers/irv-ml1/README.md |
| ana-docker | 10.250.50.70 | Anaheim | General-purpose Docker host (non-GPU VM on pfi-pve) | servers/ana-docker/README.md |
| pfi-ana-webhost | 10.250.50.52 | Anaheim | VM on pfi-pve (VMID 110) — web workload | servers/pfi-ana-webhost/README.md |
| ana-filebot | 10.250.50.53 | Anaheim | LXC on pfi-pve (CT 112) — file-task automation | servers/ana-filebot/README.md |
| pfi-pteradactyl | 10.250.50.55 | Anaheim | VM on pfi-pve (VMID 107) — Pterodactyl game panel | servers/pfi-pteradactyl/README.md |
| pfi-tacticalrmm | 10.250.50.57 | Anaheim | VM on pfi-pve (VMID 111) — TacticalRMM | servers/pfi-tacticalrmm/README.md |
| pfi-postgres | 10.250.50.80 | Anaheim | VM on pfi-pve (VMID 105) — shared Postgres (vaultwarden/gitea/paperless) | servers/pfi-postgres/README.md |
| ana-wg | 10.250.50.252 | Anaheim | LXC on pfi-pve (CT 113) — WireGuard | servers/ana-wg/README.md |
| pfi-pve | 10.250.250.31 | Anaheim | Proxmox VE hypervisor | servers/pfi-pve/README.md |
| pbs-ana | 10.250.50.90 | Anaheim | Proxmox Backup Server — fleet primary (VM on pfi-pve, NFS datastore on ana-nas) | servers/pbs-ana/README.md |
| sfsrv-ana | 10.250.250.115 | Anaheim | SureFire client (PFI-managed) — Proxmox VE hypervisor | servers/sfsrv-ana/README.md |
| sf-ana-container | 10.250.150.100 | Anaheim | SureFire client (PFI-managed) — container workload on sfsrv-ana | servers/sf-ana-container/README.md |
| sf-r630 | iDRAC 10.250.250.110 | Anaheim | SureFire client (PFI-managed) — physical Dell R630, iDRAC-managed from PFI side | servers/sf-r630/README.md |
| corviduo-dev | 10.250.50.152 | Anaheim | Worldtree-team dev VM (PFI-hosted) — runs the demo + personal + pinned Worldtree deployments vor/asset-engine talk to | servers/corviduo-dev/README.md |
| nh3-docker | 10.100.50.40 | NH3 (10.100.0.0/16) |
General-purpose Docker host (non-GPU VM on nh3-pve) | servers/nh3-docker/README.md |
| nh3-dev | 10.100.10.50 | NH3 | Dev box — fleet sidecars (egress SOCKS5 proxy, ttyd seat, mead-hall, volva) + live Claude Code sessions; not a Docker-stack host | servers/nh3-dev/README.md |
| nh3-extdev | 10.100.50.42 | NH3 | Manager / external-dev box (VM on nh3-pve, Debian 13); sudo-less infra-ops identity (user-level only, no Docker); successor to retired nh3-ansible | servers/nh3-extdev/README.md |
| nh3-pve | 10.100.250.60 | NH3 | Proxmox VE hypervisor | servers/nh3-pve/README.md |
| nh3-nas | 10.100.50.50 | NH3 | Synology RS2418+ — NFS exports, rest-server-nh3, PBS-NH3 datastore backend | servers/nh3-nas/README.md |
| pbs-nh3 | 10.100.50.90 | NH3 | Proxmox Backup Server — DR mirror (VM on nh3-pve, NFS datastore on nh3-nas); syncs from pbs-ana | servers/pbs-nh3/README.md |
| esh-docker-vm | 10.0.50.45 | ESH home lab (esteban.net, 10.0.50.0/24) |
Home-lab Docker host (VM on esh-pve) | servers/esh-docker-vm/README.md |
| vm-esh-nas | 10.0.50.154 | ESH home lab | NAS-adjacent Docker host (VM on esh-pve-nas) | servers/vm-esh-nas/README.md |
| esh-pve | 10.0.250.35 | ESH home lab | Proxmox VE hypervisor | servers/esh-pve/README.md |
| esh-pve-nas | 10.0.50.55 | ESH home lab | Proxmox VE hypervisor (storage / media) | servers/esh-pve-nas/README.md |
| esh-vm-db | 10.0.50.60 | ESH home lab | DB VM — PostgreSQL (paperless-ng) + MongoDB; bare-metal VM, no Docker | servers/esh-vm-db/README.md |
Placement rules:
- GPU-required stacks →
ana-ml2(primary, Anaheim) orirv-ml1(secondary, Irvine — bigger VRAM ceiling at 72 GB total). Access toirv-ml1requires WireGuard. - Anaheim non-GPU services →
ana-docker. - NH-site non-GPU services →
nh3-docker. - ESH home-lab workloads (
esteban.net) →esh-docker-vm(general) orvm-esh-nas(needs direct NFS mounts from 10.0.50.50). Not part of the PFI colo topology, but shares monitoring/backup tooling. - Cross-site services (e.g. Beszel hub, Dozzle hub) live on
ana-dockerand pull from agents on the other hosts. - SureFire (SF) client hosts (
sf-*,sfsrv-ana) are PFI-managed under the hosting agreement — SSH, OS ops, backups are PFI's responsibility. Hardware and data belong to the client; coordinate anything that affects data with them. - Worldtree-team dev VM (
corviduo-dev) is PFI-hosted (Anaheim subnet) but Worldtree-team-managed at the OS / application layer. PFI handles networking + emergency-ops backstop; OS configuration + deploy workflows + backup decisions live with the architect's team. Treat data-affecting work like SF hosts — coordinate before touching. - Hypervisors (
pfi-pve,nh3-pve,esh-pve,esh-pve-nas) are tracked for inventory / capacity planning. Don't deploy Docker stacks directly on them; new workloads land as VMs.server_inspect.shcaptures host-level detail only — VM/LXC/ZFS enumeration needs Proxmox-native tooling (qm list,pvesh get …,zpool list).
How to refresh a server's state
# Show help (no args)
scripts/refresh-server-info.sh
# Refresh every host discovered under servers/*/
scripts/refresh-server-info.sh all
# Refresh a specific host (must match a servers/<name>/ dir; ssh_config
# entry or servers/<name>/ssh-target handles how to reach it)
scripts/refresh-server-info.sh ana-docker
Fleet-wide runs require the literal all keyword — no-args prints help so you can't accidentally hit every host by forgetting a name.
The script pipes server_inspect.sh over SSH via stdin (no scp, no remote cleanup) and writes each servers/<host>/system-details.txt atomically — a failed run never clobbers the previous snapshot. The inspect script itself is read-only.
Each server dir can hold an ssh-target file (one line, <ip> or <user>@<ip>) as a fallback for when the dir name doesn't resolve via DNS or ~/.ssh/config. The script prefers whatever ssh would resolve normally and only consults the file when that fails.
To register a new server:
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name> # pull the first snapshot
To audit discovery without touching the network (checks permissions, unresolvable names with no fallback, missing README / system-details, malformed ssh-target):
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh --validate-only <host>
Stack tree convention (canonical vs mirror)
Two trees, distinct roles. They are NOT interchangeable.
| tree | role | git | who writes | who reads |
|---|---|---|---|---|
stacks/<stack>/ |
canonical / intent — source of truth for what we want deployed | tracked | you / Claude | deploy-stack.sh |
stacks-mirror/<host>/<stack>/ |
snapshot / reality — what's currently on each host | gitignored | sync-stacks.sh |
drift inspection |
Why two: keeps "intent" (committed, reviewable, deployed) cleanly separate from "reality on the server right now" (often drifts, useful to compare, not durable). Editing the mirror does NOT affect what gets deployed.
# Edit the canonical, then push it to the host:
# stacks/<stack>/<file> → /opt/docker/compose/<stack>/<file>
# stacks/<stack>/conf/<file> → /opt/docker/conf/<stack>/<file>
$EDITOR stacks/<stack>/compose.yaml
scripts/deploy-stack.sh <host> <stack> # diffs vs live, prompts y/N
scripts/deploy-stack.sh <host> <stack> --compose # skip conf side
scripts/deploy-stack.sh <host> <stack> --conf # skip compose side
# Pull current host state into the gitignored snapshot tree (drift check):
scripts/sync-stacks.sh # every host
scripts/sync-stacks.sh ana-docker # one host
scripts/sync-stacks.sh --dry-run # see what would change
# Compare canonical (intent) vs mirror (reality) for one stack:
diff -ru stacks/<stack>/ stacks-mirror/<host>/<stack>/
Opt-out per stack (mirror only — sync-stacks.sh skip): create stacks-mirror/<host>/<stack>/.no-sync (skip both sides) or stacks-mirror/<host>/<stack>/conf/.no-sync (skip conf only).
Always excluded in both directions (secrets / runtime state): .env, .env.*, acme.json, client_secrets.json, *.pem, *.key, *.crt, *.pfx, *.sqlite, *.sqlite3, *.db, *.log, *.log.*, *.pid, hub/, logs/.
Requires rsync installed on this workstation and every host you sync against (apt install rsync).
Layout
eshpfi-management/
├── CLAUDE.md # this file
├── README.md # human-facing overview
├── scripts/
│ └── server_inspect.sh # gather server state for compose planning
├── servers/
│ └── <name>/
│ ├── README.md
│ └── system-details.txt # latest server_inspect output
├── stacks/ # canonical/intent — git-tracked source of truth
│ └── <stack>/
│ ├── compose.yaml # deployed to /opt/docker/compose/<stack>/
│ ├── conf/<file> # deployed to /opt/docker/conf/<stack>/<file>
│ ├── .env.example # template; real .env lives on server
│ └── README.md # what this stack does, how to deploy
├── stacks-mirror/ # gitignored snapshot of live host state (drift detection)
│ └── <host>/<stack>/ # populated by sync-stacks.sh, NOT a deploy source
├── dns/ # fleet internal DNS — *.internal names
│ ├── internal.yaml # source of truth (hosts, sites, aliases)
│ └── README.md # workflow, naming, IPv6 caveat
└── docs/
└── pfi/ # general PFI infrastructure reference
Internal DNS (*.internal)
Fleet hosts have names: <host>.<site>.internal, sites ana / esh / nh3.
dns/internal.yaml is the source of truth; the AdGuard resolvers are derived
state.
$EDITOR dns/internal.yaml
scripts/dns-sync.py --dry-run # diff
scripts/dns-sync.py # apply
The sync is authoritative within .internal only — names added by hand in
the AdGuard UI get deleted, but rewrites in other zones (ESH's esteban.net
entries) are left alone. See dns/README.md, especially the IPv6 note: v6
addresses only go in the file once they are pinned statically on the host,
because SLAAC addresses rotate and a stale record is worse than none.
Working rules
- Copies, not symlinks. Files here reflect what's on the server at the time of the last sync. When you edit here, the server doesn't change until you deploy.
- Never commit secrets. Use
.env.exampletemplates; real.envfiles (with tokens, passwords) live on the server and are gitignored if/when this becomes a git repo. - Surgical edits. When fixing one stack, don't touch unrelated ones. Follow AIPA-MCP's CLAUDE.md rules about scope discipline.
- Sanity-check before deploying. Run
docker compose config(dry parse) beforedocker compose up -don the server.