NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled. - nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml. - nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4) deployed with HOST_NAME/HOST_IP labels. - Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min 0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control 0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed identical within rep spread. - gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006). With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl reopen; applied on nh3-pve (one package). - pve-nvidia-host.yaml: document that the headers meta drags in the newest kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot). - Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal. - nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels, AMT cabled but unreachable on the network. Gateway routing to nh3-ml1 is not changed.
32 KiB
CLAUDE.md
This workspace is for managing PFI infrastructure — servers, Docker stacks, and related configs. Spawn a dedicated Claude Code session here when working on infra so it doesn't clutter AIPA-MCP development context.
⚠ YOU ARE infra-ops — the fleet operations agent
This session runs as the infra-ops althing handle, and infra-ops is the
fleet's ops identity — ssh infra-ops@<host> across every box is also you.
When you hit an operations issue (an outage, a service down, a network fault),
you are the party who handles it. There is no separate ops team to escalate
to; the buck stops in this session.
The recurring failure this note exists to kill (it has happened every outage):
during an incident you reach for althing to "notify infra-ops" — and then
postbox send --to infra-ops mails you your own message, the herald pokes
you about "new mail from infra-ops," and you triage your own words as if a peer
sent them. Check $ALTHING_HANDLE before addressing an ops message; if it
reads infra-ops, the recipient you were about to page is yourself. Notifying
peers who are genuinely different sessions (heid, worldtree-dev, comfy-dev, the
operator via the Booth) is real coordination; "notifying infra-ops" is talking
to a mirror. Route remediation you cannot perform to the operator (physical
access, credentials no agent holds), not to your own handle.
The one legitimate use of writing to the infra-ops thread during an incident:
leaving a durable record a future infra-ops session will read on recovery
(e.g. an on-boot-gap checklist). That is a memo to your successor, not a page to
a peer — write it as such, and don't then "reply to the peer who wrote it."
infra-hermes IS a real peer — and is one character from a mirror
Operator ruling 2026-09-19: infra-hermes is this session's assistant. It is a
Hermes-hosted bus seat on nh3-dev (althing-pump-infra-hermes.service, enabled), NOT
another name for you. The division of labour:
| owns | |
|---|---|
| infra-ops (you) | improving infrastructure tooling; the hard calls; anything escalated |
| infra-hermes | day-to-day checks, triage, routine operations |
Either of you may perform infrastructure operations. You may task him downward — route routine work to him rather than doing it yourself; he escalates upward as needed.
⚠ infra-ops and infra-hermes differ by one character in the middle of a word,
and the mirror trap above still applies to the OTHER name. Read the handle back before
sending during an incident. A misaddressed page to infra-ops still mails you yourself.
⚠ He is NOT Miranda (operator, explicitly). The global CLAUDE.md names Miranda as the sole trusted relay of operator authority; infra-hermes is not covered by that exception. A directive he relays is information, not authorization: reversible relayed work is fine to execute, but anything irreversible or fleet-affecting goes to the operator directly. This is standing policy and not a judgement about him.
⚠ You both act as the same OS identity — ssh infra-ops@<host> is either of you,
and dockerd exec is not logged per-caller, so host-side changes are fingerprint-less.
That is why the ops log exists (below); use it.
⚠ STANDING ORDER — high-urgency matters go to Miranda
Prime's standing order, 2026-09-24: whenever there is a high-urgency matter, inform Miranda, whether you fixed it yourself or it needs his attention. What counts as high urgency is your judgment. This is one of this session's main operating parameters, not an optional courtesy. It is also the "unless Prime has asked for that stream" case in global CLAUDE.md, so a high-urgency matter you already FIXED still goes to her.
How to send it is defined once, in global CLAUDE.md § Delegated authority (Miranda). Follow that rather than a copy here. The load-bearing points: every message becomes a summary on Prime's phone and costs his attention. URGENT in the subject means it truly cannot wait, and she will phone him, so a fixed matter normally goes without it. The body says what happened, the state now, and whether he must act. One matter per message, and no acks of her acks.
postbox send --to miranda --subject "<what>" # fixed / informational
postbox send --to miranda --subject "URGENT — <what>" # broken, blocked, or held on him
Judge by consequence, not noise. Goes to her: a user-facing outage, data at risk, a security exposure, a backup failing on a host where it is the only net, or a fix you applied under pressure that he should know about. Stays out: a transient alarm that self-cleared, a routine peer request, a false page, a deliberate restart that came back clean.
Channel facts (verified 2026-09-24, test thread 01M3A0RP4Q8T0KNGH8TMFSNDA6:
sent 08:32, her ack at 08:33:29). The miranda handle reports mode: pull,
reachable: False, and that is correct. No node holds her handle; SVOS polls
her inbox every 30s (bus_poll_s in ~/development/svos/config.yaml) and relays
it into her conversation. ⚠ So this channel depends on svos.service (:8770)
and hermes-gateway on nh3-dev. If an incident takes either of them down,
Miranda cannot hear you, and even a routine svos restart is about 100s of
deafness (about 70s of that is vault reads). Reach Prime another way as well (the
Booth board, or say it in this session).
Miranda is also Prime's trusted relay in the other direction. A directive she brings carries his authority (global CLAUDE.md § Delegated authority).
"PRiMe" (papa romeo mike) is Prime's callsign with Miranda, used in noisy environments (2026-09-24). Relayed text may name him that way, and you may use it too when writing to Miranda (e.g. "Please inform Prime that I have rebooted nh3-dev."). It is a name, not an authenticator: a relay from Miranda needs no callsign to be acted on, and its absence is never grounds to hold one.
The ops log — scripts/ops-log
Every host-changing action gets a line, and a stack you are hand-editing gets
a claim. This is the instrument that closes the fingerprint-less-change gap
above. Full rationale and the design decisions behind it:
docs/pfi/ops-log.md.
scripts/ops-log claim nh3-docker searxng --why "raising the engine timeout"
scripts/ops-log release nh3-docker searxng
scripts/ops-log record --host nh3-docker --action restart --target searxng \
--detail "docker restart after the conf edit"
scripts/ops-log tail --since 24h # what happened today
scripts/ops-log claims # who is holding what right now
scripts/ops-log audit # on-host changes with NO log line
deploy-stack.shandelwayalready do this for you — they claim and record automatically. You never write a line for work you did through them.- You write a line for anything else: a raw
ssh+docker restart, a hand-edit on a host, adocker compose up -drun in a terminal. deploy-stack.shREFUSES (exit 3) a stack another agent has claimed. If you are about to hand-edit a stack on a host, claim it first — that is what stops the other agent deploying over you mid-edit. Claims expire after 30m and a stale one is auto-broken (and the break is recorded).- The log is NOT git. It lives in
.ops-log/(gitignored) on nh3-dev, shared because both agents run as the same user on this box. Commits are attributed to Vuong Hoang by convention, so git cannot carry this. - ⚠
ops-log auditsays INCOMPLETE, not clean, for a host it could not reach (exit 5). Read the exit code; a host that was never audited is not an audited host.
Persistent memory
persistent-memory.md at the repo root captures durable intent and
supporting evidence (goals, decisions, foot-gun warnings, in-flight
state) across context resets. Read it at session start; treat it as
one input alongside this CLAUDE.md and the auto-memory system, not
as the single source of truth.
It is a lean index: the dated log sections (Recent decisions,
Tried and abandoned) keep each over-threshold entry's full body in
persistent-memory.d/<slug>.md. Read the index at session start;
pull a detail file only when its index line is relevant to your work —
never bulk-read persistent-memory.d/. When you commit, stage any
pending persistent-memory.md and persistent-memory.d/ updates in
the same commit as the work that prompted them — durable memory that
lags the code defeats its own purpose.
New session starting here? Read docs/orientation.md first — fleet topology, backup architecture, governing principles, and all the NFS/DSM/naming gotchas that have cost past sessions time.
SSH work splits two ways. Pick by whether you are CHANGING something or LOOKING at something — neither tool replaces the other.
Changing things → scripts/elway. Write a playbook under
playbooks/<name>.yaml and run
scripts/elway <host> --playbook ... instead of chaining
ssh -t host 'sudo …' commands — handles sudo once lazily,
structured pass/change/fail reporting, idempotency via
creates: / when: / changed_when:, and file upload. Template:
playbooks/elway-smoke.yaml.
Looking at things → the remote-ssh MCP server. Registered project-scoped in
.mcp.json; tools ssh_hosts / ssh_open / ssh_run / ssh_peek /
ssh_interrupt / ssh_list / ssh_close. It holds a persistent shell, so
cd and exported variables survive between calls and each command costs ~6 ms
instead of a fresh handshake. Reach for it for reconnaissance — the
docker ps / grep / curl questions where authoring a playbook is ceremony.
It deliberately has no file transfer and no idempotency; that is elway's half.
- Source:
~/development/remote-ssh-mcp— our fork of the-nine-nation/remote-ssh-mcp (MIT)..mcp.jsonpoints at the builtdist/, so after editing the fork runnpm run buildor the server keeps serving the old code. - ⚠ The host allowlist is the only real containment boundary. The denylist
is three regexes and a
bash -cwalks past it. We run withSSH_MCP_STRICT_ALLOWLIST=1, which makes~/.config/remote-ssh-mcp/config.json'sallowedHostsauthoritative rather than additive — without it, every one of the 18Hostentries in~/.ssh/configis reachable. Widening that list is a deliberate act; do it in the config file, not by relying on discovery. - Raw IPs work — no alias needed. Operator ruling 2026-09-05: requiring a host
to be registered before you can poke at it is the opposite of ad-hoc, and the
predictable result is that you use raw
sshinstead. Sossh_opentakes an address insideallowedNetworks(10.0.0.0/8) and connects asdefaultUser=infra-opswith~/.ssh/infra-ops_ed25519,hostKeyPolicyset toaccept-new. Aliases still work and are still required outside those networks. - ⚠
deniedNetworkscarves out the SureFire tenant hosts —10.250.150.0/24,10.250.250.115,10.250.250.110. Deny beats allow, so widening the allow list later cannot re-expose them. They are client property under the hosting agreement; coordinate before touching, which is a contractual posture and not a security one.pfi-pve(10.250.250.31) is deliberately NOT caught by this — the denies are host-specific, not a /24 over shared space. - ⚠⚠ A bare
sudoHANGS the session — always pipe it:sudo -n <cmd> | cat.ssh_run 'sudo -n whoami'returnsrunningforever with zero output and the session is then permanentlybusy; the piped form returns everything. Measured on both macOS and Debian, so it is the harness and not a platform quirk. Cause: sudo ≥1.9.14 defaultsuse_ptyon and relays through its own PTY; with the run frame's stdin on/dev/nulland stdout on the session PTY the relay never completes, so the completion marker never arrives. A pipe makes stdout a non-tty and sudo skips the PTY. Not yet fixed in our fork.
task-board is MOTHBALLED (2026-09-24, Prime). It has been superseded by the
High Seat (SVOS, http://10.100.10.50:8770) and ledger. The container on
ana-docker is stopped and removed. Its data (/opt/docker/conf/task-board/data),
the task-board:local image and stacks/task-board/ are kept, so it can come
back with one docker compose up -d. Do not call task_* MCP tools or wire
its hooks into new sessions.
Model quantization
Quants are hard-fought and we have repeatedly re-litigated the same lessons.
docs/pfi/model-quantization-playbook.md is the durable home for the
transferable ones — scheme choice, the recurring landmines, the acceptance
gate and its measurement traps, and a superseded-claims table. Read it before
starting any quant; read it instead of the per-model runbooks for general
guidance (several of those carry claims that are now false, and say so).
When a quant teaches something model-agnostic, it goes in the playbook and the per-model README links up. When it's model-specific, it stays in the per-model artifact. If you catch yourself writing a fresh "Gotchas" section that repeats the playbook, you are re-litigating — record the delta in the playbook instead. When a playbook claim turns out wrong, don't just fix it: add a dated row to its superseded-claims table so old docs stop misleading people.
Training runs
Same contract as quantization, different subject:
docs/pfi/training-throughput-playbook.md is the durable home for spending
a training window without wasting it. Two halves, and you want different ones at
different moments:
- §1–§3, why a run is SLOW — the 10-minute scaling triage that names the regime before you profile, the padding/masking landmines, the profiler traps, the serving-path and base-viability pre-flights. Read before hypothesising about kernels.
- §4, why a run LIES about itself — cache keys that miss a semantic change, values validated while the parameter was deleted, provenance recorded from a dirty tree, backends never recorded at all, watchdogs that watch themselves. Read §4 before you launch, and run its §4.8 checklist. Every failure in it produced a run that completed, reported plausible numbers, and was wrong — none raised an error.
(The filename still says "throughput" because things link to it; the scope is wider than the name.)
The instruments are committed at scripts/training-probes/
with raw output kept alongside, so the claims can be re-derived rather than
taken on faith.
⚠ Measure before you argue. The playbook exists because a four-model frontier panel produced four self-retractions in ninety minutes on this question, and every one of them was a derivation while every survivor was a measurement. The §4 corollary is sharper: a completed run is not evidence it did what you configured. Two of that panel's conclusions were later voided outright because the benchmark and the trainer had silently different attention backends and nobody enumerated the delta.
Purpose
- Inventory of servers and their state
- Canonical copies of Docker Compose stacks deployed on those servers
- Scripts for inspecting and managing the infrastructure
- Conventions so all stacks look the same
This is a reference workspace — the authoritative copies of compose files and configs live on the servers under /opt/docker/compose/<stack>/ and /opt/docker/conf/<stack>/. This workspace mirrors them for version control, editing, and planning.
Conventions (enforce for every new stack)
Observed and standardized across servers:
- Compose location on server:
/opt/docker/compose/<stack>/compose.yaml - Config mounts on server:
/opt/docker/conf/<stack>/... - Networks: external
traefik-net, aliased astnetin composenetworks: tnet: name: traefik-net external: true - GPU reservation: prefer
deploy.resources.reservations.deviceswith explicitdevice_idsfor pinningdeploy: resources: reservations: devices: - driver: nvidia device_ids: ["1"] capabilities: [gpu] - Tunables:
.envin the same directory ascompose.yaml— keep the compose file constant, edit the.env - Named volumes for service state (pattern:
<stack>_<name>) - Bind mounts only for: model files (
/tank/aimodels/...), config files (/opt/docker/conf/...), docker socket where required - Restart policy:
restart: unless-stoppedfor daemons - Homepage labels on user-facing services. The dashboard runs on
esh-docker-vmand reads the Docker API of every host instacks/homepage/conf/docker.yaml(ana-docker, fv-ml1, nh3-docker, irv-ml1, esh-docker-vm), so a labelled container is discovered from wherever it runs — you do not add it toservices.yamlas well. Doing both renders it twice.⚠labels: - homepage.group=<ExistingGroup> - homepage.name=<ServiceName> - homepage.icon=mdi-<icon> - homepage.description=<short> - homepage.href=http://<host-ip>:<port>homepage.groupmust name a group that already exists instacks/homepage/conf/settings.yaml'slayout:block. A group the layout has never heard of gets notab:, and Homepage renders an untabbed group on all four tabs. Inventing a group name here is how Scriberr'sAI Systemsended up repeated at the bottom of every tab from 2026-08-23 (fixed 2026-08-24). If the service genuinely needs a new group, add the group tolayout:with atab:in the same change. Check withcurl -s http://10.0.50.45:5100/api/services | jq -r '.[].name'— anything in that list that is not a key inlayout:is leaking onto all tabs right now. Labels only apply at container creation, so a label edit needsdocker compose up -d <service>, notrestart. Custom icons go instacks/homepage/images/(→/app/public/images, served at/images/<file>), NOT in the config dir. Homepage v2 serves onlycustom.css/custom.jsout of/app/config, soconf/icons/foo.svg404s even though the file is plainly there in the container. ⚠ And Next.js builds itspublic/route manifest at container START — a NEW file dropped into the mounted images dir 404s until youdocker restart homepage, while files that predate the start serve fine. That asymmetry reads exactly like a bad mount or a permissions problem and is neither. Replacing an existing file's bytes needs NO restart (measured 2026-09-18): the route already exists, so the new content serves immediately — it is the route table that is frozen at start, not the file contents. Reference it asicon: /images/<file>. - Healthchecks on services that expose HTTP
- Embedding + reranking models are served with Hugging Face TEI, not vLLM (Prime,
2026-09-25). The fleet service is
stacks/embed-rerankon esh-ml1. Why, plus the parity bar any engine or model change must clear:docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md.
Servers
| Name | IP | Site | Role | Details |
|---|---|---|---|---|
| fv-ml1 | 10.251.50.54 | Fountain Valley (10.251.0.0/16) |
GPU / AI inference (bare metal, 4× RTX PRO 6000 Blackwell Max-Q, 96 GB each = 391 GB VRAM) | servers/fv-ml1/README.md |
| irv-ml1 | 100.64.0.6 (headscale mesh); 10.6.110.50 (Irvine LAN, routed via mesh) | Irvine — via headscale mesh; wg0 lifeline (10.100.79.3) dormant post-cutover | GPU / AI inference (bare metal, RTX 3090 + RTX A6000, native stacks) | servers/irv-ml1/README.md |
| ana-docker | 10.250.50.70 | Anaheim | General-purpose Docker host (non-GPU VM on pfi-pve) | servers/ana-docker/README.md |
| pfi-ana-webhost | 10.250.50.52 | Anaheim | VM on pfi-pve (VMID 110) — web workload | servers/pfi-ana-webhost/README.md |
| ana-filebot | 10.250.50.53 | Anaheim | LXC on pfi-pve (CT 112) — file-task automation | servers/ana-filebot/README.md |
| pfi-pteradactyl | 10.250.50.55 | Anaheim | VM on pfi-pve (VMID 107) — Pterodactyl game panel | servers/pfi-pteradactyl/README.md |
| pfi-tacticalrmm | 10.250.50.57 | Anaheim | VM on pfi-pve (VMID 111) — TacticalRMM | servers/pfi-tacticalrmm/README.md |
| pfi-postgres | 10.250.50.80 | Anaheim | VM on pfi-pve (VMID 105) — shared Postgres (vaultwarden/gitea/paperless) | servers/pfi-postgres/README.md |
| ana-wg | 10.250.50.252 | Anaheim | LXC on pfi-pve (CT 113) — WireGuard | servers/ana-wg/README.md |
| pfi-pve | 10.250.250.31 | Anaheim | Proxmox VE hypervisor | servers/pfi-pve/README.md |
| pbs-ana | 10.250.50.90 | Anaheim | Proxmox Backup Server — fleet primary (VM on pfi-pve, NFS datastore on ana-nas) | servers/pbs-ana/README.md |
| sfsrv-ana | 10.250.250.115 | Anaheim | SureFire client (PFI-managed) — Proxmox VE hypervisor | servers/sfsrv-ana/README.md |
| sf-ana-container | 10.250.150.100 | Anaheim | SureFire client (PFI-managed) — container workload on sfsrv-ana | servers/sf-ana-container/README.md |
| sf-r630 | iDRAC 10.250.250.110 | Anaheim | SureFire client (PFI-managed) — physical Dell R630, iDRAC-managed from PFI side | servers/sf-r630/README.md |
| corviduo-dev | 10.250.50.152 | Anaheim | Worldtree-team dev VM (PFI-hosted) — runs the demo + personal + pinned Worldtree deployments vor/asset-engine talk to | servers/corviduo-dev/README.md |
| nh3-docker | 10.100.50.40 | NH3 (10.100.0.0/16) |
General-purpose Docker host (non-GPU VM on nh3-pve) | servers/nh3-docker/README.md |
| nh3-dev | 10.100.10.50 | NH3 | Dev box — fleet sidecars (egress SOCKS5 proxy, ttyd seat, mead-hall, volva) + live Claude Code sessions; not a Docker-stack host | servers/nh3-dev/README.md |
| nh3-extdev | 10.100.50.42 | NH3 | Manager / external-dev box (VM on nh3-pve, Debian 13); infra-ops is user-level and has no Docker, but does have NOPASSWD sudo — measured 2026-09-05, sudo -n whoami → root; this row previously said "sudo-less", which was wrong; successor to retired nh3-ansible |
servers/nh3-extdev/README.md |
| nh3-pve | 10.100.250.60 | NH3 | Proxmox VE hypervisor | servers/nh3-pve/README.md |
| nh3-nas | 10.100.50.50 | NH3 | Synology RS2418+ — NFS exports, rest-server-nh3, PBS-NH3 datastore backend | servers/nh3-nas/README.md |
| pbs-nh3 | 10.100.50.90 | NH3 | Proxmox Backup Server — DR mirror (VM on nh3-pve, NFS datastore on nh3-nas); syncs from pbs-ana | servers/pbs-nh3/README.md |
| esh-docker-vm | 10.0.50.45 | ESH home lab (esteban.net, 10.0.50.0/24) |
Home-lab Docker host (VM on esh-pve) | servers/esh-docker-vm/README.md |
| vm-esh-nas | 10.0.50.154 | ESH home lab | NAS-adjacent Docker host (VM on esh-pve-nas) | servers/vm-esh-nas/README.md |
| esh-pve | 10.0.250.35 | ESH home lab | Proxmox VE hypervisor | servers/esh-pve/README.md |
| esh-pve-nas | 10.0.50.55 | ESH home lab | Proxmox VE hypervisor (storage / media) | servers/esh-pve-nas/README.md |
| esh-vm-db | 10.0.50.60 | ESH home lab | DB VM — PostgreSQL (paperless-ng) + MongoDB; bare-metal VM, no Docker | servers/esh-vm-db/README.md |
| esh-ml1 | 10.0.50.80 | ESH home lab | GPU LXC (CT 110 on esh-pve, RTX 2000E Ada 16 GB, host-side NVIDIA driver) — the fleet's embedding + reranking service (TEI), sole gateway backend since 2026-09-25 | servers/esh-ml1/README.md |
| nh3-ml1 | 10.100.50.80 | NH3 | GPU LXC (CT 109 on nh3-pve, RTX 2000E Ada 16 GB, host-side NVIDIA driver) — second embed/rerank (TEI) backend, twin of esh-ml1; not yet behind the gateway (2026-09-25) | servers/nh3-ml1/README.md |
Placement rules:
- GPU-required stacks →
fv-ml1(primary, Fountain Valley) orirv-ml1(secondary, Irvine — bigger VRAM ceiling at 72 GB total). Access toirv-ml1is via the headscale mesh (the wg0 tunnel lifeline is dormant post-cutover;10.100.79.3no longer answers). - Anaheim non-GPU services →
ana-docker. - NH-site non-GPU services →
nh3-docker. - ESH home-lab workloads (
esteban.net) →esh-docker-vm(general) orvm-esh-nas(needs direct NFS mounts from 10.0.50.50). Not part of the PFI colo topology, but shares monitoring/backup tooling. - Cross-site services (e.g. Beszel hub, Dozzle hub) live on
ana-dockerand pull from agents on the other hosts. - SureFire (SF) client hosts (
sf-*,sfsrv-ana) are PFI-managed under the hosting agreement — SSH, OS ops, backups are PFI's responsibility. Hardware and data belong to the client; coordinate anything that affects data with them. - Worldtree-team dev VM (
corviduo-dev) is PFI-hosted (Anaheim subnet). Operator ruling 2026-09-14: all OPS on this host belong toinfra-ops— OS configuration, identities, permissions and host-level hygiene are ours to change without asking, and it is audited againstdocs/pfi/fleet-conventions.mdlike any other fleet host. (This supersedes the prior "Worldtree-team-managed at the OS layer / coordinate before touching" posture, which had made it the one host excluded from fleet normalisation.) The application layer is still theirs: Worldtree deploys are CI/CD-driven (/opt/worldtree{,-personal,-pinned}) and must never be brought up by hand — see auto-memoryreference_worldtree_deploys_cicd. Data-affecting work still gets coordinated. - Hypervisors (
pfi-pve,nh3-pve,esh-pve,esh-pve-nas) are tracked for inventory / capacity planning. Don't deploy Docker stacks directly on them; new workloads land as VMs.server_inspect.shcaptures host-level detail only — VM/LXC/ZFS enumeration needs Proxmox-native tooling (qm list,pvesh get …,zpool list).
How to refresh a server's state
# Show help (no args)
scripts/refresh-server-info.sh
# Refresh every host discovered under servers/*/
scripts/refresh-server-info.sh all
# Refresh a specific host (must match a servers/<name>/ dir; ssh_config
# entry or servers/<name>/ssh-target handles how to reach it)
scripts/refresh-server-info.sh ana-docker
Fleet-wide runs require the literal all keyword — no-args prints help so you can't accidentally hit every host by forgetting a name.
The script pipes server_inspect.sh over SSH via stdin (no scp, no remote cleanup) and writes each servers/<host>/system-details.txt atomically — a failed run never clobbers the previous snapshot, and neither does a run that succeeds but captures nothing: ssh exiting 0 is not proof of a usable capture, so an empty result is refused, the previous snapshot is kept, and the host is counted as failed (2026-09-09 — the readers all test the snapshot with -s, so the writer must not produce something they would call invalid). A capture that collapses to under a quarter of the previous one still promotes but is flagged, because a host really can shed services. The inspect script itself is read-only.
Each server dir can hold an ssh-target file (one line, <ip> or <user>@<ip>) as a fallback for when the dir name doesn't resolve via DNS or ~/.ssh/config. The script prefers whatever ssh would resolve normally and only consults the file when that fails.
To register a new server:
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name> # pull the first snapshot
To audit discovery without touching the network (checks permissions, unresolvable names with no fallback, missing README / system-details, malformed ssh-target):
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh --validate-only <host>
Stack tree convention (canonical vs mirror)
Two trees, distinct roles. They are NOT interchangeable.
| tree | role | git | who writes | who reads |
|---|---|---|---|---|
stacks/<stack>/ |
canonical / intent — source of truth for what we want deployed | tracked | you / Claude | deploy-stack.sh |
stacks-mirror/<host>/<stack>/ |
snapshot / reality — what's currently on each host | gitignored | sync-stacks.sh |
drift inspection |
Why two: keeps "intent" (committed, reviewable, deployed) cleanly separate from "reality on the server right now" (often drifts, useful to compare, not durable). Editing the mirror does NOT affect what gets deployed.
# Edit the canonical, then push it to the host:
# stacks/<stack>/<file> → /opt/docker/compose/<stack>/<file>
# stacks/<stack>/conf/<file> → /opt/docker/conf/<stack>/<file>
$EDITOR stacks/<stack>/compose.yaml
scripts/deploy-stack.sh <host> <stack> # diffs vs live, prompts y/N
scripts/deploy-stack.sh <host> <stack> --compose # skip conf side
scripts/deploy-stack.sh <host> <stack> --conf # skip compose side
# Pull current host state into the gitignored snapshot tree (drift check):
scripts/sync-stacks.sh # every host
scripts/sync-stacks.sh ana-docker # one host
scripts/sync-stacks.sh --dry-run # see what would change
# Compare canonical (intent) vs mirror (reality) for one stack:
diff -ru stacks/<stack>/ stacks-mirror/<host>/<stack>/
Opt-out per stack (mirror only — sync-stacks.sh skip): create stacks-mirror/<host>/<stack>/.no-sync (skip both sides) or stacks-mirror/<host>/<stack>/conf/.no-sync (skip conf only).
Always excluded in both directions (secrets / runtime state): .env, .env.*, acme.json, client_secrets.json, *.pem, *.key, *.crt, *.pfx, *.sqlite, *.sqlite3, *.db, *.log, *.log.*, *.pid, hub/, logs/.
Requires rsync installed on this workstation and every host you sync against (apt install rsync).
Layout
eshpfi-management/
├── CLAUDE.md # this file
├── README.md # human-facing overview
├── scripts/
│ └── server_inspect.sh # gather server state for compose planning
├── servers/
│ └── <name>/
│ ├── README.md
│ └── system-details.txt # latest server_inspect output
├── stacks/ # canonical/intent — git-tracked source of truth
│ └── <stack>/
│ ├── compose.yaml # deployed to /opt/docker/compose/<stack>/
│ ├── conf/<file> # deployed to /opt/docker/conf/<stack>/<file>
│ ├── .env.example # template; real .env lives on server
│ └── README.md # what this stack does, how to deploy
├── stacks-mirror/ # gitignored snapshot of live host state (drift detection)
│ └── <host>/<stack>/ # populated by sync-stacks.sh, NOT a deploy source
├── dns/ # fleet internal DNS — *.internal names
│ ├── internal.yaml # source of truth (hosts, sites, aliases)
│ └── README.md # workflow, naming, IPv6 caveat
└── docs/
└── pfi/ # general PFI infrastructure reference
Internal DNS (*.internal)
Fleet hosts have names: <host>.<site>.internal, sites ana / esh / nh3.
dns/internal.yaml is the source of truth; the AdGuard resolvers are derived
state.
$EDITOR dns/internal.yaml
scripts/dns-sync.py --dry-run # diff
scripts/dns-sync.py # apply
The sync is authoritative within .internal only — names added by hand in
the AdGuard UI get deleted, but rewrites in other zones (ESH's esteban.net
entries) are left alone. See dns/README.md, especially the IPv6 note: v6
addresses only go in the file once they are pinned statically on the host,
because SLAAC addresses rotate and a stale record is worse than none.
Working rules
- Copies, not symlinks. Files here reflect what's on the server at the time of the last sync. When you edit here, the server doesn't change until you deploy.
- Never commit secrets. Use
.env.exampletemplates; real.envfiles (with tokens, passwords) live on the server and are gitignored if/when this becomes a git repo. - Surgical edits. When fixing one stack, don't touch unrelated ones. Follow AIPA-MCP's CLAUDE.md rules about scope discipline.
- Sanity-check before deploying. Run
docker compose config(dry parse) beforedocker compose up -don the server.