The playbook covered why a run is SLOW. It did not cover the more expensive
failure: a run that COMPLETES, reports plausible numbers, and is wrong about
itself. Seven of those turned up on the Gemma-4 ERP/RP tune between 08-24 and
08-26 and not one raised an error.
New §4, seven landmines plus a pre-launch checklist:
4.1 a cache key must cover the MEANING of the cached thing. The encode
cache missed the impersonation mask; run 2 would have reused run 1's
unmasked encodings and written impersonation_mask_sha256 into its own
manifest while doing it. No error, no count change, normal loss curve.
4.2 validating a VALUE is not validating the PARAMETER. warmup_ratio was
in range and deleted from transformers 5. Build kwargs as data and
diff the NAMES against the installed signature -- you cannot check the
argument list of a call you have already made.
4.3 record what the run RESOLVED to, never what it requested. Run 1
recorded no attention backend, so an MFU panel profiled the serving
seat under sdpa and recommended adopting flex_attention for a run that
was already using it.
4.4 never train from a dirty tree; harness_commit will name a commit that
does not describe the run. Annotate afterwards, never edit the shipped
artifact -- and state what is NOT wrong, or the note casts doubt on
every field it omits.
4.5 a watchdog whose pgrep pattern appears in its own argv can only ever
return "alive". The inert-gate shape in a liveness check.
4.6 an instrument nobody runs is not an instrument. Mutation-check any
test guarding a property that fails silently.
4.7 fix a stale measurement at the source. "~4.3 HOURS to rebuild the
encode cache" (really 145.5 s) was copied into a new launcher by the
same person who had just measured the real number.
4.8 the pre-launch honesty checklist, ten minutes.
Also:
- Header and framing widened. The file is now a training playbook with a
throughput half and an integrity half; the filename stays for inbound links.
- Sections 4-7 renumbered to 5-8. External refs are all to §1.1 and §3.4 and
are unaffected.
- Four rows added to the superseded-claims table, including the kernel table /
68% quadratic / 8.6% MFU set, which describe the serving seat rather than
the training run.
- gemma4-erp-tune-sizing.md §6 carries a correction banner with the explicit
falls/survives split, because that is the doc someone actually reads before
a run.
19 KiB
CLAUDE.md
This workspace is for managing PFI infrastructure — servers, Docker stacks, and related configs. Spawn a dedicated Claude Code session here when working on infra so it doesn't clutter AIPA-MCP development context.
Persistent memory
persistent-memory.md at the repo root captures durable intent and
supporting evidence (goals, decisions, foot-gun warnings, in-flight
state) across context resets. Read it at session start; treat it as
one input alongside this CLAUDE.md and the auto-memory system, not
as the single source of truth.
It is a lean index: the dated log sections (Recent decisions,
Tried and abandoned) keep each over-threshold entry's full body in
persistent-memory.d/<slug>.md. Read the index at session start;
pull a detail file only when its index line is relevant to your work —
never bulk-read persistent-memory.d/. When you commit, stage any
pending persistent-memory.md and persistent-memory.d/ updates in
the same commit as the work that prompted them — durable memory that
lags the code defeats its own purpose.
New session starting here? Read docs/orientation.md first — fleet topology, backup architecture, governing principles, and all the NFS/DSM/naming gotchas that have cost past sessions time.
For SSH-driven work: use scripts/elway. Write a playbook under
playbooks/<name>.yaml and run
scripts/elway <host> --playbook ... instead of chaining
ssh -t host 'sudo …' commands — handles sudo once lazily,
structured pass/change/fail reporting, idempotency via
creates: / when: / changed_when:. Template:
playbooks/elway-smoke.yaml.
Task visibility via task-board. If the Claude Code session has
the task-board plugin enabled (installed from
git@gitea.phasefinal.com:vh/task-board.git), a card at
http://10.250.50.70:7878/ tracks work in progress. Hooks flip the
card on turn boundaries automatically; call task_start /
task_update / task_wait / task_complete MCP tools to set the
activity subheader and post meaningful log entries.
When you launch a Bash tool with run_in_background: true (or any
long-running shell / monitor / poll loop), call task_set_shells
with one short description per active background shell — and call it
again with the updated list (or []) when one completes. The board
flips a waiting card to orange while the list is non-empty so the
user can tell at a glance the session is parked on background work,
not stalled on them. Hooks have no way to enumerate the bg-task list
externally, so this is on the assistant.
Model quantization
Quants are hard-fought and we have repeatedly re-litigated the same lessons.
docs/pfi/model-quantization-playbook.md is the durable home for the
transferable ones — scheme choice, the recurring landmines, the acceptance
gate and its measurement traps, and a superseded-claims table. Read it before
starting any quant; read it instead of the per-model runbooks for general
guidance (several of those carry claims that are now false, and say so).
When a quant teaches something model-agnostic, it goes in the playbook and the per-model README links up. When it's model-specific, it stays in the per-model artifact. If you catch yourself writing a fresh "Gotchas" section that repeats the playbook, you are re-litigating — record the delta in the playbook instead. When a playbook claim turns out wrong, don't just fix it: add a dated row to its superseded-claims table so old docs stop misleading people.
Training runs
Same contract as quantization, different subject:
docs/pfi/training-throughput-playbook.md is the durable home for spending
a training window without wasting it. Two halves, and you want different ones at
different moments:
- §1–§3, why a run is SLOW — the 10-minute scaling triage that names the regime before you profile, the padding/masking landmines, the profiler traps, the serving-path and base-viability pre-flights. Read before hypothesising about kernels.
- §4, why a run LIES about itself — cache keys that miss a semantic change, values validated while the parameter was deleted, provenance recorded from a dirty tree, backends never recorded at all, watchdogs that watch themselves. Read §4 before you launch, and run its §4.8 checklist. Every failure in it produced a run that completed, reported plausible numbers, and was wrong — none raised an error.
(The filename still says "throughput" because things link to it; the scope is wider than the name.)
The instruments are committed at scripts/training-probes/
with raw output kept alongside, so the claims can be re-derived rather than
taken on faith.
⚠ Measure before you argue. The playbook exists because a four-model frontier panel produced four self-retractions in ninety minutes on this question, and every one of them was a derivation while every survivor was a measurement. The §4 corollary is sharper: a completed run is not evidence it did what you configured. Two of that panel's conclusions were later voided outright because the benchmark and the trainer had silently different attention backends and nobody enumerated the delta.
Purpose
- Inventory of servers and their state
- Canonical copies of Docker Compose stacks deployed on those servers
- Scripts for inspecting and managing the infrastructure
- Conventions so all stacks look the same
This is a reference workspace — the authoritative copies of compose files and configs live on the servers under /opt/docker/compose/<stack>/ and /opt/docker/conf/<stack>/. This workspace mirrors them for version control, editing, and planning.
Conventions (enforce for every new stack)
Observed and standardized across servers:
- Compose location on server:
/opt/docker/compose/<stack>/compose.yaml - Config mounts on server:
/opt/docker/conf/<stack>/... - Networks: external
traefik-net, aliased astnetin composenetworks: tnet: name: traefik-net external: true - GPU reservation: prefer
deploy.resources.reservations.deviceswith explicitdevice_idsfor pinningdeploy: resources: reservations: devices: - driver: nvidia device_ids: ["1"] capabilities: [gpu] - Tunables:
.envin the same directory ascompose.yaml— keep the compose file constant, edit the.env - Named volumes for service state (pattern:
<stack>_<name>) - Bind mounts only for: model files (
/tank/aimodels/...), config files (/opt/docker/conf/...), docker socket where required - Restart policy:
restart: unless-stoppedfor daemons - Homepage labels on user-facing services. The dashboard runs on
esh-docker-vmand reads the Docker API of every host instacks/homepage/conf/docker.yaml(ana-docker, ana-ml2, nh3-docker, irv-ml1, esh-docker-vm), so a labelled container is discovered from wherever it runs — you do not add it toservices.yamlas well. Doing both renders it twice.⚠labels: - homepage.group=<ExistingGroup> - homepage.name=<ServiceName> - homepage.icon=mdi-<icon> - homepage.description=<short> - homepage.href=http://<host-ip>:<port>homepage.groupmust name a group that already exists instacks/homepage/conf/settings.yaml'slayout:block. A group the layout has never heard of gets notab:, and Homepage renders an untabbed group on all four tabs. Inventing a group name here is how Scriberr'sAI Systemsended up repeated at the bottom of every tab from 2026-08-23 (fixed 2026-08-24). If the service genuinely needs a new group, add the group tolayout:with atab:in the same change. Check withcurl -s http://10.0.50.45:5100/api/services | jq -r '.[].name'— anything in that list that is not a key inlayout:is leaking onto all tabs right now. Labels only apply at container creation, so a label edit needsdocker compose up -d <service>, notrestart. - Healthchecks on services that expose HTTP
Servers
| Name | IP | Site | Role | Details |
|---|---|---|---|---|
| ana-ml2 | 10.250.50.54 | Anaheim (10.250.0.0/16) |
GPU / AI inference (bare metal, dual RTX PRO 6000 Blackwell Max-Q, 96 GB each) | servers/ana-ml2/README.md |
| irv-ml1 | 10.100.79.3 (WG) | Irvine — reachable only via WireGuard tunnel from NH3 | GPU / AI inference (bare metal, RTX 3090 + RTX A6000, native stacks) | servers/irv-ml1/README.md |
| ana-docker | 10.250.50.70 | Anaheim | General-purpose Docker host (non-GPU VM on pfi-pve) | servers/ana-docker/README.md |
| pfi-ana-webhost | 10.250.50.52 | Anaheim | VM on pfi-pve (VMID 110) — web workload | servers/pfi-ana-webhost/README.md |
| ana-filebot | 10.250.50.53 | Anaheim | LXC on pfi-pve (CT 112) — file-task automation | servers/ana-filebot/README.md |
| pfi-pteradactyl | 10.250.50.55 | Anaheim | VM on pfi-pve (VMID 107) — Pterodactyl game panel | servers/pfi-pteradactyl/README.md |
| pfi-tacticalrmm | 10.250.50.57 | Anaheim | VM on pfi-pve (VMID 111) — TacticalRMM | servers/pfi-tacticalrmm/README.md |
| pfi-postgres | 10.250.50.80 | Anaheim | VM on pfi-pve (VMID 105) — shared Postgres (vaultwarden/gitea/paperless) | servers/pfi-postgres/README.md |
| ana-wg | 10.250.50.252 | Anaheim | LXC on pfi-pve (CT 113) — WireGuard | servers/ana-wg/README.md |
| pfi-pve | 10.250.250.31 | Anaheim | Proxmox VE hypervisor | servers/pfi-pve/README.md |
| pbs-ana | 10.250.50.90 | Anaheim | Proxmox Backup Server — fleet primary (VM on pfi-pve, NFS datastore on ana-nas) | servers/pbs-ana/README.md |
| sfsrv-ana | 10.250.250.115 | Anaheim | SureFire client (PFI-managed) — Proxmox VE hypervisor | servers/sfsrv-ana/README.md |
| sf-ana-container | 10.250.150.100 | Anaheim | SureFire client (PFI-managed) — container workload on sfsrv-ana | servers/sf-ana-container/README.md |
| sf-r630 | iDRAC 10.250.250.110 | Anaheim | SureFire client (PFI-managed) — physical Dell R630, iDRAC-managed from PFI side | servers/sf-r630/README.md |
| corviduo-dev | 10.250.50.152 | Anaheim | Worldtree-team dev VM (PFI-hosted) — runs the demo + personal + pinned Worldtree deployments vor/asset-engine talk to | servers/corviduo-dev/README.md |
| nh3-docker | 10.100.50.40 | NH3 (10.100.0.0/16) |
General-purpose Docker host (non-GPU VM on nh3-pve) | servers/nh3-docker/README.md |
| nh3-dev | 10.100.10.50 | NH3 | Dev box — fleet sidecars (egress SOCKS5 proxy, ttyd seat, mead-hall, volva) + live Claude Code sessions; not a Docker-stack host | servers/nh3-dev/README.md |
| nh3-extdev | 10.100.50.42 | NH3 | Manager / external-dev box (VM on nh3-pve, Debian 13); sudo-less infra-ops identity (user-level only, no Docker); successor to retired nh3-ansible | servers/nh3-extdev/README.md |
| nh3-pve | 10.100.250.60 | NH3 | Proxmox VE hypervisor | servers/nh3-pve/README.md |
| nh3-nas | 10.100.50.50 | NH3 | Synology RS2418+ — NFS exports, rest-server-nh3, PBS-NH3 datastore backend | servers/nh3-nas/README.md |
| pbs-nh3 | 10.100.50.90 | NH3 | Proxmox Backup Server — DR mirror (VM on nh3-pve, NFS datastore on nh3-nas); syncs from pbs-ana | servers/pbs-nh3/README.md |
| esh-docker-vm | 10.0.50.45 | ESH home lab (esteban.net, 10.0.50.0/24) |
Home-lab Docker host (VM on esh-pve) | servers/esh-docker-vm/README.md |
| vm-esh-nas | 10.0.50.154 | ESH home lab | NAS-adjacent Docker host (VM on esh-pve-nas) | servers/vm-esh-nas/README.md |
| esh-pve | 10.0.250.35 | ESH home lab | Proxmox VE hypervisor | servers/esh-pve/README.md |
| esh-pve-nas | 10.0.50.55 | ESH home lab | Proxmox VE hypervisor (storage / media) | servers/esh-pve-nas/README.md |
| esh-vm-db | 10.0.50.60 | ESH home lab | DB VM — PostgreSQL (paperless-ng) + MongoDB; bare-metal VM, no Docker | servers/esh-vm-db/README.md |
Placement rules:
- GPU-required stacks →
ana-ml2(primary, Anaheim) orirv-ml1(secondary, Irvine — bigger VRAM ceiling at 72 GB total). Access toirv-ml1requires WireGuard. - Anaheim non-GPU services →
ana-docker. - NH-site non-GPU services →
nh3-docker. - ESH home-lab workloads (
esteban.net) →esh-docker-vm(general) orvm-esh-nas(needs direct NFS mounts from 10.0.50.50). Not part of the PFI colo topology, but shares monitoring/backup tooling. - Cross-site services (e.g. Beszel hub, Dozzle hub) live on
ana-dockerand pull from agents on the other hosts. - SureFire (SF) client hosts (
sf-*,sfsrv-ana) are PFI-managed under the hosting agreement — SSH, OS ops, backups are PFI's responsibility. Hardware and data belong to the client; coordinate anything that affects data with them. - Worldtree-team dev VM (
corviduo-dev) is PFI-hosted (Anaheim subnet) but Worldtree-team-managed at the OS / application layer. PFI handles networking + emergency-ops backstop; OS configuration + deploy workflows + backup decisions live with the architect's team. Treat data-affecting work like SF hosts — coordinate before touching. - Hypervisors (
pfi-pve,nh3-pve,esh-pve,esh-pve-nas) are tracked for inventory / capacity planning. Don't deploy Docker stacks directly on them; new workloads land as VMs.server_inspect.shcaptures host-level detail only — VM/LXC/ZFS enumeration needs Proxmox-native tooling (qm list,pvesh get …,zpool list).
How to refresh a server's state
# Show help (no args)
scripts/refresh-server-info.sh
# Refresh every host discovered under servers/*/
scripts/refresh-server-info.sh all
# Refresh a specific host (must match a servers/<name>/ dir; ssh_config
# entry or servers/<name>/ssh-target handles how to reach it)
scripts/refresh-server-info.sh ana-docker
Fleet-wide runs require the literal all keyword — no-args prints help so you can't accidentally hit every host by forgetting a name.
The script pipes server_inspect.sh over SSH via stdin (no scp, no remote cleanup) and writes each servers/<host>/system-details.txt atomically — a failed run never clobbers the previous snapshot. The inspect script itself is read-only.
Each server dir can hold an ssh-target file (one line, <ip> or <user>@<ip>) as a fallback for when the dir name doesn't resolve via DNS or ~/.ssh/config. The script prefers whatever ssh would resolve normally and only consults the file when that fails.
To register a new server:
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name> # pull the first snapshot
To audit discovery without touching the network (checks permissions, unresolvable names with no fallback, missing README / system-details, malformed ssh-target):
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh --validate-only <host>
Stack tree convention (canonical vs mirror)
Two trees, distinct roles. They are NOT interchangeable.
| tree | role | git | who writes | who reads |
|---|---|---|---|---|
stacks/<stack>/ |
canonical / intent — source of truth for what we want deployed | tracked | you / Claude | deploy-stack.sh |
stacks-mirror/<host>/<stack>/ |
snapshot / reality — what's currently on each host | gitignored | sync-stacks.sh |
drift inspection |
Why two: keeps "intent" (committed, reviewable, deployed) cleanly separate from "reality on the server right now" (often drifts, useful to compare, not durable). Editing the mirror does NOT affect what gets deployed.
# Edit the canonical, then push it to the host:
# stacks/<stack>/<file> → /opt/docker/compose/<stack>/<file>
# stacks/<stack>/conf/<file> → /opt/docker/conf/<stack>/<file>
$EDITOR stacks/<stack>/compose.yaml
scripts/deploy-stack.sh <host> <stack> # diffs vs live, prompts y/N
scripts/deploy-stack.sh <host> <stack> --compose # skip conf side
scripts/deploy-stack.sh <host> <stack> --conf # skip compose side
# Pull current host state into the gitignored snapshot tree (drift check):
scripts/sync-stacks.sh # every host
scripts/sync-stacks.sh ana-docker # one host
scripts/sync-stacks.sh --dry-run # see what would change
# Compare canonical (intent) vs mirror (reality) for one stack:
diff -ru stacks/<stack>/ stacks-mirror/<host>/<stack>/
Opt-out per stack (mirror only — sync-stacks.sh skip): create stacks-mirror/<host>/<stack>/.no-sync (skip both sides) or stacks-mirror/<host>/<stack>/conf/.no-sync (skip conf only).
Always excluded in both directions (secrets / runtime state): .env, .env.*, acme.json, client_secrets.json, *.pem, *.key, *.crt, *.pfx, *.sqlite, *.sqlite3, *.db, *.log, *.log.*, *.pid, hub/, logs/.
Requires rsync installed on this workstation and every host you sync against (apt install rsync).
Layout
eshpfi-management/
├── CLAUDE.md # this file
├── README.md # human-facing overview
├── scripts/
│ └── server_inspect.sh # gather server state for compose planning
├── servers/
│ └── <name>/
│ ├── README.md
│ └── system-details.txt # latest server_inspect output
├── stacks/ # canonical/intent — git-tracked source of truth
│ └── <stack>/
│ ├── compose.yaml # deployed to /opt/docker/compose/<stack>/
│ ├── conf/<file> # deployed to /opt/docker/conf/<stack>/<file>
│ ├── .env.example # template; real .env lives on server
│ └── README.md # what this stack does, how to deploy
├── stacks-mirror/ # gitignored snapshot of live host state (drift detection)
│ └── <host>/<stack>/ # populated by sync-stacks.sh, NOT a deploy source
├── dns/ # fleet internal DNS — *.internal names
│ ├── internal.yaml # source of truth (hosts, sites, aliases)
│ └── README.md # workflow, naming, IPv6 caveat
└── docs/
└── pfi/ # general PFI infrastructure reference
Internal DNS (*.internal)
Fleet hosts have names: <host>.<site>.internal, sites ana / esh / nh3.
dns/internal.yaml is the source of truth; the AdGuard resolvers are derived
state.
$EDITOR dns/internal.yaml
scripts/dns-sync.py --dry-run # diff
scripts/dns-sync.py # apply
The sync is authoritative within .internal only — names added by hand in
the AdGuard UI get deleted, but rewrites in other zones (ESH's esteban.net
entries) are left alone. See dns/README.md, especially the IPv6 note: v6
addresses only go in the file once they are pinned statically on the host,
because SLAAC addresses rotate and a stale record is worse than none.
Working rules
- Copies, not symlinks. Files here reflect what's on the server at the time of the last sync. When you edit here, the server doesn't change until you deploy.
- Never commit secrets. Use
.env.exampletemplates; real.envfiles (with tokens, passwords) live on the server and are gitignored if/when this becomes a git repo. - Surgical edits. When fixing one stack, don't touch unrelated ones. Follow AIPA-MCP's CLAUDE.md rules about scope discipline.
- Sanity-check before deploying. Run
docker compose config(dry parse) beforedocker compose up -don the server.