STATUS.md:
- Mark 4b done (both Postgres migration + NFS decoupling)
- Add arch decisions for gitea remote + prefer-elway policy
- Add tooling entries for elway + tea CLI
- Document 2026-04-24 session milestones (irv-ml1 AI stacks,
elway, task-board, 4b finish)
- Expand memory-pointer list with the files added this session
CLAUDE.md:
- Tell new sessions to use elway for SSH-driven work, point at
the smoke playbook template
- Document the task-board plugin + MCP-tool contract so assistant
sessions with the plugin enabled know the assistant should call
task_start / task_update / task_wait / task_complete at
meaningful checkpoints
.claude/settings.json:
- Project-level env: TASK_BOARD_SESSION=Infra so every Claude Code
session opened here labels its task-board cards "Infra"
playbooks/decouple-pfi-postgres-from-ana-nas.yaml:
- Finishes the DB-off-NFS migration on pfi-postgres. Already ran
against prod today; fstab clean, unmounted, no systemd mnt-db
unit. Verify 3 was mis-expressed on first run (`grep -q active`
matched "inactive") — fixed to invert systemctl exit code
directly.
14 KiB
Status + Open Issues
Last updated: 2026-04-24
Snapshot of fleet state and open work. Refresh this file when a pass of significant work lands — don't let it drift quietly.
New session? Read docs/orientation.md first.
What's in place
Backup coverage (2-layer, fully operational)
- PBS fleet-wide. All 5 hypervisors (pfi-pve, nh3-pve, esh-pve,
esh-pve-nas, sfsrv-ana) →
pbs-ana(primary, VM on pfi-pve with NFSv3 datastore on ana-nas). Per-hypervisor namespaces to avoid VMID collisions. DR mirror atpbs-nh3(VM on nh3-pve with NFSv3 datastore on nh3-nas Synology) pulls daily at 06:00. Verify jobs armed on both sides. - restic 8/8 hosts with resticprofile + systemd timers at 01:00:
Host Target rest-server DB hooks ana-docker rest-server-ana (local) synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui ana-ml2 rest-server-ana (cross-site) — (no DBs) nh3-docker rest-server-nh3 (local, Synology) — esh-docker-vm rest-server-ana (cross-site) paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite vm-esh-nas rest-server-ana (cross-site) — nh3-dev (workstation) rest-server-nh3 (local) — irv-ml1 rest-server-nh3 (via WG) — esh-vm-db rest-server-ana (cross-site) pg_dumpall, mongodump - Cross-site restic rsync.
ana-nas → nh3-nasat 04:00 daily (mirrors rest-server-ana data).nh3-nas → ana-nasat 05:00 daily (mirrors rest-server-nh3 data, runs as root since DSM writes files as admin mode 400). Tracked atconfigs/rsync/{ana-nas-to-nh3,nh3-nas-to-ana}/.
Inventory
- ~22 tracked hosts under
servers/across ANA / NH3 / ESH / IRV sites. SSH config aliases in~/.ssh/configfor every host —ssh <name>just works. - Homepage at http://10.0.50.45:5100 — function-first layout (Main / Infrastructure / Toolchain tabs), per-group icons, four Infra sub-groups (ANA / NH3 / IRV / ESH). Docker auto-discovery active for 5 hosts (ana-docker, ana-ml2, nh3-docker, esh-docker-vm, irv-ml1).
- FortiGate + UniFi discovery scripts produce TSVs; gap analysis
against
servers/*/for unmanaged IPs.
Architecture decisions (durable)
- DB data on local disk, not NFS. pfi-postgres migrated 2026-04-23;
fstab + NFS mount removed 2026-04-24 (pfi-postgres now fully
decoupled from ana-nas). Mongo on esh-vm-db confirmed already local.
Removes the biggest ana-nas blast-radius risk. Memory:
project_db_migrate_off_nfs.md. - Backups must not risk production. Rule adopted after
2026-04-23 ana-nas self-backup crash. Memory:
feedback_backups_must_not_risk_production.md. - offen sidecars retired fleet-wide 2026-04-23. restic covers
equivalent scope; reclaimed 16 GB of redundant tarballs on
/mnt/backup/docker/esh-vm-docker/. - Gitea remote for this repo (2026-04-23).
originisvh/esh-pfi-infrastructureongitea.phasefinal.com. Also:vh/task-board(new 2026-04-24) hosts a Claude Code plugin + marketplace serving the assistant task-state dashboard. - Prefer elway for multi-step SSH work (2026-04-24). The new
scripts/elwaymini-playbook runner replaces chainedssh -t sudo …commands — handles sudo once up front, structured reporting, idempotency (creates/when/changed_when). Memory:feedback_use_elway.md; template playbook:playbooks/elway-smoke.yaml.
Tooling
scripts/server_inspect.sh+refresh-server-info.shfor Docker hosts.scripts/proxmox_inspect.sh+refresh-proxmox-info.shfor PVE nodes (reports VMs/LXCs/storage/backup-coverage).scripts/discover-fortigate.sh,scripts/discover-unifi.sh,scripts/discover-gaps.shfor network-level inventory discovery.scripts/deploy-stack.sh+scripts/sync-stacks.shfor compose push/pull.scripts/add-host.shfor new-host registration.scripts/elway(2026-04-24) — mini-ansible playbook runner over SSH. Playbooks live underplaybooks/. Tier 1 (creates/when) + tier 2 (changed_when) idempotency; handlers + register + multi-host fan-out deferred as gitea issues #3–#5.teaCLI at/usr/local/bin/tea— gitea CLI, already logged in asvh. Use for issue / PR work instead of inventing URLs.
Open issues
🟥 Quick wins (do next)
-
Fix Backrest's— done.esh-docker-vmURI -
Push SSH keys + pull snapshots for 6 unrefreshed hosts— done 2026-04-23. 4 VMs (pfi-ana-webhost, pfi-pteradactyl, pfi-tacticalrmm, pfi-postgres) took the workstation key viassh-copy-idwithlkraven@. 2 LXCs (ana-filebot, ana-wg) needed key installed viapct pushfrom pfi-pve becausePermitRootLogin prohibit-passwordblockedssh-copy-id. Minor: pfi-pteradactyl'sserver_inspectDocker section is blank because lkraven is not in thedockergroup there —sudo usermod -aG docker lkraven+ re-login to fix. -
Patch out— done.forgetschedules from all 6 restic profiles
3b. Discover sf-r630 OS-side IP — resolved 2026-04-23:
the R630 with iDRAC 10.250.250.110 is the same physical box
that runs sfsrv-ana (Proxmox VE at 10.250.250.115). No
separate OS IP to find. servers/sf-r630/ now clarifies it as
the hardware/BMC-only inventory entry; servers/sfsrv-ana/ is
the OS view. servers/ana-ml2/README.md updated with its own
distinct BMC IP (10.250.250.50, Supermicro) to prevent future
confusion between the two physical chassis.
- (removed — see note above)
🟧 Real work (dedicated session each)
4b. Migrate DB data directories off NFS onto local VM disk. —
done 2026-04-23/24. Postgres on pfi-postgres migrated
04-23; mongodb on esh-vm-db confirmed already local (and serves
zero user data in practice — paperless uses Postgres 15 on the
same VM). fstab entry + NFS mount for /mnt/db on pfi-postgres
removed 2026-04-24 via
playbooks/decouple-pfi-postgres-from-ana-nas.yaml. pfi-postgres
has zero remaining dependency on ana-nas. Only residue: cold
archive of pre-migration data still on ana-nas (/mnt/db/pfi-*);
harmless, can sit indefinitely.
-
Rotate exposed secrets— done 2026-04-23. All six rotated: vaultwarden/gitea/paperless-ng Postgres passwords (hardcodedcompose.yamlliterals moved to gitignored.envfiles in the process), ana-docker + ana-ml2 + esh-docker-vm rest-server htpasswd entries, and rest-server repo passphrases for ana-docker + esh-docker-vm (ana-ml2's was already rotated during prior wipe+reinit). Also discovered along the way: paperless uses a separateesh-vm-dbVM (10.0.50.60), not pfi-postgres as the old notes implied. -
Cross-site rsync— done 2026-04-23, both directions.- ana → nh3 (04:00 daily):
ana-nas:/mnt/backup/restic/repo/ana/→nh3-nas:/volume1/Backup/restic-ana-mirror/. Runs on ana-nas as lkraven. 15.6 GB initial sync completed 07:18 UTC. - nh3 → ana (05:00 daily):
nh3-nas:/volume1/Backup/restic/→ana-nas:/mnt/backup/restic-nh3-mirror/. Runs on nh3-nas as root (rest-server-nh3 container writes mode-400 files; only root can read them on Synology). 2.31 GB initial sync completed 07:42 UTC. - Tracked at
configs/rsync/ana-nas-to-nh3/andconfigs/rsync/nh3-nas-to-ana/respectively.
- ana → nh3 (04:00 daily):
6b. PBS deployment across the fleet. Runbook at
docs/runbooks/pbs-deployment.md. Phases 0–6 done (2026-04-22):
PBS-ANA on pfi-pve with NFS datastore on Ana NAS (NFSv3 for ZFS
case-insensitivity), PBS-NH3 on nh3-pve with NFS datastore on NH3
Synology (NFSv3 + Linux-mode POSIX 777 to avoid syno_acl
interference), per-hypervisor namespaces, API tokens, verify jobs
on both, one-way sync ANA → NH3 at 06:00 daily, and all 5
hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana)
onboarded to PBS-ANA.
Remaining phases:
- Phase 7-8 — one week burn-in, then retire legacy vzdump targets
on each hypervisor (keep until 2026-04-29 earliest)
- Phase 9 — homepage cards for PBS-ANA + PBS-NH3 web UIs,
memory/doc updates, servers/pbs-ana + servers/pbs-nh3 dirs
SureFire tenant backup plan decision— resolved 2026-04-23 by folding sfsrv-ana into the fleet-wide PBS deployment (dedicatedsfsrv-ananamespace on PBS-ANA, replicates to PBS-NH3 via the same sync job as the rest of the fleet). Hosting-agreement option chosen: PFI provides backup coverage as part of managed hosting.
🟨 Prereqs / polish
-
Synology SSH setup— done 2026-04-22. Dedicatedsyncuseraccount (admin-group membership) with key auth, registered asservers/nh3-nas/, reachable asssh nh3-nas. Unlocks cross-site rsync (item 6), rest-server-nh3 healthcheck deploy, and.htpasswdedits on the NH3 side. -
scripts/restic-prune.sh— temporarily flip--append-onlyoff, run forget + prune across all hosts, flip back on. Needed quarterly for disk hygiene. Not urgent; blocks only the "I need to reclaim disk space now" scenario. -
Retire— done 2026-04-23. Removed from paperless-ngx and pgadmin composes on esh-docker-vm (only hosts in the fleet that had them). 16 GB of orphan tarballs atoffen/docker-volume-backupsidecars/mnt/backup/docker/esh-vm-docker/reclaimed. Restic coverage verified equivalent:/var/lib/docker/volumesin source + pre-backup hook handles the DB dumps for paperless (against esh-vm-db), HA, pgadmin, CWA, uptime-kuma. Obsolete*_offen_backup_dataexclude also removed from the restic profile. -
Clean up retired mattermost dir on ana-docker — compose dir at
/opt/docker/compose/mattermost/may still linger. Singlerm -rfwhen convenient.
🟩 Research / deferred / intermittent
-
Backrest UI intermittent timeout. Backend is confirmed healthy (direct
curl GetConfigreturns in 200ms). UI chokes on the same endpoint. Likely cause: SSE stream wedge in browser, or service worker stale. Workaround: Ctrl+Shift+R. Root cause investigation deferred. -
UniFi controller homepage cards (ESH-UDMPM at 10.0.0.1, PFI-UDMSE at 10.100.0.1). Auto-discovered via the Site Manager API; not yet linked as homepage cards.
-
Prune + credential-rotation scripts as repeatable tooling (vs per-incident manual work).
🟦 Memory / documentation housekeeping
-
docs/organization could use a pass — multiple READMEs and reference files in different spots. Not urgent. -
STATUS.md(this file) drift. Update whenever significant work lands. Consider ascripts/status-regen.shif manual updates slip.
Session milestones — 2026-04-24 (the "tooling day")
- irv-ml1 AI stacks deployed as Docker: ComfyUI (port 8188, host- writable workflows), Parakeet ASR (port 8765, rewritten on sherpa-onnx after the Shadowfita FastAPI wrapper hit two unfixed upstream bugs), CosyVoice 3 (port 8190, GLaDOS voice cloned + live).
- irv-ml1 restic profile extended to cover
/worktank/{comfyui/basedir/{user,custom_nodes,input},cosyvoice/voices}; bulk weights + scratch dirs stay excluded. scripts/elwayshipped — ~800-line Python playbook runner with tier 1 + tier 2 idempotency; handlers / register / multi-host / content-hash are deferred (gitea #3–#6).- task-board built end-to-end (separate repo,
vh/task-boardon gitea) and shipped as a Claude Code plugin. Green/red cards per session via UserPromptSubmit + Stop hooks; four MCP tools expose explicit activity tracking. - pfi-postgres NFS decoupling finished item 4b — zero residual dependency on ana-nas for that VM.
Session milestones — 2026-04-20 / 2026-04-21
- Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
- Per-group icons + equal-height layout + 4-column grids.
- Fleet-wide label sweep (function groups, Service Networking renamed from Wiring/Plumbing due to homepage-parser slash bug).
- 6 restic clients deployed end-to-end (client creds, repo init, profile install, systemd timers, first backups verified).
- 9 hosts registered under
servers/from gap-analysis (PFI VMs + SureFire tenant, with tenancy awareness). - Discovery scripts (FortiGate SSH + UniFi cloud API + TSV diff vs inventory).
- Proxmox inspect script + fleet-wide refresh wrapper.
- Calibre-Web-Automated migration replacing calibre + calibre-web pair.
- llama-swap Qwen3.6 stock + mradermacher abliterated, 128K ctx.
- Mattermost retired (not running, compose dir cleanup pending).
- FortiGate 101F at NH3 retired; homepage card removed.
Memory pointers (for future Claude sessions)
Relevant ~/.claude/.../memory/ entries:
server_split.md— host placement rulesfeedback_ssh_sudo.md— usessh -tfor remote sudo (mostly superseded by elway, but still applies to ad-hoc ssh)feedback_git_autonomous.md— handle git commits without askingfeedback_git_commits.md— no Claude attribution in commit messagesfeedback_use_elway.md— write elway playbooks; don't chain ssh+sudofeedback_backups_must_not_risk_production.md— rule adopted after the 2026-04-23 ana-nas self-backup crashproject_backup_pipeline_gaps.md— user's explicit goal of "all hosts + configs + DBs backed up"project_db_migrate_off_nfs.md— DB-off-NFS decision + statusproject_surefire_tenant.md— SureFire tenancy boundary awarenessreference_gitea_remote.md— origin isvh/esh-pfi-infrastructureon gitea.phasefinal.com;teaCLI logged in asvhreference_task_board.md— task-board plugin + tools contractstorage_ana_nas.md— ana NAS is Debian LXC (CT 109), not TrueNASstorage_nh3_nas.md— NH3 NAS viasyncuser, notadminincident_ana_nas_spof.md— blast-radius matrix for ana-nas outages