# Status + Open Issues Last updated: 2026-04-24 Snapshot of fleet state and open work. Refresh this file when a pass of significant work lands — don't let it drift quietly. **New session? Read [`docs/orientation.md`](docs/orientation.md) first.** ## What's in place ### Backup coverage (2-layer, fully operational) - **PBS fleet-wide.** All 5 hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana) → `pbs-ana` (primary, VM on pfi-pve with NFSv3 datastore on ana-nas). Per-hypervisor namespaces to avoid VMID collisions. DR mirror at `pbs-nh3` (VM on nh3-pve with NFSv3 datastore on nh3-nas Synology) pulls daily at 06:00. Verify jobs armed on both sides. - **restic 8/8 hosts** with resticprofile + systemd timers at 01:00: | Host | Target rest-server | DB hooks | |---|---|---| | ana-docker | rest-server-ana (local) | synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui | | ana-ml2 | rest-server-ana (cross-site) | — (no DBs) | | nh3-docker | rest-server-nh3 (local, Synology) | — | | esh-docker-vm | rest-server-ana (cross-site) | paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite | | vm-esh-nas | rest-server-ana (cross-site) | — | | nh3-dev (workstation) | rest-server-nh3 (local) | — | | irv-ml1 | rest-server-nh3 (via WG) | — | | esh-vm-db | rest-server-ana (cross-site) | pg_dumpall, mongodump | - **Cross-site restic rsync.** `ana-nas → nh3-nas` at 04:00 daily (mirrors rest-server-ana data). `nh3-nas → ana-nas` at 05:00 daily (mirrors rest-server-nh3 data, runs as root since DSM writes files as admin mode 400). Tracked at `configs/rsync/{ana-nas-to-nh3,nh3-nas-to-ana}/`. ### Inventory - **~22 tracked hosts** under `servers/` across ANA / NH3 / ESH / IRV sites. SSH config aliases in `~/.ssh/config` for every host — `ssh ` just works. - **Homepage** at — function-first layout (Main / Infrastructure / Toolchain tabs), per-group icons, four Infra sub-groups (ANA / NH3 / IRV / ESH). Docker auto-discovery active for 5 hosts (ana-docker, ana-ml2, nh3-docker, esh-docker-vm, irv-ml1). - **FortiGate + UniFi discovery scripts** produce TSVs; gap analysis against `servers/*/` for unmanaged IPs. ### Architecture decisions (durable) - **DB data on local disk, not NFS.** pfi-postgres migrated 2026-04-23; fstab + NFS mount removed 2026-04-24 (pfi-postgres now fully decoupled from ana-nas). Mongo on esh-vm-db confirmed already local. Removes the biggest ana-nas blast-radius risk. Memory: `project_db_migrate_off_nfs.md`. - **Backups must not risk production.** Rule adopted after 2026-04-23 ana-nas self-backup crash. Memory: `feedback_backups_must_not_risk_production.md`. - **offen sidecars retired fleet-wide 2026-04-23.** restic covers equivalent scope; reclaimed 16 GB of redundant tarballs on `/mnt/backup/docker/esh-vm-docker/`. - **Gitea remote for this repo (2026-04-23).** `origin` is `vh/esh-pfi-infrastructure` on `gitea.phasefinal.com`. Also: `vh/task-board` (new 2026-04-24) hosts a Claude Code plugin + marketplace serving the assistant task-state dashboard. - **Prefer elway for multi-step SSH work (2026-04-24).** The new `scripts/elway` mini-playbook runner replaces chained `ssh -t sudo …` commands — handles sudo once up front, structured reporting, idempotency (creates/when/changed_when). Memory: `feedback_use_elway.md`; template playbook: `playbooks/elway-smoke.yaml`. ### Tooling - `scripts/server_inspect.sh` + `refresh-server-info.sh` for Docker hosts. - `scripts/proxmox_inspect.sh` + `refresh-proxmox-info.sh` for PVE nodes (reports VMs/LXCs/storage/backup-coverage). - `scripts/discover-fortigate.sh`, `scripts/discover-unifi.sh`, `scripts/discover-gaps.sh` for network-level inventory discovery. - `scripts/deploy-stack.sh` + `scripts/sync-stacks.sh` for compose push/pull. - `scripts/add-host.sh` for new-host registration. - **`scripts/elway`** (2026-04-24) — mini-ansible playbook runner over SSH. Playbooks live under `playbooks/`. Tier 1 (creates/when) + tier 2 (changed_when) idempotency; handlers + register + multi-host fan-out deferred as gitea issues #3–#5. - **`tea` CLI** at `/usr/local/bin/tea` — gitea CLI, already logged in as `vh`. Use for issue / PR work instead of inventing URLs. ## Open issues ### 🟥 Quick wins (do next) 1. ~~**Fix Backrest's `esh-docker-vm` URI**~~ — done. 2. ~~**Push SSH keys + pull snapshots for 6 unrefreshed hosts**~~ — **done 2026-04-23**. 4 VMs (pfi-ana-webhost, pfi-pteradactyl, pfi-tacticalrmm, pfi-postgres) took the workstation key via `ssh-copy-id` with `lkraven@`. 2 LXCs (ana-filebot, ana-wg) needed key installed via `pct push` from pfi-pve because `PermitRootLogin prohibit-password` blocked `ssh-copy-id`. Minor: pfi-pteradactyl's `server_inspect` Docker section is blank because lkraven is not in the `docker` group there — `sudo usermod -aG docker lkraven` + re-login to fix. 3. ~~**Patch out `forget` schedules from all 6 restic profiles**~~ — done. 3b. ~~**Discover sf-r630 OS-side IP**~~ — **resolved 2026-04-23**: the R630 with iDRAC `10.250.250.110` is the same physical box that runs `sfsrv-ana` (Proxmox VE at `10.250.250.115`). No separate OS IP to find. `servers/sf-r630/` now clarifies it as the hardware/BMC-only inventory entry; `servers/sfsrv-ana/` is the OS view. `servers/ana-ml2/README.md` updated with its own distinct BMC IP (`10.250.250.50`, Supermicro) to prevent future confusion between the two physical chassis. 4. _(removed — see note above)_ ### 🟧 Real work (dedicated session each) 4b. ~~**Migrate DB data directories off NFS onto local VM disk.**~~ — **done 2026-04-23/24.** Postgres on pfi-postgres migrated 04-23; mongodb on esh-vm-db confirmed already local (and serves zero user data in practice — paperless uses Postgres 15 on the same VM). fstab entry + NFS mount for `/mnt/db` on pfi-postgres removed 2026-04-24 via `playbooks/decouple-pfi-postgres-from-ana-nas.yaml`. pfi-postgres has zero remaining dependency on ana-nas. Only residue: cold archive of pre-migration data still on ana-nas (`/mnt/db/pfi-*`); harmless, can sit indefinitely. 5. ~~**Rotate exposed secrets**~~ — **done 2026-04-23**. All six rotated: vaultwarden/gitea/paperless-ng Postgres passwords (hardcoded `compose.yaml` literals moved to gitignored `.env` files in the process), ana-docker + ana-ml2 + esh-docker-vm rest-server htpasswd entries, and rest-server repo passphrases for ana-docker + esh-docker-vm (ana-ml2's was already rotated during prior wipe+reinit). Also discovered along the way: paperless uses a separate `esh-vm-db` VM (10.0.50.60), not pfi-postgres as the old notes implied. 6. ~~**Cross-site rsync**~~ — **done 2026-04-23, both directions**. - **ana → nh3** (04:00 daily): `ana-nas:/mnt/backup/restic/repo/ana/` → `nh3-nas:/volume1/Backup/restic-ana-mirror/`. Runs on ana-nas as lkraven. 15.6 GB initial sync completed 07:18 UTC. - **nh3 → ana** (05:00 daily): `nh3-nas:/volume1/Backup/restic/` → `ana-nas:/mnt/backup/restic-nh3-mirror/`. Runs on nh3-nas as root (rest-server-nh3 container writes mode-400 files; only root can read them on Synology). 2.31 GB initial sync completed 07:42 UTC. - Tracked at `configs/rsync/ana-nas-to-nh3/` and `configs/rsync/nh3-nas-to-ana/` respectively. 6b. **PBS deployment across the fleet.** Runbook at `docs/runbooks/pbs-deployment.md`. Phases 0–6 **done** (2026-04-22): PBS-ANA on pfi-pve with NFS datastore on Ana NAS (NFSv3 for ZFS case-insensitivity), PBS-NH3 on nh3-pve with NFS datastore on NH3 Synology (NFSv3 + Linux-mode POSIX 777 to avoid syno_acl interference), per-hypervisor namespaces, API tokens, verify jobs on both, one-way sync ANA → NH3 at 06:00 daily, and all 5 hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana) onboarded to PBS-ANA. Remaining phases: - Phase 7-8 — one week burn-in, then retire legacy vzdump targets on each hypervisor (keep until 2026-04-29 earliest) - Phase 9 — homepage cards for PBS-ANA + PBS-NH3 web UIs, memory/doc updates, `servers/pbs-ana` + `servers/pbs-nh3` dirs 7. ~~**SureFire tenant backup plan decision**~~ — **resolved 2026-04-23** by folding sfsrv-ana into the fleet-wide PBS deployment (dedicated `sfsrv-ana` namespace on PBS-ANA, replicates to PBS-NH3 via the same sync job as the rest of the fleet). Hosting-agreement option chosen: PFI provides backup coverage as part of managed hosting. ### 🟨 Prereqs / polish 8. ~~**Synology SSH setup**~~ — **done 2026-04-22**. Dedicated `syncuser` account (admin-group membership) with key auth, registered as `servers/nh3-nas/`, reachable as `ssh nh3-nas`. Unlocks cross-site rsync (item 6), rest-server-nh3 healthcheck deploy, and `.htpasswd` edits on the NH3 side. 9. **`scripts/restic-prune.sh`** — temporarily flip `--append-only` off, run forget + prune across all hosts, flip back on. Needed quarterly for disk hygiene. Not urgent; blocks only the "I need to reclaim disk space now" scenario. 10. ~~**Retire `offen/docker-volume-backup` sidecars**~~ — **done 2026-04-23**. Removed from paperless-ngx and pgadmin composes on esh-docker-vm (only hosts in the fleet that had them). 16 GB of orphan tarballs at `/mnt/backup/docker/esh-vm-docker/` reclaimed. Restic coverage verified equivalent: `/var/lib/docker/volumes` in source + pre-backup hook handles the DB dumps for paperless (against esh-vm-db), HA, pgadmin, CWA, uptime-kuma. Obsolete `*_offen_backup_data` exclude also removed from the restic profile. 11. **Clean up retired mattermost dir** on ana-docker — compose dir at `/opt/docker/compose/mattermost/` may still linger. Single `rm -rf` when convenient. ### 🟩 Research / deferred / intermittent 12. **Backrest UI intermittent timeout.** Backend is confirmed healthy (direct `curl GetConfig` returns in 200ms). UI chokes on the same endpoint. Likely cause: SSE stream wedge in browser, or service worker stale. Workaround: Ctrl+Shift+R. Root cause investigation deferred. 13. **UniFi controller homepage cards** (ESH-UDMPM at 10.0.0.1, PFI-UDMSE at 10.100.0.1). Auto-discovered via the Site Manager API; not yet linked as homepage cards. 14. **Prune + credential-rotation scripts** as repeatable tooling (vs per-incident manual work). ### 🟦 Memory / documentation housekeeping 15. **`docs/` organization** could use a pass — multiple READMEs and reference files in different spots. Not urgent. 16. **`STATUS.md` (this file) drift.** Update whenever significant work lands. Consider a `scripts/status-regen.sh` if manual updates slip. ## Session milestones — 2026-04-24 (the "tooling day") - **irv-ml1 AI stacks deployed** as Docker: ComfyUI (port 8188, host- writable workflows), Parakeet ASR (port 8765, rewritten on sherpa-onnx after the Shadowfita FastAPI wrapper hit two unfixed upstream bugs), CosyVoice 3 (port 8190, GLaDOS voice cloned + live). - **irv-ml1 restic profile extended** to cover `/worktank/{comfyui/basedir/{user,custom_nodes,input},cosyvoice/voices}`; bulk weights + scratch dirs stay excluded. - **`scripts/elway`** shipped — ~800-line Python playbook runner with tier 1 + tier 2 idempotency; handlers / register / multi-host / content-hash are deferred (gitea #3–#6). - **task-board** built end-to-end (separate repo, `vh/task-board` on gitea) and shipped as a Claude Code plugin. Green/red cards per session via UserPromptSubmit + Stop hooks; four MCP tools expose explicit activity tracking. - **pfi-postgres NFS decoupling** finished item 4b — zero residual dependency on ana-nas for that VM. ## Session milestones — 2026-04-20 / 2026-04-21 - Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain). - Per-group icons + equal-height layout + 4-column grids. - Fleet-wide label sweep (function groups, Service Networking renamed from Wiring/Plumbing due to homepage-parser slash bug). - 6 restic clients deployed end-to-end (client creds, repo init, profile install, systemd timers, first backups verified). - 9 hosts registered under `servers/` from gap-analysis (PFI VMs + SureFire tenant, with tenancy awareness). - Discovery scripts (FortiGate SSH + UniFi cloud API + TSV diff vs inventory). - Proxmox inspect script + fleet-wide refresh wrapper. - Calibre-Web-Automated migration replacing calibre + calibre-web pair. - llama-swap Qwen3.6 stock + mradermacher abliterated, 128K ctx. - Mattermost retired (not running, compose dir cleanup pending). - FortiGate 101F at NH3 retired; homepage card removed. ## Memory pointers (for future Claude sessions) Relevant `~/.claude/.../memory/` entries: - `server_split.md` — host placement rules - `feedback_ssh_sudo.md` — use `ssh -t` for remote sudo (mostly superseded by elway, but still applies to ad-hoc ssh) - `feedback_git_autonomous.md` — handle git commits without asking - `feedback_git_commits.md` — no Claude attribution in commit messages - `feedback_use_elway.md` — write elway playbooks; don't chain ssh+sudo - `feedback_backups_must_not_risk_production.md` — rule adopted after the 2026-04-23 ana-nas self-backup crash - `project_backup_pipeline_gaps.md` — user's explicit goal of "all hosts + configs + DBs backed up" - `project_db_migrate_off_nfs.md` — DB-off-NFS decision + status - `project_surefire_tenant.md` — SureFire tenancy boundary awareness - `reference_gitea_remote.md` — origin is `vh/esh-pfi-infrastructure` on gitea.phasefinal.com; `tea` CLI logged in as `vh` - `reference_task_board.md` — task-board plugin + tools contract - `storage_ana_nas.md` — ana NAS is Debian LXC (CT 109), not TrueNAS - `storage_nh3_nas.md` — NH3 NAS via `syncuser`, not `admin` - `incident_ana_nas_spof.md` — blast-radius matrix for ana-nas outages