Files
esh-pfi-infrastructure/STATUS.md
T
vh 1f14c6d959 scripts: restic-prune.sh — quarterly forget + prune ceremony (closes #9)
Toggles --append-only off on the rest-server via a temporary
docker-compose.override.yaml (canonical compose untouched), runs
resticprofile forget --prune --verbose on each client of that
rest-server, then restores --append-only. The restore is wrapped in
a trap so a partial-failure prune still leaves the rest-server in
its safe configuration.

ANA side is fully automated against ana-docker (5 clients:
ana-docker, ana-ml2, esh-docker-vm, vm-esh-nas, esh-vm-db).

NH3 side currently prints a manual DSM ceremony — Synology Container
Manager doesn't expose docker on the expected paths and syncuser
sudo isn't NOPASSWD, so the toggle isn't safely scriptable from
this workstation. The instructions cover the same flow in DSM web
UI + interactive ssh on each NH3 client (nh3-docker, nh3-dev,
irv-ml1).

Usage:
  scripts/restic-prune.sh ana    # ANA only (auto)
  scripts/restic-prune.sh nh3    # NH3 instructions
  scripts/restic-prune.sh all    # both
  scripts/restic-prune.sh -h     # help
  scripts/restic-prune.sh --dry-run ana   # show every command
2026-04-24 22:01:37 -07:00

18 KiB
Raw Blame History

Status + Open Issues

Last updated: 2026-04-24

Snapshot of fleet state and open work. Refresh this file when a pass of significant work lands — don't let it drift quietly.

New session? Read docs/orientation.md first.

What's in place

Backup coverage (2-layer, fully operational)

  • PBS fleet-wide. All 5 hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana) → pbs-ana (primary, VM on pfi-pve with NFSv3 datastore on ana-nas). Per-hypervisor namespaces to avoid VMID collisions. DR mirror at pbs-nh3 (VM on nh3-pve with NFSv3 datastore on nh3-nas Synology) pulls daily at 06:00. Verify jobs armed on both sides.
  • restic 8/8 hosts with resticprofile + systemd timers at 01:00:
    Host Target rest-server DB hooks
    ana-docker rest-server-ana (local) synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui
    ana-ml2 rest-server-ana (cross-site) — (no DBs)
    nh3-docker rest-server-nh3 (local, Synology)
    esh-docker-vm rest-server-ana (cross-site) paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite
    vm-esh-nas rest-server-ana (cross-site)
    nh3-dev (workstation) rest-server-nh3 (local)
    irv-ml1 rest-server-nh3 (via WG)
    esh-vm-db rest-server-ana (cross-site) pg_dumpall, mongodump
  • Cross-site restic rsync. ana-nas → nh3-nas at 04:00 daily (mirrors rest-server-ana data). nh3-nas → ana-nas at 05:00 daily (mirrors rest-server-nh3 data, runs as root since DSM writes files as admin mode 400). Tracked at configs/rsync/{ana-nas-to-nh3,nh3-nas-to-ana}/.

Inventory

  • ~22 tracked hosts under servers/ across ANA / NH3 / ESH / IRV sites. SSH config aliases in ~/.ssh/config for every host — ssh <name> just works.
  • Homepage at http://10.0.50.45:5100 — function-first layout (Main / Infrastructure / Toolchain tabs), per-group icons, four Infra sub-groups (ANA / NH3 / IRV / ESH). Docker auto-discovery active for 5 hosts (ana-docker, ana-ml2, nh3-docker, esh-docker-vm, irv-ml1).
  • FortiGate + UniFi discovery scripts produce TSVs; gap analysis against servers/*/ for unmanaged IPs.

Architecture decisions (durable)

  • DB data on local disk, not NFS. pfi-postgres migrated 2026-04-23; fstab + NFS mount removed 2026-04-24 (pfi-postgres now fully decoupled from ana-nas). Mongo on esh-vm-db confirmed already local. Removes the biggest ana-nas blast-radius risk. Memory: project_db_migrate_off_nfs.md.
  • Backups must not risk production. Rule adopted after 2026-04-23 ana-nas self-backup crash. Memory: feedback_backups_must_not_risk_production.md.
  • offen sidecars retired fleet-wide 2026-04-23. restic covers equivalent scope; reclaimed 16 GB of redundant tarballs on /mnt/backup/docker/esh-vm-docker/.
  • Gitea remote for this repo (2026-04-23). origin is vh/esh-pfi-infrastructure on gitea.phasefinal.com. Also: vh/task-board (new 2026-04-24) hosts a Claude Code plugin + marketplace serving the assistant task-state dashboard.
  • Prefer elway for multi-step SSH work (2026-04-24). The new scripts/elway mini-playbook runner replaces chained ssh -t sudo … commands — handles sudo once up front, structured reporting, idempotency (creates/when/changed_when). Memory: feedback_use_elway.md; template playbook: playbooks/elway-smoke.yaml.

Tooling

  • scripts/server_inspect.sh + refresh-server-info.sh for Docker hosts.
  • scripts/proxmox_inspect.sh + refresh-proxmox-info.sh for PVE nodes (reports VMs/LXCs/storage/backup-coverage).
  • scripts/discover-fortigate.sh, scripts/discover-unifi.sh, scripts/discover-gaps.sh for network-level inventory discovery.
  • scripts/deploy-stack.sh + scripts/sync-stacks.sh for compose push/pull.
  • scripts/add-host.sh for new-host registration.
  • scripts/elway (2026-04-24) — mini-ansible playbook runner over SSH. Playbooks live under playbooks/. Tier 1 (creates/when) + tier 2 (changed_when) idempotency; handlers + register + multi-host fan-out deferred as gitea issues #3#5.
  • tea CLI at /usr/local/bin/tea — gitea CLI, already logged in as vh. Use for issue / PR work instead of inventing URLs.

Open issues

🟥 Quick wins (do next)

  1. Fix Backrest's esh-docker-vm URI — done.

  2. Push SSH keys + pull snapshots for 6 unrefreshed hostsdone 2026-04-23. 4 VMs (pfi-ana-webhost, pfi-pteradactyl, pfi-tacticalrmm, pfi-postgres) took the workstation key via ssh-copy-id with lkraven@. 2 LXCs (ana-filebot, ana-wg) needed key installed via pct push from pfi-pve because PermitRootLogin prohibit-password blocked ssh-copy-id. Minor: pfi-pteradactyl's server_inspect Docker section is blank because lkraven is not in the docker group there — sudo usermod -aG docker lkraven + re-login to fix.

  3. Patch out forget schedules from all 6 restic profiles — done.

3b. Discover sf-r630 OS-side IPresolved 2026-04-23: the R630 with iDRAC 10.250.250.110 is the same physical box that runs sfsrv-ana (Proxmox VE at 10.250.250.115). No separate OS IP to find. servers/sf-r630/ now clarifies it as the hardware/BMC-only inventory entry; servers/sfsrv-ana/ is the OS view. servers/ana-ml2/README.md updated with its own distinct BMC IP (10.250.250.50, Supermicro) to prevent future confusion between the two physical chassis.

  1. (removed — see note above)

🟧 Real work (dedicated session each)

4b. Migrate DB data directories off NFS onto local VM disk.done 2026-04-23/24. Postgres on pfi-postgres migrated 04-23; mongodb on esh-vm-db confirmed already local (and serves zero user data in practice — paperless uses Postgres 15 on the same VM). fstab entry + NFS mount for /mnt/db on pfi-postgres removed 2026-04-24 via playbooks/decouple-pfi-postgres-from-ana-nas.yaml. pfi-postgres has zero remaining dependency on ana-nas. Only residue: cold archive of pre-migration data still on ana-nas (/mnt/db/pfi-*); harmless, can sit indefinitely.

  1. Rotate exposed secretsdone 2026-04-23. All six rotated: vaultwarden/gitea/paperless-ng Postgres passwords (hardcoded compose.yaml literals moved to gitignored .env files in the process), ana-docker + ana-ml2 + esh-docker-vm rest-server htpasswd entries, and rest-server repo passphrases for ana-docker + esh-docker-vm (ana-ml2's was already rotated during prior wipe+reinit). Also discovered along the way: paperless uses a separate esh-vm-db VM (10.0.50.60), not pfi-postgres as the old notes implied.

  2. Cross-site rsyncdone 2026-04-23, both directions.

    • ana → nh3 (04:00 daily): ana-nas:/mnt/backup/restic/repo/ana/nh3-nas:/volume1/Backup/restic-ana-mirror/. Runs on ana-nas as lkraven. 15.6 GB initial sync completed 07:18 UTC.
    • nh3 → ana (05:00 daily): nh3-nas:/volume1/Backup/restic/ana-nas:/mnt/backup/restic-nh3-mirror/. Runs on nh3-nas as root (rest-server-nh3 container writes mode-400 files; only root can read them on Synology). 2.31 GB initial sync completed 07:42 UTC.
    • Tracked at configs/rsync/ana-nas-to-nh3/ and configs/rsync/nh3-nas-to-ana/ respectively.

6b. PBS deployment across the fleet. Runbook at docs/runbooks/pbs-deployment.md. Phases 06 done (2026-04-22): PBS-ANA on pfi-pve with NFS datastore on Ana NAS (NFSv3 for ZFS case-insensitivity), PBS-NH3 on nh3-pve with NFS datastore on NH3 Synology (NFSv3 + Linux-mode POSIX 777 to avoid syno_acl interference), per-hypervisor namespaces, API tokens, verify jobs on both, one-way sync ANA → NH3 at 06:00 daily, and all 5 hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana) onboarded to PBS-ANA. Remaining phases: - Phase 7-8 — one week burn-in, then retire legacy vzdump targets on each hypervisor (keep until 2026-04-29 earliest) - Phase 9 — homepage cards for PBS-ANA + PBS-NH3 web UIs, memory/doc updates, servers/pbs-ana + servers/pbs-nh3 dirs

  1. SureFire tenant backup plan decisionresolved 2026-04-23 by folding sfsrv-ana into the fleet-wide PBS deployment (dedicated sfsrv-ana namespace on PBS-ANA, replicates to PBS-NH3 via the same sync job as the rest of the fleet). Hosting-agreement option chosen: PFI provides backup coverage as part of managed hosting.

🟨 Prereqs / polish

  1. Synology SSH setupdone 2026-04-22. Dedicated syncuser account (admin-group membership) with key auth, registered as servers/nh3-nas/, reachable as ssh nh3-nas. Unlocks cross-site rsync (item 6), rest-server-nh3 healthcheck deploy, and .htpasswd edits on the NH3 side.

  2. scripts/restic-prune.shdone 2026-04-24. Quarterly disk-hygiene tool. Drops --append-only on the rest-server (via a temporary docker-compose.override.yaml — never edits the canonical compose), runs resticprofile forget --prune --verbose on each client, restores --append-only (with trap so it runs even on partial failure). ANA side fully automated (5 clients); NH3 side prints a manual ceremony because DSM Container Manager + sudo on syncuser aren't cleanly scriptable from this workstation. Run with scripts/restic-prune.sh ana|nh3|all, optionally --dry-run.

  3. Retire offen/docker-volume-backup sidecarsdone 2026-04-23. Removed from paperless-ngx and pgadmin composes on esh-docker-vm (only hosts in the fleet that had them). 16 GB of orphan tarballs at /mnt/backup/docker/esh-vm-docker/ reclaimed. Restic coverage verified equivalent: /var/lib/docker/volumes in source + pre-backup hook handles the DB dumps for paperless (against esh-vm-db), HA, pgadmin, CWA, uptime-kuma. Obsolete *_offen_backup_data exclude also removed from the restic profile.

  4. Clean up retired mattermost dir on ana-dockerdone (verified 2026-04-24: /opt/docker/compose/mattermost/ does not exist; no mattermost containers anywhere on the host).

🟩 Research / deferred / intermittent

  1. Backrest UI intermittent timeout. Backend is confirmed healthy (direct curl GetConfig returns in 200ms). UI chokes on the same endpoint. Likely cause: SSE stream wedge in browser, or service worker stale. Workaround: Ctrl+Shift+R. Root cause investigation deferred.

  2. UniFi controller homepage cardsdone 2026-04-24. Both cards added to configs/homepage/services.yaml and pushed to esh-docker-vm: PFI-UDMSE (10.100.0.1, UDM Pro SE) under Infra - NH3 as the new edge device replacing the retired Fortigate 101F; ESH-UDMPM (10.0.0.1, UDM Pro Max) under Infra - ESH. Icon si-ubiquiti.

  3. Prune + credential-rotation scripts as repeatable tooling (vs per-incident manual work).

🟦 Memory / documentation housekeeping

  1. docs/ organizationdone 2026-04-24. First-pass landed earlier (misfiled tea-*.sha256 removed; docs/README.md nav map added). Second pass landed same day: stripped the broken YAML frontmatter from both VM-102 Matrix docs (the path: values pointed at docs/pfi-ana/... which doesn't exist in this repo, and no toolchain consumed the metadata); deleted pfi/chromadb-setup.md (referenced configs/pfi-ana/... and scripts/setup-chromadb.sh, both nonexistent — deployment is long done and the operational truth lives in docker-stack.md). Kept the two VM-102 docs separate by design (each is right-sized; a merge would push past the 500-line guideline in the README).

  2. STATUS.md drift disciplinedone 2026-04-24. Refreshed with the post-tooling-day session work; status-regen.sh idea dropped — STATUS.md is intentionally narrative, not derivable from git/code, so auto-regen would lose information. Discipline rule: refresh STATUS.md at the end of any session where a 🟧 or 🟥 item closes, or three+ smaller items land.

Session milestones — 2026-04-24 (the "tooling day" + housekeeping pm)

Morning / early afternoon — the original tooling day:

  • irv-ml1 AI stacks deployed as Docker: ComfyUI (port 8188, host- writable workflows), Parakeet ASR (port 8765, rewritten on sherpa-onnx after the Shadowfita FastAPI wrapper hit two unfixed upstream bugs), CosyVoice 3 (port 8190, GLaDOS voice cloned + live).
  • irv-ml1 restic profile extended to cover /worktank/{comfyui/basedir/{user,custom_nodes,input},cosyvoice/voices}; bulk weights + scratch dirs stay excluded.
  • scripts/elway shipped — ~800-line Python playbook runner with tier 1 + tier 2 idempotency; handlers / register / multi-host / content-hash are deferred (gitea #3#6).
  • task-board built end-to-end (separate repo, vh/task-board on gitea) and shipped as a Claude Code plugin. Green/red cards per session via UserPromptSubmit + Stop hooks; four MCP tools expose explicit activity tracking.
  • pfi-postgres NFS decoupling finished item 4b — zero residual dependency on ana-nas for that VM.

Late afternoon / evening — task-board iteration + STATUS sweep:

  • task-board v0.1.1 → v0.1.3 shipped over four iterations on the live ANA deployment:
    • v0.1.1 — hooks parse Claude Code's stdin JSON for session_id and append a short suffix when TASK_BOARD_SESSION isn't set, so two sessions in one project no longer collide on a single card. Dormant transition preserves cumulative idle time (state_entered_at = last_update_at instead of now).
    • v0.1.2 — favicon (3-column SVG in active/waiting/dormant state colors); served at /static/favicon.svg with a <link rel="icon"> and a /favicon.ico route returning the same SVG. Followup fix for an XML-illegal -- in a comment.
    • v0.1.2-followup — UI live-duration ticker bumped from 5 s → 1 s (humanDuration floors to integer seconds; cheap render).
    • v0.1.3 — case-insensitive session names. sessions.name COLLATE NOCASE; real ALTER migration (not a wipe) — keeps earliest- created_at row as canonical, reassigns child comments. Write path canonicalizes session label before inserting comments. Read path uses COLLATE NOCASE for safety on external API callers.
  • Parakeet (irv-ml1) healthcheck fix — image ships wget not curl; healthcheck swap, 2,190 failing checks → healthy.
  • AIPA-MCP project session label fix — set TASK_BOARD_SESSION= Architect in .claude/settings.json, updated the project's CLAUDE.md to specify session="Architect" for explicit MCP calls, and renamed the existing AIPA-MCP card → Architect in the live SQLite (1 session row + 25 comment rows preserved).
  • STATUS items 11 / 13 / 15 / 16 closed. Mattermost dir gone on ana-docker (verified); UniFi UDM cards added to homepage (PFI-UDMSE for NH3 edge replacing the retired Fortigate 101F, ESH-UDMPM for ESH); docs nav map added + stale chromadb-setup.md removed + VM-102 Matrix docs frontmatter stripped; STATUS.md discipline rule recorded.

Session milestones — 2026-04-20 / 2026-04-21

  • Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
  • Per-group icons + equal-height layout + 4-column grids.
  • Fleet-wide label sweep (function groups, Service Networking renamed from Wiring/Plumbing due to homepage-parser slash bug).
  • 6 restic clients deployed end-to-end (client creds, repo init, profile install, systemd timers, first backups verified).
  • 9 hosts registered under servers/ from gap-analysis (PFI VMs + SureFire tenant, with tenancy awareness).
  • Discovery scripts (FortiGate SSH + UniFi cloud API + TSV diff vs inventory).
  • Proxmox inspect script + fleet-wide refresh wrapper.
  • Calibre-Web-Automated migration replacing calibre + calibre-web pair.
  • llama-swap Qwen3.6 stock + mradermacher abliterated, 128K ctx.
  • Mattermost retired (not running, compose dir cleanup pending).
  • FortiGate 101F at NH3 retired; homepage card removed.

Memory pointers (for future Claude sessions)

Relevant ~/.claude/.../memory/ entries:

  • server_split.md — host placement rules
  • feedback_ssh_sudo.md — use ssh -t for remote sudo (mostly superseded by elway, but still applies to ad-hoc ssh)
  • feedback_git_autonomous.md — handle git commits without asking
  • feedback_git_commits.md — no Claude attribution in commit messages
  • feedback_use_elway.md — write elway playbooks; don't chain ssh+sudo
  • feedback_backups_must_not_risk_production.md — rule adopted after the 2026-04-23 ana-nas self-backup crash
  • project_backup_pipeline_gaps.md — user's explicit goal of "all hosts + configs + DBs backed up"
  • project_db_migrate_off_nfs.md — DB-off-NFS decision + status
  • project_surefire_tenant.md — SureFire tenancy boundary awareness
  • reference_gitea_remote.md — origin is vh/esh-pfi-infrastructure on gitea.phasefinal.com; tea CLI logged in as vh
  • reference_task_board.md — task-board plugin + tools contract
  • storage_ana_nas.md — ana NAS is Debian LXC (CT 109), not TrueNAS
  • storage_nh3_nas.md — NH3 NAS via syncuser, not admin
  • incident_ana_nas_spof.md — blast-radius matrix for ana-nas outages