Files
esh-pfi-infrastructure/STATUS.md
T
vh 76a0768fdb restic: drop scheduled forget across all 6 hosts
Forget against an --append-only rest-server fails every night (delete
ops blocked). The resulting daily failure cluttered service status and
logs without ever actually retiring old snapshots. Schedule is now
removed from the forget block in all six profiles; the keep-daily /
keep-weekly / keep-monthly / keep-yearly policy remains so manual
invocations (during prune ceremonies, when --append-only is
temporarily off) honor the intended retention.

Files:
  configs/restic/ana-docker/profiles.yaml
  configs/restic/ana-ml2/profiles.yaml
  configs/restic/nh3-docker/profiles.yaml
  configs/restic/esh-docker-vm/profiles.yaml
  configs/restic/vm-esh-nas/profiles.yaml
  configs/restic/nh3-dev/profiles.yaml

Each file has an inline comment marking why the schedule was dropped
so a future reader doesn't re-add it thinking it was an oversight.

STATUS.md: removed the "install Backrest nightly-restart timer" line
item. User confirmed the UI timeout hits even at startup, so periodic
restart wouldn't actually help. Root cause remains deferred.
2026-04-21 17:00:23 -07:00

7.7 KiB

Status + Open Issues

Last updated: 2026-04-21

Snapshot of fleet state and open work. Refresh this file when a pass of significant work lands — don't let it drift quietly.

What's in place

Backup coverage

  • VM-level vzdump: 24/24 guests across pfi-pve, nh3-pve, esh-pve, esh-pve-nas.
  • File-level restic: 6/6 hosts configured with systemd timers:
    Host Target rest-server DB hooks
    ana-docker rest-server-ana (local) synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui
    ana-ml2 rest-server-ana (cross-site) — (no DBs)
    nh3-docker rest-server-nh3 (local, Synology)
    esh-docker-vm rest-server-ana (cross-site) paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite
    vm-esh-nas rest-server-ana (cross-site)
    nh3-dev (workstation) rest-server-nh3 (local)

Inventory

  • 18 tracked hosts under servers/ (5 Docker + 4 Proxmox + 1 workstation + 6 PFI VMs + 3 SureFire tenant + 1 NH3-dev workstation... audit the actual count vs this when updating).
  • Homepage at http://10.0.50.45:5100 shows function-first layout with Main / Infrastructure / Toolchain tabs. Per-group icons, row-count layout, useEqualHeights.
  • FortiGate + UniFi discovery scripts produce TSVs; gap analysis against servers/* for unmanaged IPs.

Tooling

  • scripts/server_inspect.sh + refresh-server-info.sh for Docker hosts.
  • scripts/proxmox_inspect.sh + refresh-proxmox-info.sh for PVE nodes (reports VMs/LXCs/storage/backup-coverage).
  • scripts/discover-fortigate.sh, scripts/discover-unifi.sh, scripts/discover-gaps.sh for network-level inventory discovery.
  • scripts/deploy-stack.sh + scripts/sync-stacks.sh for compose push/pull.
  • scripts/add-host.sh for new-host registration.

Open issues

🟥 Quick wins (do next)

  1. Fix Backrest's esh-docker-vm URI — Backrest is pointed at 10.100.50.50 (NH3 Synology) for that repo, while the actual backup writes go to 10.250.50.70 (rest-server-ana). Edit config.json in the backrest container, swap the host portion. ~5 min.

  2. Verify SSH on the 9 newly-registered hosts and pull first snapshots:

    scripts/refresh-server-info.sh --validate-only all
    scripts/refresh-server-info.sh pfi-ana-webhost ana-filebot pfi-pteradactyl \
      pfi-tacticalrmm pfi-postgres ana-wg
    

    ssh-target files were guessed (lkraven@ for VMs, root@ for LXCs) — first run surfaces mismatches.

  3. Patch out forget schedules from all 6 restic profiles. They fail nightly against --append-only rest-servers. One-liner per host: comment out schedule: under the forget block in profiles.yaml; redeploy; run resticprofile unschedule forget on each host. ~15 min.

  1. (removed — see note above)

🟧 Real work (dedicated session each)

  1. Rotate exposed secrets (captured in session transcripts 2026-04-20/21):

    • vaultwarden Postgres password (on pfi-postgres)
    • gitea Postgres password — currently literally gitea (trivially weak)
    • paperless-ng Postgres password — currently literally paperless-ng (trivially weak)
    • ana-docker rest-server htpasswd + repo passphrase
    • ana-ml2 rest-server htpasswd (repo passphrase rotated during wipe/reinit)
    • esh-docker-vm rest-server htpasswd + repo passphrase

    For each: rotate at the source (DB ALTER USER … / htpasswd / restic key add + restic key remove) → update consumer config files (/etc/restic/restic.env, /etc/restic/dbcreds.env, compose env files, Backrest config.json) → restart consumer services. ~45 min batched.

  2. Cross-site rsync between rest-server-ana data dir (/mnt/backup/restic/repo/ana/) and NH3 Synology data dir. Planned since initial rest-server setup; not built. Needs Synology SSH access first (item 8). ~20 min once access is there.

  3. SureFire tenant backup plan decision. Three options documented in servers/sfsrv-ana/README.md:

    • Tenant handles own backups
    • PFI provides dedicated scoped repo on rest-server-ana
    • Shared vzdump target

    Blocks any actual SF backup work until the hosting-agreement side of this is clear.

🟨 Prereqs / polish

  1. Synology SSH setup — tabled earlier. Unlocks: cross-site rsync (item 6), rest-server-nh3 healthcheck deploy, easier .htpasswd edits on the NH3 side.

  2. scripts/restic-prune.sh — temporarily flip --append-only off, run forget + prune across all hosts, flip back on. Needed quarterly for disk hygiene. Not urgent; blocks only the "I need to reclaim disk space now" scenario.

  3. Retire offen/docker-volume-backup sidecars on esh-docker-vm once restic has ~1 week of clean runs. Paperless + pgadmin currently run offen sidecars that write tarballs to /mnt/backup/... redundantly.

  4. Clean up retired mattermost dir on ana-docker — compose dir at /opt/docker/compose/mattermost/ may still linger. Single rm -rf when convenient.

🟩 Research / deferred / intermittent

  1. Backrest UI intermittent timeout. Backend is confirmed healthy (direct curl GetConfig returns in 200ms). UI chokes on the same endpoint. Likely cause: SSE stream wedge in browser, or service worker stale. Workaround: Ctrl+Shift+R. Root cause investigation deferred.

  2. UniFi controller homepage cards (ESH-UDMPM at 10.0.0.1, PFI-UDMSE at 10.100.0.1). Auto-discovered via the Site Manager API; not yet linked as homepage cards.

  3. Prune + credential-rotation scripts as repeatable tooling (vs per-incident manual work).

🟦 Memory / documentation housekeeping

  1. docs/ organization could use a pass — multiple READMEs and reference files in different spots. Not urgent.

  2. STATUS.md (this file) drift. Update whenever significant work lands. Consider a scripts/status-regen.sh if manual updates slip.

Session milestones — 2026-04-20 / 2026-04-21

  • Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
  • Per-group icons + equal-height layout + 4-column grids.
  • Fleet-wide label sweep (function groups, Service Networking renamed from Wiring/Plumbing due to homepage-parser slash bug).
  • 6 restic clients deployed end-to-end (client creds, repo init, profile install, systemd timers, first backups verified).
  • 9 hosts registered under servers/ from gap-analysis (PFI VMs + SureFire tenant, with tenancy awareness).
  • Discovery scripts (FortiGate SSH + UniFi cloud API + TSV diff vs inventory).
  • Proxmox inspect script + fleet-wide refresh wrapper.
  • Calibre-Web-Automated migration replacing calibre + calibre-web pair.
  • llama-swap Qwen3.6 stock + mradermacher abliterated, 128K ctx.
  • Mattermost retired (not running, compose dir cleanup pending).
  • FortiGate 101F at NH3 retired; homepage card removed.

Memory pointers (for future Claude sessions)

Relevant ~/.claude/.../memory/ entries:

  • server_split.md — host placement rules
  • feedback_ssh_sudo.md — use ssh -t for remote sudo
  • feedback_git_autonomous.md — handle git commits without asking
  • feedback_git_commits.md — no Claude attribution in commit messages
  • project_backup_pipeline_gaps.md — user's explicit goal of "all hosts + configs + DBs backed up"
  • project_surefire_tenant.md — SureFire tenancy boundary awareness
  • storage_ana_nas.md — ana NAS is Debian, not TrueNAS