Files
esh-pfi-infrastructure/STATUS.md
T
vh 76a0768fdb restic: drop scheduled forget across all 6 hosts
Forget against an --append-only rest-server fails every night (delete
ops blocked). The resulting daily failure cluttered service status and
logs without ever actually retiring old snapshots. Schedule is now
removed from the forget block in all six profiles; the keep-daily /
keep-weekly / keep-monthly / keep-yearly policy remains so manual
invocations (during prune ceremonies, when --append-only is
temporarily off) honor the intended retention.

Files:
  configs/restic/ana-docker/profiles.yaml
  configs/restic/ana-ml2/profiles.yaml
  configs/restic/nh3-docker/profiles.yaml
  configs/restic/esh-docker-vm/profiles.yaml
  configs/restic/vm-esh-nas/profiles.yaml
  configs/restic/nh3-dev/profiles.yaml

Each file has an inline comment marking why the schedule was dropped
so a future reader doesn't re-add it thinking it was an oversight.

STATUS.md: removed the "install Backrest nightly-restart timer" line
item. User confirmed the UI timeout hits even at startup, so periodic
restart wouldn't actually help. Root cause remains deferred.
2026-04-21 17:00:23 -07:00

176 lines
7.7 KiB
Markdown

# Status + Open Issues
Last updated: 2026-04-21
Snapshot of fleet state and open work. Refresh this file when a pass of
significant work lands — don't let it drift quietly.
## What's in place
### Backup coverage
- **VM-level vzdump:** 24/24 guests across pfi-pve, nh3-pve, esh-pve, esh-pve-nas.
- **File-level restic:** 6/6 hosts configured with systemd timers:
| Host | Target rest-server | DB hooks |
|---|---|---|
| ana-docker | rest-server-ana (local) | synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui |
| ana-ml2 | rest-server-ana (cross-site) | — (no DBs) |
| nh3-docker | rest-server-nh3 (local, Synology) | — |
| esh-docker-vm | rest-server-ana (cross-site) | paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite |
| vm-esh-nas | rest-server-ana (cross-site) | — |
| nh3-dev (workstation) | rest-server-nh3 (local) | — |
### Inventory
- **18 tracked hosts** under `servers/` (5 Docker + 4 Proxmox + 1 workstation +
6 PFI VMs + 3 SureFire tenant + 1 NH3-dev workstation... audit the actual count
vs this when updating).
- **Homepage** at <http://10.0.50.45:5100> shows function-first layout with
Main / Infrastructure / Toolchain tabs. Per-group icons, row-count
layout, useEqualHeights.
- **FortiGate + UniFi discovery** scripts produce TSVs; gap analysis
against `servers/*` for unmanaged IPs.
### Tooling
- `scripts/server_inspect.sh` + `refresh-server-info.sh` for Docker hosts.
- `scripts/proxmox_inspect.sh` + `refresh-proxmox-info.sh` for PVE nodes
(reports VMs/LXCs/storage/backup-coverage).
- `scripts/discover-fortigate.sh`, `scripts/discover-unifi.sh`,
`scripts/discover-gaps.sh` for network-level inventory discovery.
- `scripts/deploy-stack.sh` + `scripts/sync-stacks.sh` for compose push/pull.
- `scripts/add-host.sh` for new-host registration.
## Open issues
### 🟥 Quick wins (do next)
1. **Fix Backrest's `esh-docker-vm` URI** — Backrest is pointed at
`10.100.50.50` (NH3 Synology) for that repo, while the actual backup
writes go to `10.250.50.70` (rest-server-ana). Edit `config.json` in
the backrest container, swap the host portion. ~5 min.
2. **Verify SSH on the 9 newly-registered hosts** and pull first snapshots:
```
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh pfi-ana-webhost ana-filebot pfi-pteradactyl \
pfi-tacticalrmm pfi-postgres ana-wg
```
ssh-target files were guessed (`lkraven@` for VMs, `root@` for LXCs) —
first run surfaces mismatches.
3. **Patch out `forget` schedules from all 6 restic profiles.** They fail
nightly against `--append-only` rest-servers. One-liner per host:
comment out `schedule:` under the `forget` block in profiles.yaml;
redeploy; run `resticprofile unschedule forget` on each host. ~15 min.
<!-- 2026-04-21: Dropped the "nightly-restart safety net" idea — user
confirmed Backrest UI chokes even right after startup, so periodic
restart wouldn't reliably help. Root cause still unknown; see
item under "Research / deferred" below. -->
4. _(removed — see note above)_
### 🟧 Real work (dedicated session each)
5. **Rotate exposed secrets** (captured in session transcripts 2026-04-20/21):
- `vaultwarden` Postgres password (on pfi-postgres)
- `gitea` Postgres password — currently literally `gitea` (trivially weak)
- `paperless-ng` Postgres password — currently literally `paperless-ng` (trivially weak)
- `ana-docker` rest-server htpasswd + repo passphrase
- `ana-ml2` rest-server htpasswd (repo passphrase rotated during wipe/reinit)
- `esh-docker-vm` rest-server htpasswd + repo passphrase
For each: rotate at the source (DB `ALTER USER …` / `htpasswd` /
`restic key add` + `restic key remove`) → update consumer config
files (`/etc/restic/restic.env`, `/etc/restic/dbcreds.env`, compose
env files, Backrest config.json) → restart consumer services. ~45 min
batched.
6. **Cross-site rsync** between `rest-server-ana` data dir
(`/mnt/backup/restic/repo/ana/`) and NH3 Synology data dir. Planned
since initial rest-server setup; not built. Needs Synology SSH
access first (item 8). ~20 min once access is there.
7. **SureFire tenant backup plan decision.** Three options documented in
`servers/sfsrv-ana/README.md`:
- Tenant handles own backups
- PFI provides dedicated scoped repo on rest-server-ana
- Shared vzdump target
Blocks any actual SF backup work until the hosting-agreement side of
this is clear.
### 🟨 Prereqs / polish
8. **Synology SSH setup** — tabled earlier. Unlocks: cross-site rsync
(item 6), rest-server-nh3 healthcheck deploy, easier `.htpasswd`
edits on the NH3 side.
9. **`scripts/restic-prune.sh`** — temporarily flip `--append-only` off,
run forget + prune across all hosts, flip back on. Needed quarterly
for disk hygiene. Not urgent; blocks only the "I need to reclaim
disk space now" scenario.
10. **Retire `offen/docker-volume-backup` sidecars on esh-docker-vm** once
restic has ~1 week of clean runs. Paperless + pgadmin currently run
offen sidecars that write tarballs to `/mnt/backup/...` redundantly.
11. **Clean up retired mattermost dir** on ana-docker — compose dir at
`/opt/docker/compose/mattermost/` may still linger. Single `rm -rf`
when convenient.
### 🟩 Research / deferred / intermittent
12. **Backrest UI intermittent timeout.** Backend is confirmed healthy
(direct `curl GetConfig` returns in 200ms). UI chokes on the same
endpoint. Likely cause: SSE stream wedge in browser, or service
worker stale. Workaround: Ctrl+Shift+R. Root cause investigation
deferred.
13. **UniFi controller homepage cards** (ESH-UDMPM at 10.0.0.1, PFI-UDMSE
at 10.100.0.1). Auto-discovered via the Site Manager API; not yet
linked as homepage cards.
14. **Prune + credential-rotation scripts** as repeatable tooling (vs
per-incident manual work).
### 🟦 Memory / documentation housekeeping
15. **`docs/` organization** could use a pass — multiple READMEs and
reference files in different spots. Not urgent.
16. **`STATUS.md` (this file) drift.** Update whenever significant work
lands. Consider a `scripts/status-regen.sh` if manual updates
slip.
## Session milestones — 2026-04-20 / 2026-04-21
- Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
- Per-group icons + equal-height layout + 4-column grids.
- Fleet-wide label sweep (function groups, Service Networking renamed from
Wiring/Plumbing due to homepage-parser slash bug).
- 6 restic clients deployed end-to-end (client creds, repo init, profile
install, systemd timers, first backups verified).
- 9 hosts registered under `servers/` from gap-analysis (PFI VMs + SureFire
tenant, with tenancy awareness).
- Discovery scripts (FortiGate SSH + UniFi cloud API + TSV diff vs inventory).
- Proxmox inspect script + fleet-wide refresh wrapper.
- Calibre-Web-Automated migration replacing calibre + calibre-web pair.
- llama-swap Qwen3.6 stock + mradermacher abliterated, 128K ctx.
- Mattermost retired (not running, compose dir cleanup pending).
- FortiGate 101F at NH3 retired; homepage card removed.
## Memory pointers (for future Claude sessions)
Relevant `~/.claude/.../memory/` entries:
- `server_split.md` — host placement rules
- `feedback_ssh_sudo.md` — use `ssh -t` for remote sudo
- `feedback_git_autonomous.md` — handle git commits without asking
- `feedback_git_commits.md` — no Claude attribution in commit messages
- `project_backup_pipeline_gaps.md` — user's explicit goal of "all hosts +
configs + DBs backed up"
- `project_surefire_tenant.md` — SureFire tenancy boundary awareness
- `storage_ana_nas.md` — ana NAS is Debian, not TrueNAS