@@ -1,16 +1,23 @@
# Status + Open Issues
Last updated: 2026-04-21
Last updated: 2026-04-24
Snapshot of fleet state and open work. Refresh this file when a pass of
significant work lands — don't let it drift quietly.
**New session? Read [`docs/orientation.md`](docs/orientation.md) first. **
## What's in place
### Backup coverage
### Backup coverage (2-layer, fully operational)
- **VM-level vzdump:** 24/24 guests acros s pfi-pve, nh3-pve, esh-pve, esh-pve-nas.
- **File-level restic:** 6/6 hosts configured with systemd timers:
- **PBS fleet-wide.** All 5 hypervisor s ( pfi-pve, nh3-pve, esh-pve,
esh-pve-nas, sfsrv-ana) → `pbs-ana` (primary, VM on pfi-pve with
NFSv3 datastore on ana-nas). Per-hypervisor namespaces to avoid
VMID collisions. DR mirror at `pbs-nh3` (VM on nh3-pve with NFSv3
datastore on nh3-nas Synology) pulls daily at 06:00. Verify jobs
armed on both sides.
- **restic 8/8 hosts** with resticprofile + systemd timers at 01:00:
| Host | Target rest-server | DB hooks |
|---|---|---|
| ana-docker | rest-server-ana (local) | synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui |
@@ -19,17 +26,48 @@ significant work lands — don't let it drift quietly.
| esh-docker-vm | rest-server-ana (cross-site) | paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite |
| vm-esh-nas | rest-server-ana (cross-site) | — |
| nh3-dev (workstation) | rest-server-nh3 (local) | — |
| irv-ml1 | rest-server-nh3 (via WG) | — |
| esh-vm-db | rest-server-ana (cross-site) | pg_dumpall, mongodump |
- **Cross-site restic rsync.** `ana-nas → nh3-nas` at 04:00 daily
(mirrors rest-server-ana data). `nh3-nas → ana-nas` at 05:00 daily
(mirrors rest-server-nh3 data, runs as root since DSM writes files
as admin mode 400). Tracked at `configs/rsync/{ana-nas-to-nh3,nh3-nas-to-ana}/` .
### Inventory
- **18 tracked hosts** under `servers/` (5 Docker + 4 Proxmox + 1 workstation +
6 PFI VMs + 3 SureFire tenant + 1 NH3-dev workstation... audit the actual count
vs this when updating).
- **Homepage** at <http://10.0.50.45:5100> shows function-first layout with
Main / Infrastructure / Toolchain tabs. Per-group icons, row-count
layout, useEqualHeights.
- **FortiGate + UniFi discovery** scripts produce TSVs; gap analysis
against `servers/*` for unmanaged IPs.
- **~22 tracked hosts** under `servers/` across ANA / NH3 / ESH / IRV sites.
SSH config aliases in `~/.ssh/config` for every host — `ssh <name>` just works.
- **Homepage** at <http://10.0.50.45:5100> — function-first layout
(Main / Infrastructure / Toolchain tabs), per-group icons, four
Infra sub-groups (ANA / NH3 / IRV / ESH). Docker auto-discovery
active for 5 hosts (ana-docker, ana-ml2, nh3-docker, esh-docker-vm,
irv-ml1).
- **FortiGate + UniFi discovery scripts** produce TSVs; gap analysis
against `servers/*/` for unmanaged IPs.
### Architecture decisions (durable)
- **DB data on local disk, not NFS.** pfi-postgres migrated 2026-04-23;
fstab + NFS mount removed 2026-04-24 (pfi-postgres now fully
decoupled from ana-nas). Mongo on esh-vm-db confirmed already local.
Removes the biggest ana-nas blast-radius risk. Memory:
`project_db_migrate_off_nfs.md` .
- **Backups must not risk production.** Rule adopted after
2026-04-23 ana-nas self-backup crash. Memory:
`feedback_backups_must_not_risk_production.md` .
- **offen sidecars retired fleet-wide 2026-04-23.** restic covers
equivalent scope; reclaimed 16 GB of redundant tarballs on
`/mnt/backup/docker/esh-vm-docker/` .
- **Gitea remote for this repo (2026-04-23).** `origin` is
`vh/esh-pfi-infrastructure` on `gitea.phasefinal.com` . Also:
`vh/task-board` (new 2026-04-24) hosts a Claude Code plugin +
marketplace serving the assistant task-state dashboard.
- **Prefer elway for multi-step SSH work (2026-04-24).** The new
`scripts/elway` mini-playbook runner replaces chained `ssh -t sudo …`
commands — handles sudo once up front, structured reporting,
idempotency (creates/when/changed_when). Memory:
`feedback_use_elway.md` ; template playbook:
`playbooks/elway-smoke.yaml` .
### Tooling
@@ -40,29 +78,39 @@ significant work lands — don't let it drift quietly.
`scripts/discover-gaps.sh` for network-level inventory discovery.
- `scripts/deploy-stack.sh` + `scripts/sync-stacks.sh` for compose push/pull.
- `scripts/add-host.sh` for new-host registration.
- **`scripts/elway` ** (2026-04-24) — mini-ansible playbook runner over
SSH. Playbooks live under `playbooks/` . Tier 1 (creates/when) + tier
2 (changed_when) idempotency; handlers + register + multi-host
fan-out deferred as gitea issues #3 – #5 .
- **`tea` CLI** at `/usr/local/bin/tea` — gitea CLI, already logged in
as `vh` . Use for issue / PR work instead of inventing URLs.
## Open issues
### 🟥 Quick wins (do next)
1. **Fix Backrest's `esh-docker-vm` URI ** — Backrest is pointed at
`10.100.50.50` (NH3 Synology) for that repo, while the actual backup
writes go to `10.250.50.70` (rest-server-ana). Edit `config.json` in
the backrest container, swap the host portion. ~5 min.
1. ~~ **Fix Backrest's `esh-docker-vm` URI**~~ — done.
2. **Verify SSH on the 9 newly-registered hos ts** and pull first snaps hots:
```
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh pfi-ana-webhost ana-filebot pfi-pteradactyl \
pfi-tacticalrmm pfi-postgres ana-wg
` ``
ssh-target files were guessed (` lkraven@` for VMs, ` root@` for LXCs) —
first run surfaces mismatches.
2. ~~**Push SSH keys + pull snaps hots for 6 unrefreshed hos ts**~~ —
**done 2026-04-23 ** . 4 VMs (pfi-ana-webhost, pfi-pteradactyl,
pfi-tacticalrmm, pfi-postgres) took the workstation key via
`ssh-copy-id` with `lkraven@` . 2 LXCs (ana-filebot, ana-wg)
needed key installed via `pct push` from pfi-pve because
`PermitRootLogin prohibit-password` blocked `ssh-copy-id` .
Minor: pfi-pteradactyl's `server_inspect` Docker section is blank
because lkraven is not in the `docker` group there — `sudo usermod
-aG docker lkraven` + re-login to fix.
3. **Patch out ` forget` schedules from all 6 restic profiles. ** They fail
nightly against ` --append-only` rest-servers. One-liner per host:
comment out ` schedule:` under the ` forget` block in profiles.yaml;
redeploy; run ` resticprofile unschedule forget` on each host. ~15 min.
3. ~~ **Patch out `forget` schedules from all 6 restic profiles**~~ — done.
3b. ~~**Discover sf-r630 OS-side IP**~~ — **resolved 2026-04-23 ** :
the R630 with iDRAC `10.250.250.110` is the same physical box
that runs `sfsrv-ana` (Proxmox VE at `10.250.250.115` ). No
separate OS IP to find. `servers/sf-r630/` now clarifies it as
the hardware/BMC-only inventory entry; `servers/sfsrv-ana/` is
the OS view. `servers/ana-ml2/README.md` updated with its own
distinct BMC IP (`10.250.250.50` , Supermicro) to prevent future
confusion between the two physical chassis.
<!-- 2026-04-21: Dropped the "nightly-restart safety net" idea — user
confirmed Backrest UI chokes even right after startup, so periodic
@@ -73,63 +121,81 @@ significant work lands — don't let it drift quietly.
### 🟧 Real work (dedicated session each)
5. **Rotate exposed secrets** (captured in session transcripts 2026-04-20/21):
- ` vaultwarden ` Postgres password ( on pfi-postgres)
- ` gitea` Postgres password — currently literally ` gitea` (trivially weak)
- ` paperless-ng` Postgres password — currently literally ` paperless-ng` (trivially weak)
- ` ana-docker` rest-server htpasswd + repo passphrase
- ` ana-ml2` rest-server htpasswd (repo passphrase rotated during wipe/reinit)
- ` esh-docker-vm` rest-server htpasswd + repo passphrase
4b. ~~**Migrate DB data directories off NFS onto local VM disk.**~~ —
**done 2026-04-23/24. ** Postgres on pfi-postgres migrated
04-23; mongodb on esh-vm-db confirmed already local (and serves
zero user data in practice — paperless uses Postgres 15 on the
same VM). fstab entry + NFS mount for `/mnt/db` on pfi-postgres
removed 2026-04-24 via
`playbooks/decouple-pfi-postgres-from-ana-nas.yaml` . pfi-postgres
has zero remaining dependency on ana-nas. Only residue: cold
archive of pre-migration data still on ana-nas (`/mnt/db/pfi-*` );
harmless, can sit indefinitely.
For each: rotate at the source (DB ` ALTER USER …` / ` htpasswd` /
` restic key add` + ` restic key remove`) → update consumer config
files (` /etc/restic/restic.env`, ` /etc/restic/dbcreds.env`, compose
env files, Backrest config.json) → restart consumer services. ~45 min
batched.
5. ~~**Rotate exposed secrets**~~ — **done 2026-04-23 ** . All six
rotated: vaultwarden/gitea/paperless-ng Postgres passwords
(hardcoded `compose.yaml` literals moved to gitignored `.env`
files in the process), ana-docker + ana-ml2 + esh-docker-vm
rest-server htpasswd entries, and rest-server repo passphrases for
ana-docker + esh-docker-vm (ana-ml2's was already rotated during
prior wipe+reinit). Also discovered along the way: paperless uses
a separate `esh-vm-db` VM (10.0.50.60), not pfi-postgres as the
old notes implied.
6. **Cross-site rsync** between ` rest-server-ana ` data dir
(` /mnt/backup/restic/repo/ana/`) and NH3 Synology data dir. Planned
since initial rest-server setup; not built. Needs Synology SSH
access first (item 8). ~20 min once access is there .
- Note: cross-site mirroring for the VM-image layer is being
addressed by the PBS deployment (item 6b) — this rsync is now
scoped to restic repos only.
6. ~~ **Cross-site rsync**~~ — **done 2026-04-23, both directions ** .
- **ana → nh3** (04:00 daily): `ana-nas:/mnt/backup/restic/repo/ana/`
→ `nh3-nas:/volume1/Backup/restic-ana-mirror/` . Runs on ana-nas
as lkraven. 15.6 GB initial sync completed 07:18 UTC .
- **nh3 → ana** (05:00 daily): `nh3-nas:/volume1/Backup/restic/`
→ `ana-nas:/mnt/backup/restic-nh3-mirror/` . Runs on nh3-nas as
root (rest-server-nh3 container writes mode-400 files; only
root can read them on Synology). 2.31 GB initial sync completed
07:42 UTC.
- Tracked at `configs/rsync/ana-nas-to-nh3/` and
`configs/rsync/nh3-nas-to-ana/` respectively.
6b. **PBS deployment across the fleet.** Runbook lives at
` docs/runbooks/pbs-deployment.md`. Primary at ANA (VM on pfi-pve,
NFS datastore on Debian NAS), DR mirror at NH3 (VM on nh3-pve,
local datastore). One-way sync from ANA → NH3 nightly. All 5
6b. **PBS deployment across the fleet. ** Runbook at
`docs/runbooks/pbs-deployment.md` . Phases 0– 6 **done ** (2026-04-22):
PBS-ANA on pfi-pve with NFS datastore on Ana NAS (NFSv3 for ZFS
case-insensitivity), PBS-NH3 on nh3-pve with NFS datastore on NH3
Synology (NFSv3 + Linux-mode POSIX 777 to avoid syno_acl
interference), per-hypervisor namespaces, API tokens, verify jobs
on both, one-way sync ANA → NH3 at 06:00 daily, and all 5
hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana)
migrate off local-dump vzdump jobs on to PBS.
- Closes: SureFire's zero-coverage gap (currently no vzdump at all on sfsrv-ana)
- Closes: "cross-site redundancy for VM images" scope item
- Provides: cross-fleet block-level dedupe, dirty-bitmap incrementals,
real safe-prune retention, per-host encryption keys
- Est. 4– 6h spread across sittings. 9 phases in the runbook.
onboarded to PBS-ANA .
Remaining phases:
- Phase 7-8 — one week burn-in, then retire legacy vzdump targets
on each hypervisor (keep until 2026-04-29 earliest)
- Phase 9 — homepage cards for PBS-ANA + PBS-NH3 web UIs,
memory/doc updates, `servers/pbs-ana` + `servers/pbs-nh3` dirs
7. **SureFire tenant backup plan decision. ** Three options documented in
` servers/sfsrv-ana/README.md`:
- Tenant handles own backups
- PFI provides dedicated scoped repo on rest-server-ana
- Shared vzdump target
Blocks any actual SF backup work until the hosting-agreement side of
this is clear.
7. ~~ **SureFire tenant backup plan decision**~~ — **resolved 2026-04-23 **
by folding sfsrv-ana into the fleet-wide PBS deployment (dedicated
`sfsrv-ana` namespace on PBS-ANA, replicates to PBS-NH3 via the
same sync job as the rest of the fleet). Hosting-agreement option
chosen: PFI provides backup coverage as part of managed hosting.
### 🟨 Prereqs / polish
8. **Synology SSH setup** — tabled earlier. Unlocks: cross-site rsync
(item 6), rest-server-nh3 healthcheck deploy, easier ` .htpasswd`
edits on the NH3 side .
8. ~~ **Synology SSH setup**~~ — **done 2026-04-22 ** . Dedicated
`syncuser` account (admin-group membership) with key auth,
registered as `servers/nh3-nas/` , reachable as `ssh nh3-nas` .
Unlocks cross-site rsync (item 6), rest-server-nh3 healthcheck
deploy, and `.htpasswd` edits on the NH3 side.
9. * * `scripts/restic-prune.sh` ** — temporarily flip `--append-only` off,
run forget + prune across all hosts, flip back on. Needed quarterly
for disk hygiene. Not urgent; blocks only the "I need to reclaim
disk space now" scenario.
10. **Retire ` offen/docker-volume-backup` sidecars on esh-docker-vm ** onc e
restic has ~1 week of clean runs. Paperless + pgadmin currently ru n
offen sidecars that write tarballs to ` /mnt/backup/...` redundantly.
10. ~~ **Retire `offen/docker-volume-backup` sidecars**~~ — **d one
2026-04-23**. Removed from paperless-ngx and pgadmin composes o n
esh-docker-vm (only hosts in the fleet that had them). 16 GB of
orphan tarballs at `/mnt/backup/docker/esh-vm-docker/` reclaimed.
Restic coverage verified equivalent: `/var/lib/docker/volumes` in
source + pre-backup hook handles the DB dumps for paperless
(against esh-vm-db), HA, pgadmin, CWA, uptime-kuma. Obsolete
`*_offen_backup_data` exclude also removed from the restic profile.
11. **Clean up retired mattermost dir ** on ana-docker — compose dir at
`/opt/docker/compose/mattermost/` may still linger. Single `rm -rf`
@@ -159,6 +225,25 @@ significant work lands — don't let it drift quietly.
lands. Consider a `scripts/status-regen.sh` if manual updates
slip.
## Session milestones — 2026-04-24 (the "tooling day")
- **irv-ml1 AI stacks deployed** as Docker: ComfyUI (port 8188, host-
writable workflows), Parakeet ASR (port 8765, rewritten on
sherpa-onnx after the Shadowfita FastAPI wrapper hit two unfixed
upstream bugs), CosyVoice 3 (port 8190, GLaDOS voice cloned + live).
- **irv-ml1 restic profile extended** to cover
`/worktank/{comfyui/basedir/{user,custom_nodes,input},cosyvoice/voices}` ;
bulk weights + scratch dirs stay excluded.
- **`scripts/elway` ** shipped — ~800-line Python playbook runner with
tier 1 + tier 2 idempotency; handlers / register / multi-host /
content-hash are deferred (gitea #3 – #6 ).
- **task-board** built end-to-end (separate repo, `vh/task-board` on
gitea) and shipped as a Claude Code plugin. Green/red cards per
session via UserPromptSubmit + Stop hooks; four MCP tools expose
explicit activity tracking.
- **pfi-postgres NFS decoupling** finished item 4b — zero residual
dependency on ana-nas for that VM.
## Session milestones — 2026-04-20 / 2026-04-21
- Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
@@ -181,10 +266,20 @@ significant work lands — don't let it drift quietly.
Relevant `~/.claude/.../memory/` entries:
- `server_split.md` — host placement rules
- ` feedback_ssh_sudo.md` — use ` ssh -t` for remote sudo
- ` feedback_ssh_sudo.md` — use `ssh -t` for remote sudo (mostly
superseded by elway, but still applies to ad-hoc ssh)
- `feedback_git_autonomous.md` — handle git commits without asking
- `feedback_git_commits.md` — no Claude attribution in commit messages
- ` project_backup_pipeline_gaps .md` — user's explicit goal of "all hosts +
configs + DBs backed up"
- `feedback_use_elway .md` — write elway playbooks; don't chain ssh+sudo
- `feedback_backups_must_not_risk_production.md` — rule adopted after
the 2026-04-23 ana-nas self-backup crash
- `project_backup_pipeline_gaps.md` — user's explicit goal of "all
hosts + configs + DBs backed up"
- `project_db_migrate_off_nfs.md` — DB-off-NFS decision + status
- `project_surefire_tenant.md` — SureFire tenancy boundary awareness
- ` storage_ana_nas .md` — ana NAS is Debian, not TrueNAS
- `reference_gitea_remote .md` — origin is `vh/esh-pfi-infrastructure`
on gitea.phasefinal.com; `tea` CLI logged in as `vh`
- `reference_task_board.md` — task-board plugin + tools contract
- `storage_ana_nas.md` — ana NAS is Debian LXC (CT 109), not TrueNAS
- `storage_nh3_nas.md` — NH3 NAS via `syncuser` , not `admin`
- `incident_ana_nas_spof.md` — blast-radius matrix for ana-nas outages