paperwork: refresh STATUS.md + CLAUDE.md; finalize 4b

STATUS.md:
  - Mark 4b done (both Postgres migration + NFS decoupling)
  - Add arch decisions for gitea remote + prefer-elway policy
  - Add tooling entries for elway + tea CLI
  - Document 2026-04-24 session milestones (irv-ml1 AI stacks,
    elway, task-board, 4b finish)
  - Expand memory-pointer list with the files added this session

CLAUDE.md:
  - Tell new sessions to use elway for SSH-driven work, point at
    the smoke playbook template
  - Document the task-board plugin + MCP-tool contract so assistant
    sessions with the plugin enabled know the assistant should call
    task_start / task_update / task_wait / task_complete at
    meaningful checkpoints

.claude/settings.json:
  - Project-level env: TASK_BOARD_SESSION=Infra so every Claude Code
    session opened here labels its task-board cards "Infra"

playbooks/decouple-pfi-postgres-from-ana-nas.yaml:
  - Finishes the DB-off-NFS migration on pfi-postgres. Already ran
    against prod today; fstab clean, unmounted, no systemd mnt-db
    unit. Verify 3 was mis-expressed on first run (`grep -q active`
    matched "inactive") — fixed to invert systemctl exit code
    directly.
This commit is contained in:
vh
2026-04-24 16:31:41 -07:00
parent 42fbc4b4c4
commit f7a8b668d3
4 changed files with 297 additions and 77 deletions
+170 -75
View File
@@ -1,16 +1,23 @@
# Status + Open Issues
Last updated: 2026-04-21
Last updated: 2026-04-24
Snapshot of fleet state and open work. Refresh this file when a pass of
significant work lands — don't let it drift quietly.
**New session? Read [`docs/orientation.md`](docs/orientation.md) first.**
## What's in place
### Backup coverage
### Backup coverage (2-layer, fully operational)
- **VM-level vzdump:** 24/24 guests across pfi-pve, nh3-pve, esh-pve, esh-pve-nas.
- **File-level restic:** 6/6 hosts configured with systemd timers:
- **PBS fleet-wide.** All 5 hypervisors (pfi-pve, nh3-pve, esh-pve,
esh-pve-nas, sfsrv-ana) → `pbs-ana` (primary, VM on pfi-pve with
NFSv3 datastore on ana-nas). Per-hypervisor namespaces to avoid
VMID collisions. DR mirror at `pbs-nh3` (VM on nh3-pve with NFSv3
datastore on nh3-nas Synology) pulls daily at 06:00. Verify jobs
armed on both sides.
- **restic 8/8 hosts** with resticprofile + systemd timers at 01:00:
| Host | Target rest-server | DB hooks |
|---|---|---|
| ana-docker | rest-server-ana (local) | synapse, seafile, vaultwarden-pg, gitea (native dump), openwebui |
@@ -19,17 +26,48 @@ significant work lands — don't let it drift quietly.
| esh-docker-vm | rest-server-ana (cross-site) | paperless-pg, HA SQLite, CWA SQLite, pgadmin SQLite, uptime-kuma SQLite |
| vm-esh-nas | rest-server-ana (cross-site) | — |
| nh3-dev (workstation) | rest-server-nh3 (local) | — |
| irv-ml1 | rest-server-nh3 (via WG) | — |
| esh-vm-db | rest-server-ana (cross-site) | pg_dumpall, mongodump |
- **Cross-site restic rsync.** `ana-nas → nh3-nas` at 04:00 daily
(mirrors rest-server-ana data). `nh3-nas → ana-nas` at 05:00 daily
(mirrors rest-server-nh3 data, runs as root since DSM writes files
as admin mode 400). Tracked at `configs/rsync/{ana-nas-to-nh3,nh3-nas-to-ana}/`.
### Inventory
- **18 tracked hosts** under `servers/` (5 Docker + 4 Proxmox + 1 workstation +
6 PFI VMs + 3 SureFire tenant + 1 NH3-dev workstation... audit the actual count
vs this when updating).
- **Homepage** at <http://10.0.50.45:5100> shows function-first layout with
Main / Infrastructure / Toolchain tabs. Per-group icons, row-count
layout, useEqualHeights.
- **FortiGate + UniFi discovery** scripts produce TSVs; gap analysis
against `servers/*` for unmanaged IPs.
- **~22 tracked hosts** under `servers/` across ANA / NH3 / ESH / IRV sites.
SSH config aliases in `~/.ssh/config` for every host — `ssh <name>` just works.
- **Homepage** at <http://10.0.50.45:5100> — function-first layout
(Main / Infrastructure / Toolchain tabs), per-group icons, four
Infra sub-groups (ANA / NH3 / IRV / ESH). Docker auto-discovery
active for 5 hosts (ana-docker, ana-ml2, nh3-docker, esh-docker-vm,
irv-ml1).
- **FortiGate + UniFi discovery scripts** produce TSVs; gap analysis
against `servers/*/` for unmanaged IPs.
### Architecture decisions (durable)
- **DB data on local disk, not NFS.** pfi-postgres migrated 2026-04-23;
fstab + NFS mount removed 2026-04-24 (pfi-postgres now fully
decoupled from ana-nas). Mongo on esh-vm-db confirmed already local.
Removes the biggest ana-nas blast-radius risk. Memory:
`project_db_migrate_off_nfs.md`.
- **Backups must not risk production.** Rule adopted after
2026-04-23 ana-nas self-backup crash. Memory:
`feedback_backups_must_not_risk_production.md`.
- **offen sidecars retired fleet-wide 2026-04-23.** restic covers
equivalent scope; reclaimed 16 GB of redundant tarballs on
`/mnt/backup/docker/esh-vm-docker/`.
- **Gitea remote for this repo (2026-04-23).** `origin` is
`vh/esh-pfi-infrastructure` on `gitea.phasefinal.com`. Also:
`vh/task-board` (new 2026-04-24) hosts a Claude Code plugin +
marketplace serving the assistant task-state dashboard.
- **Prefer elway for multi-step SSH work (2026-04-24).** The new
`scripts/elway` mini-playbook runner replaces chained `ssh -t sudo …`
commands — handles sudo once up front, structured reporting,
idempotency (creates/when/changed_when). Memory:
`feedback_use_elway.md`; template playbook:
`playbooks/elway-smoke.yaml`.
### Tooling
@@ -40,29 +78,39 @@ significant work lands — don't let it drift quietly.
`scripts/discover-gaps.sh` for network-level inventory discovery.
- `scripts/deploy-stack.sh` + `scripts/sync-stacks.sh` for compose push/pull.
- `scripts/add-host.sh` for new-host registration.
- **`scripts/elway`** (2026-04-24) — mini-ansible playbook runner over
SSH. Playbooks live under `playbooks/`. Tier 1 (creates/when) + tier
2 (changed_when) idempotency; handlers + register + multi-host
fan-out deferred as gitea issues #3–#5.
- **`tea` CLI** at `/usr/local/bin/tea` — gitea CLI, already logged in
as `vh`. Use for issue / PR work instead of inventing URLs.
## Open issues
### 🟥 Quick wins (do next)
1. **Fix Backrest's `esh-docker-vm` URI** — Backrest is pointed at
`10.100.50.50` (NH3 Synology) for that repo, while the actual backup
writes go to `10.250.50.70` (rest-server-ana). Edit `config.json` in
the backrest container, swap the host portion. ~5 min.
1. ~~**Fix Backrest's `esh-docker-vm` URI**~~ — done.
2. **Verify SSH on the 9 newly-registered hosts** and pull first snapshots:
```
scripts/refresh-server-info.sh --validate-only all
scripts/refresh-server-info.sh pfi-ana-webhost ana-filebot pfi-pteradactyl \
pfi-tacticalrmm pfi-postgres ana-wg
```
ssh-target files were guessed (`lkraven@` for VMs, `root@` for LXCs) —
first run surfaces mismatches.
2. ~~**Push SSH keys + pull snapshots for 6 unrefreshed hosts**~~ —
**done 2026-04-23**. 4 VMs (pfi-ana-webhost, pfi-pteradactyl,
pfi-tacticalrmm, pfi-postgres) took the workstation key via
`ssh-copy-id` with `lkraven@`. 2 LXCs (ana-filebot, ana-wg)
needed key installed via `pct push` from pfi-pve because
`PermitRootLogin prohibit-password` blocked `ssh-copy-id`.
Minor: pfi-pteradactyl's `server_inspect` Docker section is blank
because lkraven is not in the `docker` group there — `sudo usermod
-aG docker lkraven` + re-login to fix.
3. **Patch out `forget` schedules from all 6 restic profiles.** They fail
nightly against `--append-only` rest-servers. One-liner per host:
comment out `schedule:` under the `forget` block in profiles.yaml;
redeploy; run `resticprofile unschedule forget` on each host. ~15 min.
3. ~~**Patch out `forget` schedules from all 6 restic profiles**~~ — done.
3b. ~~**Discover sf-r630 OS-side IP**~~ — **resolved 2026-04-23**:
the R630 with iDRAC `10.250.250.110` is the same physical box
that runs `sfsrv-ana` (Proxmox VE at `10.250.250.115`). No
separate OS IP to find. `servers/sf-r630/` now clarifies it as
the hardware/BMC-only inventory entry; `servers/sfsrv-ana/` is
the OS view. `servers/ana-ml2/README.md` updated with its own
distinct BMC IP (`10.250.250.50`, Supermicro) to prevent future
confusion between the two physical chassis.
<!-- 2026-04-21: Dropped the "nightly-restart safety net" idea — user
confirmed Backrest UI chokes even right after startup, so periodic
@@ -73,63 +121,81 @@ significant work lands — don't let it drift quietly.
### 🟧 Real work (dedicated session each)
5. **Rotate exposed secrets** (captured in session transcripts 2026-04-20/21):
- `vaultwarden` Postgres password (on pfi-postgres)
- `gitea` Postgres password — currently literally `gitea` (trivially weak)
- `paperless-ng` Postgres password — currently literally `paperless-ng` (trivially weak)
- `ana-docker` rest-server htpasswd + repo passphrase
- `ana-ml2` rest-server htpasswd (repo passphrase rotated during wipe/reinit)
- `esh-docker-vm` rest-server htpasswd + repo passphrase
4b. ~~**Migrate DB data directories off NFS onto local VM disk.**~~ —
**done 2026-04-23/24.** Postgres on pfi-postgres migrated
04-23; mongodb on esh-vm-db confirmed already local (and serves
zero user data in practice — paperless uses Postgres 15 on the
same VM). fstab entry + NFS mount for `/mnt/db` on pfi-postgres
removed 2026-04-24 via
`playbooks/decouple-pfi-postgres-from-ana-nas.yaml`. pfi-postgres
has zero remaining dependency on ana-nas. Only residue: cold
archive of pre-migration data still on ana-nas (`/mnt/db/pfi-*`);
harmless, can sit indefinitely.
For each: rotate at the source (DB `ALTER USER …` / `htpasswd` /
`restic key add` + `restic key remove`) → update consumer config
files (`/etc/restic/restic.env`, `/etc/restic/dbcreds.env`, compose
env files, Backrest config.json) → restart consumer services. ~45 min
batched.
5. ~~**Rotate exposed secrets**~~ — **done 2026-04-23**. All six
rotated: vaultwarden/gitea/paperless-ng Postgres passwords
(hardcoded `compose.yaml` literals moved to gitignored `.env`
files in the process), ana-docker + ana-ml2 + esh-docker-vm
rest-server htpasswd entries, and rest-server repo passphrases for
ana-docker + esh-docker-vm (ana-ml2's was already rotated during
prior wipe+reinit). Also discovered along the way: paperless uses
a separate `esh-vm-db` VM (10.0.50.60), not pfi-postgres as the
old notes implied.
6. **Cross-site rsync** between `rest-server-ana` data dir
(`/mnt/backup/restic/repo/ana/`) and NH3 Synology data dir. Planned
since initial rest-server setup; not built. Needs Synology SSH
access first (item 8). ~20 min once access is there.
- Note: cross-site mirroring for the VM-image layer is being
addressed by the PBS deployment (item 6b) — this rsync is now
scoped to restic repos only.
6. ~~**Cross-site rsync**~~ — **done 2026-04-23, both directions**.
- **ana → nh3** (04:00 daily): `ana-nas:/mnt/backup/restic/repo/ana/`
→ `nh3-nas:/volume1/Backup/restic-ana-mirror/`. Runs on ana-nas
as lkraven. 15.6 GB initial sync completed 07:18 UTC.
- **nh3 → ana** (05:00 daily): `nh3-nas:/volume1/Backup/restic/`
→ `ana-nas:/mnt/backup/restic-nh3-mirror/`. Runs on nh3-nas as
root (rest-server-nh3 container writes mode-400 files; only
root can read them on Synology). 2.31 GB initial sync completed
07:42 UTC.
- Tracked at `configs/rsync/ana-nas-to-nh3/` and
`configs/rsync/nh3-nas-to-ana/` respectively.
6b. **PBS deployment across the fleet.** Runbook lives at
`docs/runbooks/pbs-deployment.md`. Primary at ANA (VM on pfi-pve,
NFS datastore on Debian NAS), DR mirror at NH3 (VM on nh3-pve,
local datastore). One-way sync from ANA → NH3 nightly. All 5
6b. **PBS deployment across the fleet.** Runbook at
`docs/runbooks/pbs-deployment.md`. Phases 0–6 **done** (2026-04-22):
PBS-ANA on pfi-pve with NFS datastore on Ana NAS (NFSv3 for ZFS
case-insensitivity), PBS-NH3 on nh3-pve with NFS datastore on NH3
Synology (NFSv3 + Linux-mode POSIX 777 to avoid syno_acl
interference), per-hypervisor namespaces, API tokens, verify jobs
on both, one-way sync ANA → NH3 at 06:00 daily, and all 5
hypervisors (pfi-pve, nh3-pve, esh-pve, esh-pve-nas, sfsrv-ana)
migrate off local-dump vzdump jobs onto PBS.
- Closes: SureFire's zero-coverage gap (currently no vzdump at all on sfsrv-ana)
- Closes: "cross-site redundancy for VM images" scope item
- Provides: cross-fleet block-level dedupe, dirty-bitmap incrementals,
real safe-prune retention, per-host encryption keys
- Est. 4–6h spread across sittings. 9 phases in the runbook.
onboarded to PBS-ANA.
Remaining phases:
- Phase 7-8 — one week burn-in, then retire legacy vzdump targets
on each hypervisor (keep until 2026-04-29 earliest)
- Phase 9 — homepage cards for PBS-ANA + PBS-NH3 web UIs,
memory/doc updates, `servers/pbs-ana` + `servers/pbs-nh3` dirs
7. **SureFire tenant backup plan decision.** Three options documented in
`servers/sfsrv-ana/README.md`:
- Tenant handles own backups
- PFI provides dedicated scoped repo on rest-server-ana
- Shared vzdump target
Blocks any actual SF backup work until the hosting-agreement side of
this is clear.
7. ~~**SureFire tenant backup plan decision**~~ — **resolved 2026-04-23**
by folding sfsrv-ana into the fleet-wide PBS deployment (dedicated
`sfsrv-ana` namespace on PBS-ANA, replicates to PBS-NH3 via the
same sync job as the rest of the fleet). Hosting-agreement option
chosen: PFI provides backup coverage as part of managed hosting.
### 🟨 Prereqs / polish
8. **Synology SSH setup** — tabled earlier. Unlocks: cross-site rsync
(item 6), rest-server-nh3 healthcheck deploy, easier `.htpasswd`
edits on the NH3 side.
8. ~~**Synology SSH setup**~~ — **done 2026-04-22**. Dedicated
`syncuser` account (admin-group membership) with key auth,
registered as `servers/nh3-nas/`, reachable as `ssh nh3-nas`.
Unlocks cross-site rsync (item 6), rest-server-nh3 healthcheck
deploy, and `.htpasswd` edits on the NH3 side.
9. **`scripts/restic-prune.sh`** — temporarily flip `--append-only` off,
run forget + prune across all hosts, flip back on. Needed quarterly
for disk hygiene. Not urgent; blocks only the "I need to reclaim
disk space now" scenario.
10. **Retire `offen/docker-volume-backup` sidecars on esh-docker-vm** once
restic has ~1 week of clean runs. Paperless + pgadmin currently run
offen sidecars that write tarballs to `/mnt/backup/...` redundantly.
10. ~~**Retire `offen/docker-volume-backup` sidecars**~~ — **done
2026-04-23**. Removed from paperless-ngx and pgadmin composes on
esh-docker-vm (only hosts in the fleet that had them). 16 GB of
orphan tarballs at `/mnt/backup/docker/esh-vm-docker/` reclaimed.
Restic coverage verified equivalent: `/var/lib/docker/volumes` in
source + pre-backup hook handles the DB dumps for paperless
(against esh-vm-db), HA, pgadmin, CWA, uptime-kuma. Obsolete
`*_offen_backup_data` exclude also removed from the restic profile.
11. **Clean up retired mattermost dir** on ana-docker — compose dir at
`/opt/docker/compose/mattermost/` may still linger. Single `rm -rf`
@@ -159,6 +225,25 @@ significant work lands — don't let it drift quietly.
lands. Consider a `scripts/status-regen.sh` if manual updates
slip.
## Session milestones — 2026-04-24 (the "tooling day")
- **irv-ml1 AI stacks deployed** as Docker: ComfyUI (port 8188, host-
writable workflows), Parakeet ASR (port 8765, rewritten on
sherpa-onnx after the Shadowfita FastAPI wrapper hit two unfixed
upstream bugs), CosyVoice 3 (port 8190, GLaDOS voice cloned + live).
- **irv-ml1 restic profile extended** to cover
`/worktank/{comfyui/basedir/{user,custom_nodes,input},cosyvoice/voices}`;
bulk weights + scratch dirs stay excluded.
- **`scripts/elway`** shipped — ~800-line Python playbook runner with
tier 1 + tier 2 idempotency; handlers / register / multi-host /
content-hash are deferred (gitea #3–#6).
- **task-board** built end-to-end (separate repo, `vh/task-board` on
gitea) and shipped as a Claude Code plugin. Green/red cards per
session via UserPromptSubmit + Stop hooks; four MCP tools expose
explicit activity tracking.
- **pfi-postgres NFS decoupling** finished item 4b — zero residual
dependency on ana-nas for that VM.
## Session milestones — 2026-04-20 / 2026-04-21
- Homepage reorganized to function-first with tabs (Main / Infrastructure / Toolchain).
@@ -181,10 +266,20 @@ significant work lands — don't let it drift quietly.
Relevant `~/.claude/.../memory/` entries:
- `server_split.md` — host placement rules
- `feedback_ssh_sudo.md` — use `ssh -t` for remote sudo
- `feedback_ssh_sudo.md` — use `ssh -t` for remote sudo (mostly
superseded by elway, but still applies to ad-hoc ssh)
- `feedback_git_autonomous.md` — handle git commits without asking
- `feedback_git_commits.md` — no Claude attribution in commit messages
- `project_backup_pipeline_gaps.md` — user's explicit goal of "all hosts +
configs + DBs backed up"
- `feedback_use_elway.md` — write elway playbooks; don't chain ssh+sudo
- `feedback_backups_must_not_risk_production.md` — rule adopted after
the 2026-04-23 ana-nas self-backup crash
- `project_backup_pipeline_gaps.md` — user's explicit goal of "all
hosts + configs + DBs backed up"
- `project_db_migrate_off_nfs.md` — DB-off-NFS decision + status
- `project_surefire_tenant.md` — SureFire tenancy boundary awareness
- `storage_ana_nas.md` — ana NAS is Debian, not TrueNAS
- `reference_gitea_remote.md` — origin is `vh/esh-pfi-infrastructure`
on gitea.phasefinal.com; `tea` CLI logged in as `vh`
- `reference_task_board.md` — task-board plugin + tools contract
- `storage_ana_nas.md` — ana NAS is Debian LXC (CT 109), not TrueNAS
- `storage_nh3_nas.md` — NH3 NAS via `syncuser`, not `admin`
- `incident_ana_nas_spof.md` — blast-radius matrix for ana-nas outages