Files
esh-pfi-infrastructure/README.md
T
vh 91bda3c480 fv-ml1: complete the cutover — rename, renumber, DNS, and the LiteLLM repoint
The box is physically at Fountain Valley, renamed, renumbered onto 10.251/16,
and serving inference again. This lands the repo half of that.

Host: hostname ana-ml2 -> fv-ml1, pinned to 10.251.50.54 by a dnsmasq
reservation so the address the runbook, DNS and LiteLLM all assume is the
address it actually has. Its headscale node is renamed too.

The sweep ran from scripts/fv-ml1-rename-sweep.sh, whose allowlist is the
reason this diff touches current-state files and not the record. Dated
persistent-memory entries, archival-memory and incident notes still say
ana-ml2 in 31 and 62 places respectively, because that is what the box was
when those things happened. Rewriting them would make the history lie.

LiteLLM was the load-bearing piece and needed more than the api_base sed the
runbook describes. Twenty api_base entries repointed, but a grep-and-verify
pass also caught a LIVE pass_through_endpoints target for the scalar-judge
reward route still on the old address -- an api_base-only substitution would
have left it dead. Four prose references describing current state were
repointed as well; one historical note recording where a hand-test was run
is deliberately left pointing at 10.250.50.54.

Two facts in the server tables were wrong and are corrected here. The site is
Fountain Valley, not Anaheim. And the box has FOUR RTX PRO 6000 Blackwell
Max-Q, not two -- verified by nvidia-smi -L and independently by PCI
enumeration of four GB202GL devices. That is 391 GB of VRAM rather than 196,
which changes what fits on it.

DNS: fv-ml1, fv-ml1-bmc and fv-gw added under the fv site via the piggyback
approach, scriberr re-homed, and the ana-ml2 records removed. Applied to all
three resolvers. The BMC record carries a warning that its 802.1q VLAN tag
must stay disabled -- it shipped tagging VLAN 250 into an untagged port,
which made it invisible to every network-side diagnostic and is the reason
it appeared dead through several cable changes.

Verified end to end: summarizer and sec both answer through the Anaheim
gateway across the mesh to FV seats on different ports.
2026-09-12 22:00:50 -07:00

190 lines
9.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# eshpfi-management
Infrastructure management workspace for the PFI fleet (plus the ESH home-lab host). Tracks server state, canonical Docker Compose stacks, per-host configs, and the tooling that moves them around.
See **[CLAUDE.md](CLAUDE.md)** for the full set of conventions and the rules Claude Code sessions follow when working here.
## The fleet
**Docker hosts:**
| Host | IP | Site | Role |
|---|---|---|---|
| fv-ml1 | `10.251.50.54` | Fountain Valley (`10.251.0.0/16`) | GPU / AI inference (bare metal, 4× RTX PRO 6000 Blackwell Max-Q) |
| ana-docker | `10.250.50.70` | Anaheim | General-purpose Docker + cross-site hubs (VM on pfi-pve) |
| nh3-docker | `10.100.50.40` | NH3 (`10.100.0.0/16`) | General-purpose Docker (VM on nh3-pve) |
| esh-docker-vm | `10.0.50.45` | ESH home lab (`esteban.net`) | Home-lab Docker (VM on esh-pve, non-PFI scope) |
| vm-esh-nas | `10.0.50.154` | ESH home lab | NAS-adjacent Docker, NFS-mounted shares (VM on esh-pve-nas, non-PFI scope) |
**Proxmox hypervisors** (tracked for inventory; not Docker targets):
| Host | IP | Site | Role |
|---|---|---|---|
| pfi-pve | `10.250.250.31` | Anaheim | Proxmox VE (188 GB / Xeon Silver 4310) |
| nh3-pve | `10.100.250.60` | NH3 | Proxmox VE (62 GB / i9-13900H) |
| esh-pve | `10.0.250.35` | ESH home lab | Proxmox VE (62 GB / i9-13900H) |
| esh-pve-nas | `10.0.50.55` | ESH home lab | Proxmox VE, storage-dedicated (125 GB / Xeon W-1250) |
Per-host snapshots of the running system live under `servers/<host>/system-details.txt`, refreshed via `scripts/refresh-server-info.sh`.
## Layout
```
.
├── CLAUDE.md # conventions; loaded by Claude Code sessions
├── README.md # this file
├── scripts/ # workstation tooling
│ ├── server_inspect.sh # read-only diagnostic, runs on remote via stdin
│ ├── proxmox_inspect.sh # Proxmox-aware probe (VMs, LXCs, storage, backup coverage)
│ ├── refresh-server-info.sh # pull fresh system-details.txt for one/all hosts
│ ├── refresh-proxmox-info.sh # pull fresh proxmox-details.txt for one/all PVE nodes
│ ├── add-host.sh # register a new server (writes servers/<name>/ssh-target)
│ ├── sync-stacks.sh # pull /opt/docker/{compose,conf}/ → stacks-mirror/
│ ├── deploy-stack.sh # push stacks-mirror/<host>/<stack>/ with diff + prompt
│ ├── discover-fortigate.sh # DHCP lease list from a FortiGate via SSH
│ ├── discover-unifi.sh # client list from a UniFi Controller via REST
│ └── discover-gaps.sh # find IPs in discovery TSVs not tracked in servers/
├── servers/ # per-host notes + latest snapshot + ssh-target fallback
│ └── <host>/
│ ├── README.md
│ ├── system-details.txt # regenerate on demand
│ └── ssh-target # <ip> or <user>@<ip>, used when DNS fails
├── stacks/ # canonical compose files (source of truth)
│ └── <stack>/
│ ├── compose.yaml
│ ├── .env.example
│ └── README.md
├── stacks-mirror/ # gitignored — live mirror from sync-stacks.sh
├── configs/ # host-level config files that aren't docker-compose
│ ├── homepage/ # canonical config for the fleet dashboard (on esh-docker-vm)
│ └── restic/<host>/ # resticprofile configs + pre-backup hooks
└── docs/ # general reference (network, models, proxmox, etc.)
└── pfi/
```
## Current stacks
**GPU (fv-ml1):**
- `llama-swap` — GGUF model swapper via llama.cpp (port 9292)
- `vllm` — embeddings (8001) + reranker (8002) + Skywork reward classifier (8003) via vLLM
**Anaheim non-GPU (ana-docker):**
- `traefik`, `crowdsec`, `gitea`, `vaultwarden`, `synapse`, `seafile`, `searxng`, `openwebui`, `sillytavern`, `mailrise`, `rustdesk`, `dockge`, `it-tools`
- (`mattermost` retired 2026-04-21 — compose dir may still linger, containers gone)
- Notes / feeds: `miniflux` (RSS, 8080), `nevermore` (twice-daily LLM-curated brief, 8181, multi-tenant — extracted to its own repo at [`vh/nevermore`](https://gitea.phasefinal.com/vh/nevermore)), `memos` (note server, 5230)
- Assistant tooling: `task-board` (MCP + dashboard for assistant task state, 7878)
- Fleet services: `beszel` (metrics hub, port 8090), `dozzle-hub` (log viewer, 8088), `backrest` (restic UI, 9898)
- Backup target: `rest-server-ana` on port 8000
**GPU fv-ml1 (non-canonical for now):**
- `comfyui`, `kokoro`, `parakeet`, `vibevoice` alongside the canonical `llama-swap` + `vllm`
**NH3 (nh3-docker):**
- `adguard`, `dockge`, plus Beszel/Dozzle agents
**NH3 (Synology `10.100.50.50`):**
- `rest-server-nh3` — restic backup target (port 8000)
**ESH home lab (esh-docker-vm):**
- `adguard`, `homeassistant` (macvlan), `esphome`, `mosquitto`, `paperless-ngx`, `pgadmin`, `calibre-web-automated`, `drawio`, `traefik`, `homepage`, `uptime-kuma`, plus Beszel/Dozzle agents
**ESH home lab (vm-esh-nas):**
- `filezilla` (web UI on port 5800), `dockge`, plus Beszel/Dozzle agents. Mounts `/mnt/{share,music,books,media}` from the Debian NAS at 10.0.50.50.
## Common tasks
**Refresh one host's snapshot:**
```bash
scripts/refresh-server-info.sh ana-docker
```
**Refresh all hosts:**
```bash
scripts/refresh-server-info.sh all
```
**Refresh all Proxmox nodes (separate flow — captures VM/LXC/backup-coverage):**
```bash
scripts/refresh-proxmox-info.sh all
```
**Add a new host:**
```bash
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name>
```
**Validate discovery (without hitting the network):**
```bash
scripts/refresh-server-info.sh --validate-only all
```
**Push a stack to a host (with diff + confirm):**
```bash
scripts/deploy-stack.sh <host> <stack>
```
**Pull every server's compose/conf trees into stacks-mirror/ (not committed — see `.gitignore`):**
```bash
scripts/sync-stacks.sh all
```
**Bootstrap a new fleet repo from this one:**
```bash
scripts/fork-fleet.sh ~/development/<other-fleet>-management
# optional: override the fleet-name baked into the skeleton CLAUDE.md / README
scripts/fork-fleet.sh ~/development/<dest> <fleet-name>
```
Mirrors the reusable tooling (everything under `scripts/`, the
generic playbook templates, `.gitignore`, conventions section of
`CLAUDE.md`) into the destination dir, strips fleet-specific content
(`servers/`, `stacks/`, `configs/`, fleet-named playbooks
(`deploy-*`, `decouple-*`), `docs/orientation.md`, `docs/runbooks/`,
`docs/pfi/`, `STATUS.md`, host-pinned helper scripts), regenerates
skeleton `CLAUDE.md` / `README.md` / `STATUS.md` with the new fleet
name, and initializes a fresh git history with one "Initial commit".
The script does **not** create a gitea remote or push — namespace +
repo name are an explicit manual step. The "next steps" output prints
the exact `tea repo create` + `git remote add` + `git push` commands
to wire it up when ready.
## Backup pipeline
Backups are driven by per-host `resticprofile` configs under `configs/restic/<host>/`, scheduled via systemd timers on each host:
- **Writes**: each host backs up to its site-local rest-server (`rest-server-ana` or `rest-server-nh3`), over HTTP basic-auth.
- **Authentication**: shared `.htpasswd` file on both rest-servers, one entry per host; credentials stored in `/etc/restic/restic.env` on each client host.
- **Encryption**: per-host client-side passphrase in `/etc/restic/password` (unique per repo; losing it = losing that host's backups).
- **Visibility**: Backrest (`http://10.250.50.70:9898`) shows every repo for browsing/restore.
- **Schedule**: backup at 01:00 daily, `forget` at 03:00 daily, weekly `check --read-data-subset 10%` on Sundays.
- **Prune**: manual ceremony (rest-server runs with `--append-only`, which blocks destructive prune ops).
- **Off-site**: cross-site rsync between the two rest-server data dirs is planned (not yet implemented).
### Coverage status (as of 2026-04-29)
Goal: **every Docker host + configs + every database** covered, not just VM images.
| Layer | State |
|---|---|
| VM-level (Proxmox vzdump) | ✅ All running guests covered across pfi-pve / nh3-pve / esh-pve-nas; esh-pve has VMID 108 uncovered |
| ana-docker restic (host files + DBs) | ✅ `configs/restic/ana-docker/` with pre-backup hooks for synapse / seafile / vaultwarden-pg / gitea (native dump) / openwebui |
| fv-ml1 restic | ✅ `configs/restic/fv-ml1/` — bare-metal host files (no DB hooks needed) |
| nh3-docker restic | ✅ Light — no DB hooks needed |
| esh-docker-vm restic | ✅ With DB hooks for paperless-postgres (external), home-assistant + pgadmin + uptime-kuma (host-side sqlite3), calibre-web-automated (in-container sqlite3) |
| vm-esh-nas restic | ✅ Light — NFS mounts explicitly excluded |
| nh3-dev (workstation) restic | ✅ `/home/lkraven` + `/etc` with language-toolchain and build-output excludes |
| irv-ml1 restic | ✅ Cross-site to rest-server-nh3 via WireGuard |
| esh-vm-db restic | ✅ With pg_dumpall + mongodump pre-backup hooks |
| Cross-site redundancy | ✅ `ana-nas → nh3-nas` rsync at 04:00 daily; `nh3-nas → ana-nas` at 05:00 daily (`configs/rsync/`) |
| Prune ceremony | ✅ Quarterly manual ritual documented in `docs/runbooks/nh3-prune-ritual.md` (DSM-mediated `--append-only` toggle) |
| offen/docker-volume-backup sidecars on esh-docker-vm | ❌ Remove once restic proves itself (~1 week of clean runs) |
## Authoritative vs. mirror
- **Authoritative:** files on each server under `/opt/docker/compose/<stack>/` and `/opt/docker/conf/<stack>/`.
- **This workspace:** source-of-truth copies under `stacks/<name>/` (hand-curated), and a gitignored mirror under `stacks-mirror/` pulled by `sync-stacks.sh`.
Edit in `stacks/`, push with `deploy-stack.sh`. Never commit `stacks-mirror/` — it can contain embedded plaintext secrets from upstream compose files that haven't been audited yet.