Files
esh-pfi-infrastructure/README.md
T
vh 7e7130172e vllm: rename stack from vllm-qwen3 → vllm + add Skywork reward classifier
Two related changes shipped together. The stack rename is independent
but adding `vllm-reward` to the existing `vllm-qwen3` would have made
that name actively misleading.

**Rename:** `stacks/vllm-qwen3/ → stacks/vllm/`. Updated all in-repo
references (README.md root, servers/ana-ml2/, stacks/llama-swap/,
configs/restic/ana-ml2/, docs/runbooks/disaster-recovery.md). Two
intentional history mentions retained (servers/ana-ml2 + stacks/vllm
README).

**Add `vllm-reward` service:** serves Skywork-Reward-V2-Llama-3.1-8B-AWQ
on port 8003. The AWQ output is a locally-quantized model (not from HF),
so bind-mounts `/tank/aimodels/llm:/local-models:ro` rather than the
shared HF cache. Model config.json declares LlamaForSequenceClassification
which vLLM's pooling runner picks up automatically — produces a single
reward score per input via /classify.

**Flag note:** the user's spec listed `--task classify`, but vLLM 0.19.1
deprecated --task in favor of --runner pooling (model architecture in
config.json drives the classification head). Compose uses --runner
pooling with a comment explaining the substitution.

**GPU memory:** no rebalance needed — production had already tuned
EMBED/RERANK down from 0.40 to 0.20 each (canonical .env.example now
matches reality). Adding REWARD at 0.30 totals 0.70, leaving ~14 GB
headroom on the 48 GB Ada.

**Server-side:** brought existing vllm-qwen3 down, mv'd
/opt/docker/compose/vllm-qwen3 → /opt/docker/compose/vllm, appended
REWARD_* lines to existing .env (preserving API_KEY/HF_TOKEN), deployed
new compose via scripts/deploy-stack.sh, brought all 3 services up.

**Smoke tests:**
- /health on 8001/8002/8003 → 200
- /v1/models on 8003 → lists Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
  with max_model_len 16384
- /classify with a sample conversation → returns LABEL_0 with prob 0.9999
  (single-output regression-style reward score, expected shape for a
  reward model)
2026-05-13 22:00:26 -07:00

190 lines
9.9 KiB
Markdown

# eshpfi-management
Infrastructure management workspace for the PFI fleet (plus the ESH home-lab host). Tracks server state, canonical Docker Compose stacks, per-host configs, and the tooling that moves them around.
See **[CLAUDE.md](CLAUDE.md)** for the full set of conventions and the rules Claude Code sessions follow when working here.
## The fleet
**Docker hosts:**
| Host | IP | Site | Role |
|---|---|---|---|
| ana-ml2 | `10.250.50.54` | Anaheim (`10.250.0.0/16`) | GPU / AI inference (bare metal) |
| ana-docker | `10.250.50.70` | Anaheim | General-purpose Docker + cross-site hubs (VM on pfi-pve) |
| nh3-docker | `10.100.50.40` | NH3 (`10.100.0.0/16`) | General-purpose Docker (VM on nh3-pve) |
| esh-docker-vm | `10.0.50.45` | ESH home lab (`esteban.net`) | Home-lab Docker (VM on esh-pve, non-PFI scope) |
| vm-esh-nas | `10.0.50.154` | ESH home lab | NAS-adjacent Docker, NFS-mounted shares (VM on esh-pve-nas, non-PFI scope) |
**Proxmox hypervisors** (tracked for inventory; not Docker targets):
| Host | IP | Site | Role |
|---|---|---|---|
| pfi-pve | `10.250.250.31` | Anaheim | Proxmox VE (188 GB / Xeon Silver 4310) |
| nh3-pve | `10.100.250.60` | NH3 | Proxmox VE (62 GB / i9-13900H) |
| esh-pve | `10.0.250.35` | ESH home lab | Proxmox VE (62 GB / i9-13900H) |
| esh-pve-nas | `10.0.50.55` | ESH home lab | Proxmox VE, storage-dedicated (125 GB / Xeon W-1250) |
Per-host snapshots of the running system live under `servers/<host>/system-details.txt`, refreshed via `scripts/refresh-server-info.sh`.
## Layout
```
.
├── CLAUDE.md # conventions; loaded by Claude Code sessions
├── README.md # this file
├── scripts/ # workstation tooling
│ ├── server_inspect.sh # read-only diagnostic, runs on remote via stdin
│ ├── proxmox_inspect.sh # Proxmox-aware probe (VMs, LXCs, storage, backup coverage)
│ ├── refresh-server-info.sh # pull fresh system-details.txt for one/all hosts
│ ├── refresh-proxmox-info.sh # pull fresh proxmox-details.txt for one/all PVE nodes
│ ├── add-host.sh # register a new server (writes servers/<name>/ssh-target)
│ ├── sync-stacks.sh # pull /opt/docker/{compose,conf}/ → stacks-mirror/
│ ├── deploy-stack.sh # push stacks-mirror/<host>/<stack>/ with diff + prompt
│ ├── discover-fortigate.sh # DHCP lease list from a FortiGate via SSH
│ ├── discover-unifi.sh # client list from a UniFi Controller via REST
│ └── discover-gaps.sh # find IPs in discovery TSVs not tracked in servers/
├── servers/ # per-host notes + latest snapshot + ssh-target fallback
│ └── <host>/
│ ├── README.md
│ ├── system-details.txt # regenerate on demand
│ └── ssh-target # <ip> or <user>@<ip>, used when DNS fails
├── stacks/ # canonical compose files (source of truth)
│ └── <stack>/
│ ├── compose.yaml
│ ├── .env.example
│ └── README.md
├── stacks-mirror/ # gitignored — live mirror from sync-stacks.sh
├── configs/ # host-level config files that aren't docker-compose
│ ├── homepage/ # canonical config for the fleet dashboard (on esh-docker-vm)
│ └── restic/<host>/ # resticprofile configs + pre-backup hooks
└── docs/ # general reference (network, models, proxmox, etc.)
└── pfi/
```
## Current stacks
**GPU (ana-ml2):**
- `llama-swap` — GGUF model swapper via llama.cpp (port 9292)
- `vllm` — embeddings (8001) + reranker (8002) + Skywork reward classifier (8003) via vLLM
**Anaheim non-GPU (ana-docker):**
- `traefik`, `crowdsec`, `gitea`, `vaultwarden`, `synapse`, `seafile`, `searxng`, `openwebui`, `sillytavern`, `mailrise`, `rustdesk`, `dockge`, `it-tools`
- (`mattermost` retired 2026-04-21 — compose dir may still linger, containers gone)
- Notes / feeds: `miniflux` (RSS, 8080), `nevermore` (twice-daily LLM-curated brief, 8181, multi-tenant — extracted to its own repo at [`vh/nevermore`](https://gitea.phasefinal.com/vh/nevermore)), `memos` (note server, 5230)
- Assistant tooling: `task-board` (MCP + dashboard for assistant task state, 7878)
- Fleet services: `beszel` (metrics hub, port 8090), `dozzle-hub` (log viewer, 8088), `backrest` (restic UI, 9898)
- Backup target: `rest-server-ana` on port 8000
**GPU ana-ml2 (non-canonical for now):**
- `comfyui`, `kokoro`, `parakeet`, `vibevoice` alongside the canonical `llama-swap` + `vllm`
**NH3 (nh3-docker):**
- `adguard`, `dockge`, plus Beszel/Dozzle agents
**NH3 (Synology `10.100.50.50`):**
- `rest-server-nh3` — restic backup target (port 8000)
**ESH home lab (esh-docker-vm):**
- `adguard`, `homeassistant` (macvlan), `esphome`, `mosquitto`, `paperless-ngx`, `pgadmin`, `calibre-web-automated`, `drawio`, `traefik`, `homepage`, `uptime-kuma`, plus Beszel/Dozzle agents
**ESH home lab (vm-esh-nas):**
- `filezilla` (web UI on port 5800), `dockge`, plus Beszel/Dozzle agents. Mounts `/mnt/{share,music,books,media}` from the Debian NAS at 10.0.50.50.
## Common tasks
**Refresh one host's snapshot:**
```bash
scripts/refresh-server-info.sh ana-docker
```
**Refresh all hosts:**
```bash
scripts/refresh-server-info.sh all
```
**Refresh all Proxmox nodes (separate flow — captures VM/LXC/backup-coverage):**
```bash
scripts/refresh-proxmox-info.sh all
```
**Add a new host:**
```bash
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name>
```
**Validate discovery (without hitting the network):**
```bash
scripts/refresh-server-info.sh --validate-only all
```
**Push a stack to a host (with diff + confirm):**
```bash
scripts/deploy-stack.sh <host> <stack>
```
**Pull every server's compose/conf trees into stacks-mirror/ (not committed — see `.gitignore`):**
```bash
scripts/sync-stacks.sh all
```
**Bootstrap a new fleet repo from this one:**
```bash
scripts/fork-fleet.sh ~/development/<other-fleet>-management
# optional: override the fleet-name baked into the skeleton CLAUDE.md / README
scripts/fork-fleet.sh ~/development/<dest> <fleet-name>
```
Mirrors the reusable tooling (everything under `scripts/`, the
generic playbook templates, `.gitignore`, conventions section of
`CLAUDE.md`) into the destination dir, strips fleet-specific content
(`servers/`, `stacks/`, `configs/`, fleet-named playbooks
(`deploy-*`, `decouple-*`), `docs/orientation.md`, `docs/runbooks/`,
`docs/pfi/`, `STATUS.md`, host-pinned helper scripts), regenerates
skeleton `CLAUDE.md` / `README.md` / `STATUS.md` with the new fleet
name, and initializes a fresh git history with one "Initial commit".
The script does **not** create a gitea remote or push — namespace +
repo name are an explicit manual step. The "next steps" output prints
the exact `tea repo create` + `git remote add` + `git push` commands
to wire it up when ready.
## Backup pipeline
Backups are driven by per-host `resticprofile` configs under `configs/restic/<host>/`, scheduled via systemd timers on each host:
- **Writes**: each host backs up to its site-local rest-server (`rest-server-ana` or `rest-server-nh3`), over HTTP basic-auth.
- **Authentication**: shared `.htpasswd` file on both rest-servers, one entry per host; credentials stored in `/etc/restic/restic.env` on each client host.
- **Encryption**: per-host client-side passphrase in `/etc/restic/password` (unique per repo; losing it = losing that host's backups).
- **Visibility**: Backrest (`http://10.250.50.70:9898`) shows every repo for browsing/restore.
- **Schedule**: backup at 01:00 daily, `forget` at 03:00 daily, weekly `check --read-data-subset 10%` on Sundays.
- **Prune**: manual ceremony (rest-server runs with `--append-only`, which blocks destructive prune ops).
- **Off-site**: cross-site rsync between the two rest-server data dirs is planned (not yet implemented).
### Coverage status (as of 2026-04-29)
Goal: **every Docker host + configs + every database** covered, not just VM images.
| Layer | State |
|---|---|
| VM-level (Proxmox vzdump) | ✅ All running guests covered across pfi-pve / nh3-pve / esh-pve-nas; esh-pve has VMID 108 uncovered |
| ana-docker restic (host files + DBs) | ✅ `configs/restic/ana-docker/` with pre-backup hooks for synapse / seafile / vaultwarden-pg / gitea (native dump) / openwebui |
| ana-ml2 restic | ✅ `configs/restic/ana-ml2/` — bare-metal host files (no DB hooks needed) |
| nh3-docker restic | ✅ Light — no DB hooks needed |
| esh-docker-vm restic | ✅ With DB hooks for paperless-postgres (external), home-assistant + pgadmin + uptime-kuma (host-side sqlite3), calibre-web-automated (in-container sqlite3) |
| vm-esh-nas restic | ✅ Light — NFS mounts explicitly excluded |
| nh3-dev (workstation) restic | ✅ `/home/lkraven` + `/etc` with language-toolchain and build-output excludes |
| irv-ml1 restic | ✅ Cross-site to rest-server-nh3 via WireGuard |
| esh-vm-db restic | ✅ With pg_dumpall + mongodump pre-backup hooks |
| Cross-site redundancy | ✅ `ana-nas → nh3-nas` rsync at 04:00 daily; `nh3-nas → ana-nas` at 05:00 daily (`configs/rsync/`) |
| Prune ceremony | ✅ Quarterly manual ritual documented in `docs/runbooks/nh3-prune-ritual.md` (DSM-mediated `--append-only` toggle) |
| offen/docker-volume-backup sidecars on esh-docker-vm | ❌ Remove once restic proves itself (~1 week of clean runs) |
## Authoritative vs. mirror
- **Authoritative:** files on each server under `/opt/docker/compose/<stack>/` and `/opt/docker/conf/<stack>/`.
- **This workspace:** source-of-truth copies under `stacks/<name>/` (hand-curated), and a gitignored mirror under `stacks-mirror/` pulled by `sync-stacks.sh`.
Edit in `stacks/`, push with `deploy-stack.sh`. Never commit `stacks-mirror/` — it can contain embedded plaintext secrets from upstream compose files that haven't been audited yet.