Three deploy iterations + four backend attempts (subprocess CUDA, resident-server CUDA, Vulkan rebuild) all failed to deliver speedup over fish-s2: * CUDA path: ggml_cuda_init succeeded, weights loaded onto GPU per s2's logs, but nvidia-smi showed 0% utilization during synthesis. Wall time 20s/long phrase vs fish-s2's 7.5s. The "CUDA get_rows unsupported for type q6_K" warning hints at incomplete op coverage in s2.cpp's alpha CUDA backend for fish-speech architecture. * Vulkan path: vk::IncompatibleDriverError on container init. NVIDIA Vulkan ICD not accessible inside the container despite NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics. Would need host-side nvidia-utils-vulkan installation or manual ICD bind mount. Didn't pursue. Both are fixable — CUDA needs op coverage upstream (author actively working on it; "selective embedding dequant" commit landed 16 days ago), Vulkan needs host-side ICD setup. Neither is a config-flip, both are real work for marginal-or-zero return. Better to delete the stack and revisit when s2.cpp matures or when we tackle FP8 quantization on ana-ml2's RTX 6000 Ada (sm_89, native FP8 hardware). Local image rmi'd, /opt/docker/compose/fish-cpp removed on irv-ml1. /worktank/fish-cpp left for user-side sudo cleanup. Future Fish acceleration paths (in order of decreasing certainty): 1. Wait for s2.cpp CUDA op coverage to mature (track upstream commits). 2. Quantize Fish BF16 → FP8 via TransformerEngine, deploy on ana-ml2's RTX 6000 Ada (Ada has native FP8 tensor cores, A6000 doesn't). ~2x speedup if it works. 3. vLLM port of Fish (no upstream support today).
eshpfi-management
Infrastructure management workspace for the PFI fleet (plus the ESH home-lab host). Tracks server state, canonical Docker Compose stacks, per-host configs, and the tooling that moves them around.
See CLAUDE.md for the full set of conventions and the rules Claude Code sessions follow when working here.
The fleet
Docker hosts:
| Host | IP | Site | Role |
|---|---|---|---|
| ana-ml2 | 10.250.50.54 |
Anaheim (10.250.0.0/16) |
GPU / AI inference (bare metal) |
| ana-docker | 10.250.50.70 |
Anaheim | General-purpose Docker + cross-site hubs (VM on pfi-pve) |
| nh3-docker | 10.100.50.40 |
NH3 (10.100.0.0/16) |
General-purpose Docker (VM on nh3-pve) |
| esh-docker-vm | 10.0.50.45 |
ESH home lab (esteban.net) |
Home-lab Docker (VM on esh-pve, non-PFI scope) |
| vm-esh-nas | 10.0.50.154 |
ESH home lab | NAS-adjacent Docker, NFS-mounted shares (VM on esh-pve-nas, non-PFI scope) |
Proxmox hypervisors (tracked for inventory; not Docker targets):
| Host | IP | Site | Role |
|---|---|---|---|
| pfi-pve | 10.250.250.31 |
Anaheim | Proxmox VE (188 GB / Xeon Silver 4310) |
| nh3-pve | 10.100.250.60 |
NH3 | Proxmox VE (62 GB / i9-13900H) |
| esh-pve | 10.0.250.35 |
ESH home lab | Proxmox VE (62 GB / i9-13900H) |
| esh-pve-nas | 10.0.50.55 |
ESH home lab | Proxmox VE, storage-dedicated (125 GB / Xeon W-1250) |
Per-host snapshots of the running system live under servers/<host>/system-details.txt, refreshed via scripts/refresh-server-info.sh.
Layout
.
├── CLAUDE.md # conventions; loaded by Claude Code sessions
├── README.md # this file
├── scripts/ # workstation tooling
│ ├── server_inspect.sh # read-only diagnostic, runs on remote via stdin
│ ├── proxmox_inspect.sh # Proxmox-aware probe (VMs, LXCs, storage, backup coverage)
│ ├── refresh-server-info.sh # pull fresh system-details.txt for one/all hosts
│ ├── refresh-proxmox-info.sh # pull fresh proxmox-details.txt for one/all PVE nodes
│ ├── add-host.sh # register a new server (writes servers/<name>/ssh-target)
│ ├── sync-stacks.sh # pull /opt/docker/{compose,conf}/ → stacks-mirror/
│ ├── deploy-stack.sh # push stacks-mirror/<host>/<stack>/ with diff + prompt
│ ├── discover-fortigate.sh # DHCP lease list from a FortiGate via SSH
│ ├── discover-unifi.sh # client list from a UniFi Controller via REST
│ └── discover-gaps.sh # find IPs in discovery TSVs not tracked in servers/
├── servers/ # per-host notes + latest snapshot + ssh-target fallback
│ └── <host>/
│ ├── README.md
│ ├── system-details.txt # regenerate on demand
│ └── ssh-target # <ip> or <user>@<ip>, used when DNS fails
├── stacks/ # canonical compose files (source of truth)
│ └── <stack>/
│ ├── compose.yaml
│ ├── .env.example
│ └── README.md
├── stacks-mirror/ # gitignored — live mirror from sync-stacks.sh
├── configs/ # host-level config files that aren't docker-compose
│ ├── homepage/ # canonical config for the fleet dashboard (on esh-docker-vm)
│ └── restic/<host>/ # resticprofile configs + pre-backup hooks
└── docs/ # general reference (network, models, proxmox, etc.)
└── pfi/
Current stacks
GPU (ana-ml2):
llama-swap— GGUF model swapper via llama.cpp (port 9292)vllm-qwen3— embeddings (8001) + reranker (8002) via vLLM
Anaheim non-GPU (ana-docker):
traefik,crowdsec,gitea,vaultwarden,synapse,seafile,searxng,openwebui,sillytavern,mailrise,rustdesk,dockge,it-tools- (
mattermostretired 2026-04-21 — compose dir may still linger, containers gone) - Fleet services:
beszel(metrics hub, port 8090),dozzle-hub(log viewer, 8088),backrest(restic UI, 9898) - Backup target:
rest-server-anaon port 8000
GPU ana-ml2 (non-canonical for now):
comfyui,kokoro,parakeet,vibevoicealongside the canonicalllama-swap+vllm-qwen3
NH3 (nh3-docker):
adguard,dockge, plus Beszel/Dozzle agents
NH3 (Synology 10.100.50.50):
rest-server-nh3— restic backup target (port 8000)
ESH home lab (esh-docker-vm):
adguard,homeassistant(macvlan),esphome,mosquitto,paperless-ngx,pgadmin,calibre-web-automated,drawio,traefik,homepage,uptime-kuma, plus Beszel/Dozzle agents
ESH home lab (vm-esh-nas):
filezilla(web UI on port 5800),dockge, plus Beszel/Dozzle agents. Mounts/mnt/{share,music,books,media}from the Debian NAS at 10.0.50.50.
Common tasks
Refresh one host's snapshot:
scripts/refresh-server-info.sh ana-docker
Refresh all hosts:
scripts/refresh-server-info.sh all
Refresh all Proxmox nodes (separate flow — captures VM/LXC/backup-coverage):
scripts/refresh-proxmox-info.sh all
Add a new host:
scripts/add-host.sh <name> <ip-or-user@ip>
scripts/refresh-server-info.sh <name>
Validate discovery (without hitting the network):
scripts/refresh-server-info.sh --validate-only all
Push a stack to a host (with diff + confirm):
scripts/deploy-stack.sh <host> <stack>
Pull every server's compose/conf trees into stacks-mirror/ (not committed — see .gitignore):
scripts/sync-stacks.sh all
Backup pipeline
Backups are driven by per-host resticprofile configs under configs/restic/<host>/, scheduled via systemd timers on each host:
- Writes: each host backs up to its site-local rest-server (
rest-server-anaorrest-server-nh3), over HTTP basic-auth. - Authentication: shared
.htpasswdfile on both rest-servers, one entry per host; credentials stored in/etc/restic/restic.envon each client host. - Encryption: per-host client-side passphrase in
/etc/restic/password(unique per repo; losing it = losing that host's backups). - Visibility: Backrest (
http://10.250.50.70:9898) shows every repo for browsing/restore. - Schedule: backup at 01:00 daily,
forgetat 03:00 daily, weeklycheck --read-data-subset 10%on Sundays. - Prune: manual ceremony (rest-server runs with
--append-only, which blocks destructive prune ops). - Off-site: cross-site rsync between the two rest-server data dirs is planned (not yet implemented).
Coverage status (as of 2026-04-20)
Goal: every Docker host + configs + every database covered, not just VM images.
| Layer | State |
|---|---|
| VM-level (Proxmox vzdump) | ✅ All running guests covered across pfi-pve / nh3-pve / esh-pve-nas; esh-pve has VMID 108 uncovered |
| ana-docker restic (host files + DBs) | ✅ configs/restic/ana-docker/ with pre-backup hooks for synapse / seafile / vaultwarden |
| ana-ml2 restic | ❌ Bare metal — no vzdump, no restic yet. Highest-priority gap |
| nh3-docker restic | ✅ Light — no DB hooks needed |
| esh-docker-vm restic | ✅ With DB hooks for paperless-postgres (external), home-assistant + pgadmin + uptime-kuma (host-side sqlite3), calibre-web-automated (in-container sqlite3) |
| vm-esh-nas restic | ✅ Light — NFS mounts explicitly excluded |
| nh3-dev (workstation) restic | ✅ /home/lkraven + /etc with language-toolchain and build-output excludes |
| DB dumps: mattermost, openwebui, gitea, beszel-hub | ❌ Pre-backup hooks not written yet |
| Cross-site redundancy | ❌ rsync between rest-server-ana ↔ rest-server-nh3 planned |
| Prune ceremony | ❌ scripts/restic-prune.sh planned (rest-server --append-only blocks direct prune) |
| offen/docker-volume-backup sidecars on esh-docker-vm | ❌ Remove once restic proves itself (~1 week of clean runs) |
Authoritative vs. mirror
- Authoritative: files on each server under
/opt/docker/compose/<stack>/and/opt/docker/conf/<stack>/. - This workspace: source-of-truth copies under
stacks/<name>/(hand-curated), and a gitignored mirror understacks-mirror/pulled bysync-stacks.sh.
Edit in stacks/, push with deploy-stack.sh. Never commit stacks-mirror/ — it can contain embedded plaintext secrets from upstream compose files that haven't been audited yet.