comfyui: add stack + deploy to irv-ml1

New stack mirroring PFI convention (stacks/comfyui/) using
mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-20260312. Both GPUs
exposed, pinned to CUDA 12.8 to match the host's 570.x driver and the
native cuda-toolkit already in place.

Layout — single tree under /worktank/comfyui/ (462G dedicated, 1%
used pre-deploy):
  - basedir/  → /basedir   user state (models, workflows, custom_nodes,
                           input, output); owned 1000:1000 so external
                           tools can edit workflow JSON directly.
  - run/      → /comfy/mnt  ComfyUI source + venv + pip cache (~7.8G
                           after bootstrap). Bind mount instead of
                           named volume — the image refuses to chown
                           mounted paths at startup, so keeping this
                           lkraven-owned avoids the sudo dance.

servers/irv-ml1/README.md refreshed: Docker upgraded to 29.4.1 with
traefik-net in place; dockge + beszel-agent + dozzle-agent already
present; /storetank dropped 92% → 64%; restic coverage to
rest-server-nh3 is operational (not "currently none" as prior text).
This commit is contained in:
vh
2026-04-23 22:56:21 -07:00
parent ba55f5e902
commit 06476745f1
5 changed files with 723 additions and 0 deletions
+105
View File
@@ -0,0 +1,105 @@
# irv-ml1
Secondary AI/ML inference host at the Irvine site. Formerly known as
`ana-ml1` when colocated at Anaheim; moved to Irvine and slated for
hostname rename to `irv-ml1` (OS-side rename pending — see below).
## Network
- **Reachable IP:** `10.100.79.3` (WireGuard tunnel endpoint)
- **No direct LAN access** — this host is reachable **only** via
WireGuard. Tunnel terminates at the NH3 site (10.100.0.0/16 WG
subnet). If WG is down, `scripts/refresh-server-info.sh irv-ml1`
will fail with "No route to host" — that's a WG issue, not a host
issue.
- **SSH:** `ssh irv-ml1` (config alias → `lkraven@10.100.79.3`,
key auth).
## Pending hostname rename
OS hostname still reports `ana-ml1` (both in `hostnamectl` and in
`system-details.txt`). To finish the rename:
```bash
ssh -t irv-ml1 'sudo hostnamectl set-hostname irv-ml1; \
sudo sed -i "s/ana-ml1/irv-ml1/g" /etc/hosts; \
cat /etc/hosts; hostname'
```
Then refresh the inventory snapshot so it reflects the new identity.
Not blocking anything — services don't care about the kernel's idea
of hostname.
## Hardware
- **Chassis:** (TBD — captured on next physical inspection)
- **CPU:** AMD Ryzen Threadripper 3970X (32 cores / 64 threads)
- **RAM:** 251.6 GB
- **GPUs:** 2× (unlike ana-ml2's matched pair):
- GPU 0: **NVIDIA GeForce RTX 3090** (24 GB VRAM)
- GPU 1: **NVIDIA RTX A6000** (48 GB VRAM)
- Total VRAM: 72 GB across both
- **OS:** Debian 12 (bookworm), kernel 6.1.0-37
- **Storage:**
- `/` on `/dev/nvme0n1p2` — 1.8 TB (78% used, ~393 GB free)
- `/worktank` — 462 GB (1% used — dedicated to Docker stacks' user state, e.g. ComfyUI models + workflows)
- `/storetank` — 1.8 TB (64% used, ~660 GB free)
## What it runs
### Native toolchain (`/opt`, owned by `llmuser`)
Predates the PFI docker convention; still the primary runtime for the
generative-AI stack:
- ComfyUI, SillyTavern, SDNext, fluxgym (image gen / SD)
- alltalk, alltalkv2, bark, kokoro, Orpheus-FastAPI, stablediffusion (TTS + voice)
- llama.cpp, llama-swap, koboldcpp, aphrodite (LLM inference)
- ollama (port 11434 listening on all interfaces) — native binary, not the container
- ai-toolkit, chat-ui, h2ogpt, o-textgen, lollms, bitsandbytes (misc ML frameworks)
- sillytavern-extras, simple-proxy-for-tavern (lkraven-owned)
### Docker stacks (`/opt/docker/compose/`, owned by `lkraven`)
Docker 29.4.1 with `nvidia` and `runc` runtimes. `lkraven` is in the
`docker` group. `traefik-net` external network exists for stacks that
need it.
| Stack | Port | Role |
|-------|------|------|
| dockge | 5001 | Per-host Compose UI |
| beszel-agent-irv | 45876 | Metrics agent → Beszel hub on ana-docker (token mode through WG) |
| dozzle-agent-irv | 7007 | Log agent → Dozzle hub on ana-docker |
| comfyui | 8188 | ComfyUI (node-based SD/Flux) — runs independently of `/opt/ComfyUI` native install |
Exposed Docker socket on `*:2375` (for the homepage integration hub on
esh-docker-vm, which auto-discovers containers on this host).
## Storage watch
Nothing acute. `/storetank` dropped from 92% → 64% after a prune pass
on the native-toolchain side; keep an eye on it since model weights
and training outputs accumulate steadily (misbehavior starts around
~95% on either ext4 or ZFS).
## Backup coverage
Restic via `resticprofile` + systemd timer (01:00 daily) → `rest-server-nh3`
(local to the WG endpoint site; lower latency than crossing back to
ana-side). Profile tracked at `configs/restic/irv-ml1/profiles.yaml`.
Excludes HuggingFace caches and bulk model files on `/storetank`
(regenerable from HF Hub).
When the ComfyUI stack ships, `/worktank/comfyui/basedir/{user,custom_nodes,input}`
should be added to the source set (workflows + hand-installed nodes);
`/worktank/comfyui/basedir/{models,output}` stay excluded (bulk /
regenerable).
## Refresh state
```bash
scripts/refresh-server-info.sh irv-ml1
```
**Caveat:** requires the WG tunnel to be up. If the refresh shows
"No route to host", bring WG up before retrying.
File diff suppressed because one or more lines are too long