comfyui: add stack + deploy to irv-ml1
New stack mirroring PFI convention (stacks/comfyui/) using
mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-20260312. Both GPUs
exposed, pinned to CUDA 12.8 to match the host's 570.x driver and the
native cuda-toolkit already in place.
Layout — single tree under /worktank/comfyui/ (462G dedicated, 1%
used pre-deploy):
- basedir/ → /basedir user state (models, workflows, custom_nodes,
input, output); owned 1000:1000 so external
tools can edit workflow JSON directly.
- run/ → /comfy/mnt ComfyUI source + venv + pip cache (~7.8G
after bootstrap). Bind mount instead of
named volume — the image refuses to chown
mounted paths at startup, so keeping this
lkraven-owned avoids the sudo dance.
servers/irv-ml1/README.md refreshed: Docker upgraded to 29.4.1 with
traefik-net in place; dockge + beszel-agent + dozzle-agent already
present; /storetank dropped 92% → 64%; restic coverage to
rest-server-nh3 is operational (not "currently none" as prior text).
This commit is contained in:
@@ -0,0 +1,105 @@
|
||||
# irv-ml1
|
||||
|
||||
Secondary AI/ML inference host at the Irvine site. Formerly known as
|
||||
`ana-ml1` when colocated at Anaheim; moved to Irvine and slated for
|
||||
hostname rename to `irv-ml1` (OS-side rename pending — see below).
|
||||
|
||||
## Network
|
||||
|
||||
- **Reachable IP:** `10.100.79.3` (WireGuard tunnel endpoint)
|
||||
- **No direct LAN access** — this host is reachable **only** via
|
||||
WireGuard. Tunnel terminates at the NH3 site (10.100.0.0/16 WG
|
||||
subnet). If WG is down, `scripts/refresh-server-info.sh irv-ml1`
|
||||
will fail with "No route to host" — that's a WG issue, not a host
|
||||
issue.
|
||||
- **SSH:** `ssh irv-ml1` (config alias → `lkraven@10.100.79.3`,
|
||||
key auth).
|
||||
|
||||
## Pending hostname rename
|
||||
|
||||
OS hostname still reports `ana-ml1` (both in `hostnamectl` and in
|
||||
`system-details.txt`). To finish the rename:
|
||||
|
||||
```bash
|
||||
ssh -t irv-ml1 'sudo hostnamectl set-hostname irv-ml1; \
|
||||
sudo sed -i "s/ana-ml1/irv-ml1/g" /etc/hosts; \
|
||||
cat /etc/hosts; hostname'
|
||||
```
|
||||
|
||||
Then refresh the inventory snapshot so it reflects the new identity.
|
||||
Not blocking anything — services don't care about the kernel's idea
|
||||
of hostname.
|
||||
|
||||
## Hardware
|
||||
|
||||
- **Chassis:** (TBD — captured on next physical inspection)
|
||||
- **CPU:** AMD Ryzen Threadripper 3970X (32 cores / 64 threads)
|
||||
- **RAM:** 251.6 GB
|
||||
- **GPUs:** 2× (unlike ana-ml2's matched pair):
|
||||
- GPU 0: **NVIDIA GeForce RTX 3090** (24 GB VRAM)
|
||||
- GPU 1: **NVIDIA RTX A6000** (48 GB VRAM)
|
||||
- Total VRAM: 72 GB across both
|
||||
- **OS:** Debian 12 (bookworm), kernel 6.1.0-37
|
||||
- **Storage:**
|
||||
- `/` on `/dev/nvme0n1p2` — 1.8 TB (78% used, ~393 GB free)
|
||||
- `/worktank` — 462 GB (1% used — dedicated to Docker stacks' user state, e.g. ComfyUI models + workflows)
|
||||
- `/storetank` — 1.8 TB (64% used, ~660 GB free)
|
||||
|
||||
## What it runs
|
||||
|
||||
### Native toolchain (`/opt`, owned by `llmuser`)
|
||||
|
||||
Predates the PFI docker convention; still the primary runtime for the
|
||||
generative-AI stack:
|
||||
|
||||
- ComfyUI, SillyTavern, SDNext, fluxgym (image gen / SD)
|
||||
- alltalk, alltalkv2, bark, kokoro, Orpheus-FastAPI, stablediffusion (TTS + voice)
|
||||
- llama.cpp, llama-swap, koboldcpp, aphrodite (LLM inference)
|
||||
- ollama (port 11434 listening on all interfaces) — native binary, not the container
|
||||
- ai-toolkit, chat-ui, h2ogpt, o-textgen, lollms, bitsandbytes (misc ML frameworks)
|
||||
- sillytavern-extras, simple-proxy-for-tavern (lkraven-owned)
|
||||
|
||||
### Docker stacks (`/opt/docker/compose/`, owned by `lkraven`)
|
||||
|
||||
Docker 29.4.1 with `nvidia` and `runc` runtimes. `lkraven` is in the
|
||||
`docker` group. `traefik-net` external network exists for stacks that
|
||||
need it.
|
||||
|
||||
| Stack | Port | Role |
|
||||
|-------|------|------|
|
||||
| dockge | 5001 | Per-host Compose UI |
|
||||
| beszel-agent-irv | 45876 | Metrics agent → Beszel hub on ana-docker (token mode through WG) |
|
||||
| dozzle-agent-irv | 7007 | Log agent → Dozzle hub on ana-docker |
|
||||
| comfyui | 8188 | ComfyUI (node-based SD/Flux) — runs independently of `/opt/ComfyUI` native install |
|
||||
|
||||
Exposed Docker socket on `*:2375` (for the homepage integration hub on
|
||||
esh-docker-vm, which auto-discovers containers on this host).
|
||||
|
||||
## Storage watch
|
||||
|
||||
Nothing acute. `/storetank` dropped from 92% → 64% after a prune pass
|
||||
on the native-toolchain side; keep an eye on it since model weights
|
||||
and training outputs accumulate steadily (misbehavior starts around
|
||||
~95% on either ext4 or ZFS).
|
||||
|
||||
## Backup coverage
|
||||
|
||||
Restic via `resticprofile` + systemd timer (01:00 daily) → `rest-server-nh3`
|
||||
(local to the WG endpoint site; lower latency than crossing back to
|
||||
ana-side). Profile tracked at `configs/restic/irv-ml1/profiles.yaml`.
|
||||
Excludes HuggingFace caches and bulk model files on `/storetank`
|
||||
(regenerable from HF Hub).
|
||||
|
||||
When the ComfyUI stack ships, `/worktank/comfyui/basedir/{user,custom_nodes,input}`
|
||||
should be added to the source set (workflows + hand-installed nodes);
|
||||
`/worktank/comfyui/basedir/{models,output}` stay excluded (bulk /
|
||||
regenerable).
|
||||
|
||||
## Refresh state
|
||||
|
||||
```bash
|
||||
scripts/refresh-server-info.sh irv-ml1
|
||||
```
|
||||
|
||||
**Caveat:** requires the WG tunnel to be up. If the refresh shows
|
||||
"No route to host", bring WG up before retrying.
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user