comfyui: add stack + deploy to irv-ml1
New stack mirroring PFI convention (stacks/comfyui/) using
mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-20260312. Both GPUs
exposed, pinned to CUDA 12.8 to match the host's 570.x driver and the
native cuda-toolkit already in place.
Layout — single tree under /worktank/comfyui/ (462G dedicated, 1%
used pre-deploy):
- basedir/ → /basedir user state (models, workflows, custom_nodes,
input, output); owned 1000:1000 so external
tools can edit workflow JSON directly.
- run/ → /comfy/mnt ComfyUI source + venv + pip cache (~7.8G
after bootstrap). Bind mount instead of
named volume — the image refuses to chown
mounted paths at startup, so keeping this
lkraven-owned avoids the sudo dance.
servers/irv-ml1/README.md refreshed: Docker upgraded to 29.4.1 with
traefik-net in place; dockge + beszel-agent + dozzle-agent already
present; /storetank dropped 92% → 64%; restic coverage to
rest-server-nh3 is operational (not "currently none" as prior text).
This commit is contained in:
@@ -0,0 +1,105 @@
|
|||||||
|
# irv-ml1
|
||||||
|
|
||||||
|
Secondary AI/ML inference host at the Irvine site. Formerly known as
|
||||||
|
`ana-ml1` when colocated at Anaheim; moved to Irvine and slated for
|
||||||
|
hostname rename to `irv-ml1` (OS-side rename pending — see below).
|
||||||
|
|
||||||
|
## Network
|
||||||
|
|
||||||
|
- **Reachable IP:** `10.100.79.3` (WireGuard tunnel endpoint)
|
||||||
|
- **No direct LAN access** — this host is reachable **only** via
|
||||||
|
WireGuard. Tunnel terminates at the NH3 site (10.100.0.0/16 WG
|
||||||
|
subnet). If WG is down, `scripts/refresh-server-info.sh irv-ml1`
|
||||||
|
will fail with "No route to host" — that's a WG issue, not a host
|
||||||
|
issue.
|
||||||
|
- **SSH:** `ssh irv-ml1` (config alias → `lkraven@10.100.79.3`,
|
||||||
|
key auth).
|
||||||
|
|
||||||
|
## Pending hostname rename
|
||||||
|
|
||||||
|
OS hostname still reports `ana-ml1` (both in `hostnamectl` and in
|
||||||
|
`system-details.txt`). To finish the rename:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh -t irv-ml1 'sudo hostnamectl set-hostname irv-ml1; \
|
||||||
|
sudo sed -i "s/ana-ml1/irv-ml1/g" /etc/hosts; \
|
||||||
|
cat /etc/hosts; hostname'
|
||||||
|
```
|
||||||
|
|
||||||
|
Then refresh the inventory snapshot so it reflects the new identity.
|
||||||
|
Not blocking anything — services don't care about the kernel's idea
|
||||||
|
of hostname.
|
||||||
|
|
||||||
|
## Hardware
|
||||||
|
|
||||||
|
- **Chassis:** (TBD — captured on next physical inspection)
|
||||||
|
- **CPU:** AMD Ryzen Threadripper 3970X (32 cores / 64 threads)
|
||||||
|
- **RAM:** 251.6 GB
|
||||||
|
- **GPUs:** 2× (unlike ana-ml2's matched pair):
|
||||||
|
- GPU 0: **NVIDIA GeForce RTX 3090** (24 GB VRAM)
|
||||||
|
- GPU 1: **NVIDIA RTX A6000** (48 GB VRAM)
|
||||||
|
- Total VRAM: 72 GB across both
|
||||||
|
- **OS:** Debian 12 (bookworm), kernel 6.1.0-37
|
||||||
|
- **Storage:**
|
||||||
|
- `/` on `/dev/nvme0n1p2` — 1.8 TB (78% used, ~393 GB free)
|
||||||
|
- `/worktank` — 462 GB (1% used — dedicated to Docker stacks' user state, e.g. ComfyUI models + workflows)
|
||||||
|
- `/storetank` — 1.8 TB (64% used, ~660 GB free)
|
||||||
|
|
||||||
|
## What it runs
|
||||||
|
|
||||||
|
### Native toolchain (`/opt`, owned by `llmuser`)
|
||||||
|
|
||||||
|
Predates the PFI docker convention; still the primary runtime for the
|
||||||
|
generative-AI stack:
|
||||||
|
|
||||||
|
- ComfyUI, SillyTavern, SDNext, fluxgym (image gen / SD)
|
||||||
|
- alltalk, alltalkv2, bark, kokoro, Orpheus-FastAPI, stablediffusion (TTS + voice)
|
||||||
|
- llama.cpp, llama-swap, koboldcpp, aphrodite (LLM inference)
|
||||||
|
- ollama (port 11434 listening on all interfaces) — native binary, not the container
|
||||||
|
- ai-toolkit, chat-ui, h2ogpt, o-textgen, lollms, bitsandbytes (misc ML frameworks)
|
||||||
|
- sillytavern-extras, simple-proxy-for-tavern (lkraven-owned)
|
||||||
|
|
||||||
|
### Docker stacks (`/opt/docker/compose/`, owned by `lkraven`)
|
||||||
|
|
||||||
|
Docker 29.4.1 with `nvidia` and `runc` runtimes. `lkraven` is in the
|
||||||
|
`docker` group. `traefik-net` external network exists for stacks that
|
||||||
|
need it.
|
||||||
|
|
||||||
|
| Stack | Port | Role |
|
||||||
|
|-------|------|------|
|
||||||
|
| dockge | 5001 | Per-host Compose UI |
|
||||||
|
| beszel-agent-irv | 45876 | Metrics agent → Beszel hub on ana-docker (token mode through WG) |
|
||||||
|
| dozzle-agent-irv | 7007 | Log agent → Dozzle hub on ana-docker |
|
||||||
|
| comfyui | 8188 | ComfyUI (node-based SD/Flux) — runs independently of `/opt/ComfyUI` native install |
|
||||||
|
|
||||||
|
Exposed Docker socket on `*:2375` (for the homepage integration hub on
|
||||||
|
esh-docker-vm, which auto-discovers containers on this host).
|
||||||
|
|
||||||
|
## Storage watch
|
||||||
|
|
||||||
|
Nothing acute. `/storetank` dropped from 92% → 64% after a prune pass
|
||||||
|
on the native-toolchain side; keep an eye on it since model weights
|
||||||
|
and training outputs accumulate steadily (misbehavior starts around
|
||||||
|
~95% on either ext4 or ZFS).
|
||||||
|
|
||||||
|
## Backup coverage
|
||||||
|
|
||||||
|
Restic via `resticprofile` + systemd timer (01:00 daily) → `rest-server-nh3`
|
||||||
|
(local to the WG endpoint site; lower latency than crossing back to
|
||||||
|
ana-side). Profile tracked at `configs/restic/irv-ml1/profiles.yaml`.
|
||||||
|
Excludes HuggingFace caches and bulk model files on `/storetank`
|
||||||
|
(regenerable from HF Hub).
|
||||||
|
|
||||||
|
When the ComfyUI stack ships, `/worktank/comfyui/basedir/{user,custom_nodes,input}`
|
||||||
|
should be added to the source set (workflows + hand-installed nodes);
|
||||||
|
`/worktank/comfyui/basedir/{models,output}` stay excluded (bulk /
|
||||||
|
regenerable).
|
||||||
|
|
||||||
|
## Refresh state
|
||||||
|
|
||||||
|
```bash
|
||||||
|
scripts/refresh-server-info.sh irv-ml1
|
||||||
|
```
|
||||||
|
|
||||||
|
**Caveat:** requires the WG tunnel to be up. If the refresh shows
|
||||||
|
"No route to host", bring WG up before retrying.
|
||||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,51 @@
|
|||||||
|
# ComfyUI stack tunables. Copy to `.env` on irv-ml1 before deploying.
|
||||||
|
#
|
||||||
|
# cp .env.example .env
|
||||||
|
# # edit as needed
|
||||||
|
# docker compose up -d
|
||||||
|
|
||||||
|
# Image tag. Format is `ubuntu<os>_cuda<ver>-<date>`; the floating
|
||||||
|
# equivalent is `ubuntu<os>_cuda<ver>-latest`.
|
||||||
|
#
|
||||||
|
# irv-ml1 runtime constraints:
|
||||||
|
# - Host driver 570.x → max runtime CUDA 12.8 (per nvidia-smi). Tags
|
||||||
|
# above 12.8 will fail at runtime on driver/CUDA mismatch even
|
||||||
|
# though the manifest pulls fine.
|
||||||
|
# - Host cuda-toolkit is 12.8 to match native Python venvs (ComfyUI,
|
||||||
|
# SDNext, fluxgym, llama.cpp, etc.) — keep the container in sync.
|
||||||
|
#
|
||||||
|
# See https://hub.docker.com/r/mmartial/comfyui-nvidia-docker/tags.
|
||||||
|
COMFYUI_VERSION=ubuntu24_cuda12.8-20260312
|
||||||
|
|
||||||
|
# Host port for the ComfyUI web UI (container listens on 8188)
|
||||||
|
COMFYUI_PORT=8188
|
||||||
|
|
||||||
|
# Bind address. 0.0.0.0 exposes on all interfaces including the WG
|
||||||
|
# tunnel IP (10.100.79.3) so homepage / the hub side can reach it.
|
||||||
|
# Use 127.0.0.1 to restrict to local / SSH-forwarded access only.
|
||||||
|
COMFYUI_BIND=0.0.0.0
|
||||||
|
|
||||||
|
# UID/GID the container user runs as. Must match the owner of
|
||||||
|
# COMFYUI_BASEDIR so workflows, custom nodes, input, and output are
|
||||||
|
# read/write from both the container and external host tooling.
|
||||||
|
# On irv-ml1: lkraven = 1000:1000.
|
||||||
|
COMFYUI_UID=1000
|
||||||
|
COMFYUI_GID=1000
|
||||||
|
|
||||||
|
# Host path for all user state (models, workflows, custom_nodes,
|
||||||
|
# input, output). Must exist and be chowned to COMFYUI_UID:COMFYUI_GID
|
||||||
|
# before the first `docker compose up`.
|
||||||
|
COMFYUI_BASEDIR=/worktank/comfyui/basedir
|
||||||
|
|
||||||
|
# Host path for container runtime (ComfyUI source + Python venv +
|
||||||
|
# pip cache, ~5 GB after bootstrap). Disposable — safe to `rm -rf`
|
||||||
|
# on version bumps to force a fresh re-bootstrap. Must be owned by
|
||||||
|
# COMFYUI_UID:COMFYUI_GID.
|
||||||
|
COMFYUI_RUNDIR=/worktank/comfyui/run
|
||||||
|
|
||||||
|
# ComfyUI-Manager security level:
|
||||||
|
# normal → default; installs custom nodes from vetted sources only.
|
||||||
|
# weak → permits arbitrary pip installs via Manager. Use only if
|
||||||
|
# you trust every workflow author you'll import.
|
||||||
|
# strict → blocks pip installs via Manager entirely.
|
||||||
|
COMFYUI_SECURITY_LEVEL=normal
|
||||||
@@ -0,0 +1,121 @@
|
|||||||
|
# ComfyUI
|
||||||
|
|
||||||
|
Node-based Stable Diffusion / Flux inference UI.
|
||||||
|
|
||||||
|
**Server:** irv-ml1 (Irvine, WireGuard-only)
|
||||||
|
**Port:** 8188 (configurable via `.env`)
|
||||||
|
**GPUs:** both (RTX 3090 + RTX A6000 both exposed; select per-workflow)
|
||||||
|
**Image:** `mmartial/comfyui-nvidia-docker`
|
||||||
|
**Native coexistence:** runs independently of `/opt/ComfyUI` (the existing native install). Migrate models by copy or move — no cross-binding.
|
||||||
|
|
||||||
|
## Path layout
|
||||||
|
|
||||||
|
All user state lives under a single host tree on `/worktank`, owned
|
||||||
|
by `lkraven:lkraven` (UID/GID 1000) so external tooling can read and
|
||||||
|
write workflow files directly. ComfyUI-Manager puts downloaded weights
|
||||||
|
into the same tree, so everything is visible on the host.
|
||||||
|
|
||||||
|
| Host path | Container path | Purpose | Restic? |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `/worktank/comfyui/basedir/models/` | `/basedir/models/` | Checkpoints, loras, vae, unet, controlnet, upscale, etc. | excluded (bulk, regenerable) |
|
||||||
|
| `/worktank/comfyui/basedir/user/default/workflows/` | `/basedir/user/default/workflows/` | Workflow JSON files — **external tools edit here** | included |
|
||||||
|
| `/worktank/comfyui/basedir/custom_nodes/` | `/basedir/custom_nodes/` | ComfyUI-Manager-installed nodes | included |
|
||||||
|
| `/worktank/comfyui/basedir/input/` | `/basedir/input/` | Uploaded source images | included |
|
||||||
|
| `/worktank/comfyui/basedir/output/` | `/basedir/output/` | Generated images | excluded (regenerable) |
|
||||||
|
| `/worktank/comfyui/run/` | `/comfy/mnt` | ComfyUI source + venv + pip cache (~5 GB) | excluded (disposable) |
|
||||||
|
|
||||||
|
Standard ComfyUI model subdirs the image creates under `models/`:
|
||||||
|
`checkpoints`, `clip`, `clip_vision`, `configs`, `controlnet`,
|
||||||
|
`diffusers`, `embeddings`, `gligen`, `hypernetworks`, `loras`,
|
||||||
|
`photomaker`, `style_models`, `unet`, `upscale_models`, `vae`,
|
||||||
|
`vae_approx`.
|
||||||
|
|
||||||
|
## First-time deploy on irv-ml1
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Push compose + env template from workstation
|
||||||
|
scripts/deploy-stack.sh irv-ml1 comfyui
|
||||||
|
|
||||||
|
# 2. Create basedir + run on the host (the image expects both to exist
|
||||||
|
# as the target UID/GID before first up; it refuses to chown
|
||||||
|
# mounted paths). One-time sudo because /worktank is root-owned.
|
||||||
|
ssh -t irv-ml1 'sudo mkdir -p /worktank/comfyui/{basedir,run} && \
|
||||||
|
sudo chown -R lkraven:lkraven /worktank/comfyui'
|
||||||
|
|
||||||
|
# 3. Create .env and bring up. First `up` takes several minutes —
|
||||||
|
# the image installs ~5 GB of Python packages on first boot.
|
||||||
|
ssh irv-ml1 '
|
||||||
|
cd /opt/docker/compose/comfyui && \
|
||||||
|
cp -n .env.example .env && \
|
||||||
|
docker compose config >/dev/null && \
|
||||||
|
docker compose up -d && \
|
||||||
|
docker compose logs -f --tail=30
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
Once the logs settle on "Starting server" / "To see the GUI go to …",
|
||||||
|
the UI is reachable at `http://10.100.79.3:8188` (from the WG-tunnel
|
||||||
|
side). The homepage hub on ana-docker auto-discovers it via the Dozzle
|
||||||
|
agent + the homepage Docker socket integration already wired on
|
||||||
|
irv-ml1.
|
||||||
|
|
||||||
|
## Populate models
|
||||||
|
|
||||||
|
Native install lives at `/opt/ComfyUI/models/` (owned `llmuser`).
|
||||||
|
Copy or move into the container's tree — paths match 1:1:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Example: copy one subdir at a time, preserving timestamps
|
||||||
|
ssh -t irv-ml1 'sudo -u lkraven rsync -av \
|
||||||
|
/opt/ComfyUI/models/checkpoints/ \
|
||||||
|
/worktank/comfyui/basedir/models/checkpoints/'
|
||||||
|
```
|
||||||
|
|
||||||
|
Repeat per subdir. `rsync -av --remove-source-files` if you want a
|
||||||
|
true move (then clean up empty source dirs). The native install will
|
||||||
|
continue to work with whatever remains under `/opt/ComfyUI/models/`.
|
||||||
|
|
||||||
|
## Editing workflows from outside the container
|
||||||
|
|
||||||
|
Workflows live as JSON under
|
||||||
|
`/worktank/comfyui/basedir/user/default/workflows/` on the host,
|
||||||
|
owned `lkraven:lkraven`. External tools running as `lkraven` can
|
||||||
|
read and write them without coordination — ComfyUI picks up file
|
||||||
|
changes on the next workflow load (no container restart needed).
|
||||||
|
|
||||||
|
## Deploy updates
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# After editing compose.yaml or .env.example here
|
||||||
|
scripts/deploy-stack.sh irv-ml1 comfyui
|
||||||
|
ssh irv-ml1 'cd /opt/docker/compose/comfyui && docker compose up -d'
|
||||||
|
```
|
||||||
|
|
||||||
|
## Image version bump
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Pick a new tag from
|
||||||
|
# https://hub.docker.com/r/mmartial/comfyui-nvidia-docker/tags
|
||||||
|
ssh -t irv-ml1 '
|
||||||
|
cd /opt/docker/compose/comfyui && \
|
||||||
|
sed -i "s/^COMFYUI_VERSION=.*/COMFYUI_VERSION=<new-tag>/" .env && \
|
||||||
|
docker compose pull && \
|
||||||
|
docker compose up -d
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
Runtime lives at `/worktank/comfyui/run/`. For a fresh bootstrap
|
||||||
|
(force venv + pip cache rebuild against the new image):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh irv-ml1 '
|
||||||
|
cd /opt/docker/compose/comfyui && \
|
||||||
|
docker compose down && \
|
||||||
|
rm -rf /worktank/comfyui/run/* && \
|
||||||
|
docker compose pull && \
|
||||||
|
docker compose up -d
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
User state under `/worktank/comfyui/basedir/` is untouched by either
|
||||||
|
pattern.
|
||||||
@@ -0,0 +1,63 @@
|
|||||||
|
# ComfyUI — node-based Stable Diffusion / Flux inference UI.
|
||||||
|
#
|
||||||
|
# Runs on irv-ml1 (dual GPU: RTX 3090 + RTX A6000). Both GPUs are
|
||||||
|
# exposed via NVIDIA_VISIBLE_DEVICES=all so GPU selection happens
|
||||||
|
# per-workflow inside ComfyUI.
|
||||||
|
#
|
||||||
|
# All user state — models, workflows, custom_nodes, input, output —
|
||||||
|
# lives under a single BASE_DIRECTORY tree on /worktank (462 GB
|
||||||
|
# dedicated), owned by lkraven:lkraven (1000:1000) so external
|
||||||
|
# tooling can read and write workflow files directly on the host.
|
||||||
|
#
|
||||||
|
# Runtime state (ComfyUI source, venv, pip cache) lives in a bind
|
||||||
|
# mount at ${COMFYUI_RUNDIR} — disposable (can be wiped on version
|
||||||
|
# bumps to force re-bootstrap), but owned by the host user so no
|
||||||
|
# sudo dance is needed. (Named volumes would be created root-owned
|
||||||
|
# and the image refuses to chown a mounted path.)
|
||||||
|
#
|
||||||
|
# First-run prerequisite: both ${COMFYUI_BASEDIR} and ${COMFYUI_RUNDIR}
|
||||||
|
# must exist on the host with ownership matching COMFYUI_UID:COMFYUI_GID
|
||||||
|
# before `up`. See README for the bootstrap command.
|
||||||
|
#
|
||||||
|
# All tunables live in .env — edit that, not this file.
|
||||||
|
|
||||||
|
services:
|
||||||
|
comfyui:
|
||||||
|
image: mmartial/comfyui-nvidia-docker:${COMFYUI_VERSION}
|
||||||
|
container_name: comfyui
|
||||||
|
restart: unless-stopped
|
||||||
|
runtime: nvidia
|
||||||
|
ports:
|
||||||
|
- "${COMFYUI_BIND:-0.0.0.0}:${COMFYUI_PORT}:8188"
|
||||||
|
environment:
|
||||||
|
- NVIDIA_VISIBLE_DEVICES=all
|
||||||
|
- WANTED_UID=${COMFYUI_UID}
|
||||||
|
- WANTED_GID=${COMFYUI_GID}
|
||||||
|
- BASE_DIRECTORY=/basedir
|
||||||
|
- SECURITY_LEVEL=${COMFYUI_SECURITY_LEVEL:-normal}
|
||||||
|
- USE_UV=true
|
||||||
|
volumes:
|
||||||
|
- ${COMFYUI_BASEDIR}:/basedir
|
||||||
|
- ${COMFYUI_RUNDIR}:/comfy/mnt
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8188/ >/dev/null || exit 1"]
|
||||||
|
interval: 30s
|
||||||
|
timeout: 10s
|
||||||
|
retries: 3
|
||||||
|
# First boot installs ~5 GB of Python packages; allow generous
|
||||||
|
# start_period so the container isn't marked unhealthy during
|
||||||
|
# bootstrap. Subsequent starts are fast.
|
||||||
|
start_period: 600s
|
||||||
|
networks:
|
||||||
|
- tnet
|
||||||
|
labels:
|
||||||
|
- homepage.group=AI Systems
|
||||||
|
- homepage.name=ComfyUI
|
||||||
|
- homepage.icon=mdi-image-auto-adjust
|
||||||
|
- homepage.description=Node-based SD/Flux inference (irv-ml1)
|
||||||
|
- homepage.href=http://10.100.79.3:${COMFYUI_PORT}
|
||||||
|
|
||||||
|
networks:
|
||||||
|
tnet:
|
||||||
|
name: traefik-net
|
||||||
|
external: true
|
||||||
Reference in New Issue
Block a user