The 2026-09-06 headscale cutover retired irv-ml1's wg0 tunnel IP 10.100.79.3 (now 10.6.110.50). Repointed all LIVE canonical refs to the DNS NAME so the next move can't re-break them: homepage.href/siteMonitor labels across 25 stack composes, load-bearing env defaults (asset-engine INFERENCE_HOST, open-webui AUDIO_TTS_OPENAI_API_BASE_URL, skaldsong SKALDSONG_TTS_BASE_URL, zonos-gateway ZONOS_URL, dia), homepage services.yaml manual cards (Voice Design Studio, IRV-ML1), and servers/irv-ml1/ssh-target. Updated the stale 'WG tunnel' comment to the mesh reality. Left as-is: README curl-examples and .env.example comments (docs), and historical mentions in CLAUDE.md/persistent-memory. NOTE: applying the label repoints to the RUNNING irv-ml1 containers needs a recreate per service (labels read at creation); deployed .env values are separate from these canonical defaults.
ComfyUI
Node-based Stable Diffusion / Flux inference UI.
Server: irv-ml1 (Irvine, WireGuard-only)
Port: 8188 (configurable via .env)
GPUs: both (RTX 3090 + RTX A6000 both exposed; select per-workflow)
Image: mmartial/comfyui-nvidia-docker
Native coexistence: runs independently of /opt/ComfyUI (the existing native install). Migrate models by copy or move — no cross-binding.
Path layout
All user state lives under a single host tree on /worktank, owned
by lkraven:lkraven (UID/GID 1000) so external tooling can read and
write workflow files directly. ComfyUI-Manager puts downloaded weights
into the same tree, so everything is visible on the host.
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
/worktank/comfyui/basedir/models/ |
/basedir/models/ |
Checkpoints, loras, vae, unet, controlnet, upscale, etc. | excluded (bulk, regenerable) |
/worktank/comfyui/basedir/user/default/workflows/ |
/basedir/user/default/workflows/ |
Workflow JSON files — external tools edit here | included |
/worktank/comfyui/basedir/custom_nodes/ |
/basedir/custom_nodes/ |
ComfyUI-Manager-installed nodes | included |
/worktank/comfyui/basedir/input/ |
/basedir/input/ |
Uploaded source images | included |
/worktank/comfyui/basedir/output/ |
/basedir/output/ |
Generated images | excluded (regenerable) |
/worktank/comfyui/run/ |
/comfy/mnt |
ComfyUI source + venv + pip cache (~5 GB) | excluded (disposable) |
Standard ComfyUI model subdirs the image creates under models/:
checkpoints, clip, clip_vision, configs, controlnet,
diffusers, embeddings, gligen, hypernetworks, loras,
photomaker, style_models, unet, upscale_models, vae,
vae_approx.
First-time deploy on irv-ml1
# 1. Push compose + env template from workstation
scripts/deploy-stack.sh irv-ml1 comfyui
# 2. Create basedir + run on the host (the image expects both to exist
# as the target UID/GID before first up; it refuses to chown
# mounted paths). One-time sudo because /worktank is root-owned.
ssh -t irv-ml1 'sudo mkdir -p /worktank/comfyui/{basedir,run} && \
sudo chown -R lkraven:lkraven /worktank/comfyui'
# 3. Create .env and bring up. First `up` takes several minutes —
# the image installs ~5 GB of Python packages on first boot.
ssh irv-ml1 '
cd /opt/docker/compose/comfyui && \
cp -n .env.example .env && \
docker compose config >/dev/null && \
docker compose up -d && \
docker compose logs -f --tail=30
'
Once the logs settle on "Starting server" / "To see the GUI go to …",
the UI is reachable at http://10.100.79.3:8188 (from the WG-tunnel
side). The homepage hub on ana-docker auto-discovers it via the Dozzle
agent + the homepage Docker socket integration already wired on
irv-ml1.
Populate models
Native install lives at /opt/ComfyUI/models/ (owned llmuser).
Copy or move into the container's tree — paths match 1:1:
# Example: copy one subdir at a time, preserving timestamps
ssh -t irv-ml1 'sudo -u lkraven rsync -av \
/opt/ComfyUI/models/checkpoints/ \
/worktank/comfyui/basedir/models/checkpoints/'
Repeat per subdir. rsync -av --remove-source-files if you want a
true move (then clean up empty source dirs). The native install will
continue to work with whatever remains under /opt/ComfyUI/models/.
Editing workflows from outside the container
Workflows live as JSON under
/worktank/comfyui/basedir/user/default/workflows/ on the host,
owned lkraven:lkraven. External tools running as lkraven can
read and write them without coordination — ComfyUI picks up file
changes on the next workflow load (no container restart needed).
Deploy updates
# After editing compose.yaml or .env.example here
scripts/deploy-stack.sh irv-ml1 comfyui
ssh irv-ml1 'cd /opt/docker/compose/comfyui && docker compose up -d'
Image version bump
# Pick a new tag from
# https://hub.docker.com/r/mmartial/comfyui-nvidia-docker/tags
ssh -t irv-ml1 '
cd /opt/docker/compose/comfyui && \
sed -i "s/^COMFYUI_VERSION=.*/COMFYUI_VERSION=<new-tag>/" .env && \
docker compose pull && \
docker compose up -d
'
Runtime lives at /worktank/comfyui/run/. For a fresh bootstrap
(force venv + pip cache rebuild against the new image):
ssh irv-ml1 '
cd /opt/docker/compose/comfyui && \
docker compose down && \
rm -rf /worktank/comfyui/run/* && \
docker compose pull && \
docker compose up -d
'
User state under /worktank/comfyui/basedir/ is untouched by either
pattern. But a run/* wipe drops any pip deps that live only in the
venv — see RTX VSR below; re-run its pip step after every fresh bootstrap.
Custom nodes needing extra provisioning — RTX Video Super Resolution
The RTX Video Super Resolution node (Nvidia_RTX_Nodes_ComfyUI, Comfy-Org)
needs the NVIDIA MAXINE Video Effects SDK bindings (nvidia-vfx, import name
nvvfx), served from NVIDIA's package index, not public PyPI. Two pieces with
different durability:
| Piece | Lives in | Survives a run/* venv wipe? |
|---|---|---|
the node Nvidia_RTX_Nodes_ComfyUI |
basedir/custom_nodes/ (persistent, restic-included) |
yes |
the nvidia-vfx wheel |
the venv under run/ (disposable, restic-excluded) |
NO — re-install it |
So the node itself is durable, but any venv rebuild (rm -rf run/*, the
fresh-bootstrap on an image bump) drops nvidia-vfx and the node breaks on
import nvvfx until it's reinstalled. irv-ml1 already meets the node's
requirements (A6000 Ampere+ Tensor cores; driver 595.58 ≥ 590.44; container OS
fine).
Provision from scratch — both commands run as UID 1000; a root / -u 0
install corrupts venv ownership and crash-loops the boot torch step (see the
chown -R 1000:1000 venv lesson):
# 1. The node — only if basedir/custom_nodes was wiped (normally persistent):
ssh irv-ml1 'git clone https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI \
/worktank/comfyui/basedir/custom_nodes/Nvidia_RTX_Nodes_ComfyUI'
# 2. The pip dep — RE-RUN AFTER EVERY venv rebuild (it lives in the disposable venv):
ssh irv-ml1 'docker exec -u 1000:1000 comfyui \
/comfy/mnt/venv/bin/pip install nvidia-vfx --extra-index-url https://pypi.nvidia.com'
# Verify:
ssh irv-ml1 'docker exec -u 1000:1000 comfyui \
/comfy/mnt/venv/bin/python -c "import nvvfx; print(\"nvvfx OK\")"'
The --extra-index-url is kept scoped to this one install, deliberately NOT
promoted to a global PIP_EXTRA_INDEX_URL in compose: a fleet-wide extra index
could perturb the delicate pinned torch/cuda + SageAttention resolution the boot
bootstrap depends on (DISABLE_UPGRADES=true, torch 2.12.1). The node's own
requirements.txt lists nvidia-vfx, so if ComfyUI ever auto-installs node
requirements it will need this index reachable — but scoping it to the manual
step keeps the boot bootstrap untouched.