# ComfyUI — node-based Stable Diffusion / Flux inference UI. # # Runs on irv-ml1 (dual GPU: RTX 3090 + RTX A6000). PINNED to the A6000 # (device 1) via NVIDIA_VISIBLE_DEVICES=1 — the 3090 hosts the audio/TTS # zoo (chatterbox, parakeet, vibevoice, ytvc, kokoro) so ComfyUI gets the # full 48 GB A6000 to itself (operator consolidation 2026-06-18). # # All user state — models, workflows, custom_nodes, input, output — # lives under a single BASE_DIRECTORY tree on /worktank (462 GB # dedicated), owned by lkraven:lkraven (1000:1000) so external # tooling can read and write workflow files directly on the host. # # Runtime state (ComfyUI source, venv, pip cache) lives in a bind # mount at ${COMFYUI_RUNDIR} — disposable (can be wiped on version # bumps to force re-bootstrap), but owned by the host user so no # sudo dance is needed. (Named volumes would be created root-owned # and the image refuses to chown a mounted path.) # # First-run prerequisite: both ${COMFYUI_BASEDIR} and ${COMFYUI_RUNDIR} # must exist on the host with ownership matching COMFYUI_UID:COMFYUI_GID # before `up`. See README for the bootstrap command. # # All tunables live in .env — edit that, not this file. services: comfyui: image: mmartial/comfyui-nvidia-docker:${COMFYUI_VERSION} container_name: comfyui restart: unless-stopped runtime: nvidia ports: - "${COMFYUI_BIND:-0.0.0.0}:${COMFYUI_PORT}:8188" environment: - NVIDIA_VISIBLE_DEVICES=1 # Pin torch at the current 2.12.1+cu129 so the boot script stops # auto-upgrading it — compiled SageAttention kernels must not drift # (comfy-dev torch-pin, operator-approved 2026-06-18). - DISABLE_UPGRADES=true - WANTED_UID=${COMFYUI_UID} - WANTED_GID=${COMFYUI_GID} - BASE_DIRECTORY=/basedir - SECURITY_LEVEL=${COMFYUI_SECURITY_LEVEL:-normal} - USE_UV=true # Extra ComfyUI launch flags (image appends these to main.py, then adds # --base-directory + --enable-manager itself): # --fp8_e4m3fn-text-enc — load the FLUX.2 Qwen3-8B text encoder as fp8 # (~8.7 GB) instead of upcasting the fp8 file to fp16 (~16 GB). Matches # the box's Ampere-fp8 posture; the encoder runs once per gen so the # upcast-on-compute cost is negligible. # --use-sage-attention — 0.24.1's NATIVE attention selection (the node-based # BlehGlobalSageAttention is dead on 0.24.1: "does not support the new # ComfyUI attention changes"). Binds the in-image sageattention v2.2.0 # sm_86 build (rebuilt against the pinned torch 2.12.1). Global speedup # across Flux/SDXL/Wan (comfy-dev benchmarking, 2026-06-18). # # ALLOCATOR (2026-07-19 — comfy-dev A/B, operator-run — cudaMallocAsync WON). # We deliberately DO NOT pass --disable-cuda-malloc, so ComfyUI keeps CUDA's # default async allocator (cudaMallocAsync). History: --disable-cuda-malloc + # PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True were added when the A6000 was # SHARED with the TTS zoo (the native allocator dodged a cudaMallocAsync # phantom-OOM). TTS moved to the 3090 (2026-06-18), removing that trigger; the # LTX-2.3 v1.5.0 LoRA stack (DMD+OmniNFT+act LoRAs patching the DiT + the 12B # Gemma text-encoder) then began hitting the 48 GB ceiling under the native # allocator, which fragments/over-reserves (~45 GB allocated+reserved-but- # unallocated) and OOMs at TE load. cudaMallocAsync packs tighter + promptly # returns freed blocks, so the same job now FITS: the operator's previously- # OOMing stress test peaks ~82% VRAM (~40/48 GB) with headroom, 0 OOM/errors. # COUPLING: PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True is native-allocator- # only, so it is REMOVED here (inert/invalid under cudaMallocAsync) — revert BOTH # together. If the TTS zoo ever moves back onto the A6000, re-evaluate the pair. - COMFY_CMDLINE_EXTRA=--fp8_e4m3fn-text-enc --use-sage-attention volumes: - ${COMFYUI_BASEDIR}:/basedir # models/ overlaid from storetank. The ~325 GB model tree was migrated # off the near-full worktank NVMe (2026-06-13) to /storetank/arbo (roomy # SATA SSD). This nested mount shadows the models subdir of /basedir; # everything else (custom_nodes, output, input, user, workflows) stays on # worktank. Inventory: docs/arbo-comfyui-model-catalog.md. - ${COMFYUI_MODELS_DIR:-/storetank/arbo/models}:/basedir/models - ${COMFYUI_RUNDIR}:/comfy/mnt healthcheck: test: ["CMD-SHELL", "curl -fsS http://localhost:8188/ >/dev/null || exit 1"] interval: 30s timeout: 10s retries: 3 # First boot installs ~5 GB of Python packages; allow generous # start_period so the container isn't marked unhealthy during # bootstrap. Subsequent starts are fast. start_period: 600s networks: - tnet labels: - homepage.group=AI - Image & Media - homepage.name=ComfyUI - homepage.icon=mdi-image-auto-adjust - homepage.description=Node-based SD/Flux inference (irv-ml1) - homepage.href=http://10.100.79.3:${COMFYUI_PORT} networks: tnet: name: traefik-net external: true