From 7ae7193c2167977d89223c8e14176e5072e44856 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Sun, 27 Sep 2026 13:57:35 -0700 Subject: [PATCH] feat(blender): Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand (Prime) Uses linuxserver/blender (Selkies Wayland desktop, NVENC), digest-pinned. The container sees only GPU 3, has restart "no", and runs only while in use, because GPU 3 is the reserve card for a full-size vLLM seat. Web desktop on :3001 with basic auth (vault fv-ml1/blender-web-password). /work is on /tank and is not backed up; /config lives under /opt/docker (restic). Acceptance: Cycles finds the card on OptiX and CUDA (sm_120 kernels ship in the build). The heavy self-test renders in 3.54 s on OptiX vs 22.71 s on CPU, a functional check with n=1. Web auth answers 401 without credentials and with a wrong password, and 200 with the right one. --- servers/fv-ml1/README.md | 4 +++ stacks/blender/.env.example | 6 +++++ stacks/blender/README.md | 51 +++++++++++++++++++++++++++++++++++++ stacks/blender/compose.yaml | 50 ++++++++++++++++++++++++++++++++++++ 4 files changed, 111 insertions(+) create mode 100644 stacks/blender/.env.example create mode 100644 stacks/blender/README.md create mode 100644 stacks/blender/compose.yaml diff --git a/servers/fv-ml1/README.md b/servers/fv-ml1/README.md index a411926..566517e 100644 --- a/servers/fv-ml1/README.md +++ b/servers/fv-ml1/README.md @@ -224,6 +224,10 @@ eats into a future big seat's profiling margin. Small seats go on GPU 0, which h the most uncommitted headroom (its seats commit util 0.88; GPU 1 is at 0.975 and GPU 2 at 0.96). +**GPU 3 on-demand tenant (Prime, 2026-09-27): `blender`** (`stacks/blender/`). It is up only +while in use (`restart: "no"`, about 270 MiB when idle with the desktop running, 0 when down). +The reserve still stands: whenever a full-size seat takes GPU 3, Blender stays down. + **Retired:** - `llama-swap` (former GGUF multiplexer on :9292) — replaced by dedicated per-model seats (e.g. `llama-charrp`); no longer running. diff --git a/stacks/blender/.env.example b/stacks/blender/.env.example new file mode 100644 index 0000000..5de9fc9 --- /dev/null +++ b/stacks/blender/.env.example @@ -0,0 +1,6 @@ +# blender — copy to /opt/docker/compose/blender/.env on fv-ml1 (mode 0600). +IMAGE=lscr.io/linuxserver/blender:5.2.2-ls241@sha256:9216c77ab2bf38a758390f802c5777554df605c53456dbcad828509a27a8f28b +HTTPS_PORT=3001 +# Basic auth for the web desktop. PASSWORD source of truth: secret get fv-ml1/blender-web-password +CUSTOM_USER=blender +PASSWORD= diff --git a/stacks/blender/README.md b/stacks/blender/README.md new file mode 100644 index 0000000..f518366 --- /dev/null +++ b/stacks/blender/README.md @@ -0,0 +1,51 @@ +# blender + +**Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand.** Agents drive it; there is also a browser +desktop to watch it or take over. Prime, 2026-09-27: "go ahead with gpu 3, both". He does not +use Blender himself, so the agent side (MCP) is the primary interface. + +| | | +|---|---| +| **Desktop** | `https://10.251.50.54:3001` (self-signed cert). Basic auth: user `blender`, password `secret get fv-ml1/blender-web-password`. | +| **Image** | `lscr.io/linuxserver/blender:5.2.2-ls241@sha256:9216c77a…` (Selkies 2.0 Wayland desktop, NVENC stream). | +| **GPU** | GPU 3 only (`NVIDIA_VISIBLE_DEVICES=3`). About 270 MiB is held while the desktop runs. | +| **Files** | `/work` → `/tank/blender` (projects, assets, renders; **not backed up**). `/config` → `/opt/docker/data/blender` (prefs, add-ons; restic). Files are owned by infra-ops (uid 1002), so agents can `scp` in and out. | + +## ⚠ On demand: GPU 3 is borrowed + +GPU 3 is the fleet's reserve card for a full-size vLLM seat (`servers/fv-ml1/README.md`). +Blender uses it only while in use: + +```bash +ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose up -d' # start +ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose down' # stop, card back to 0 +``` + +`restart: "no"`, so a reboot never brings it back. **When a big seat moves onto GPU 3, Blender +stays down.** + +## Headless rendering (no desktop needed) + +```bash +ssh infra-ops@10.251.50.54 'docker exec -u abc blender blender -b /work/.blend -E CYCLES -o /work/out/frame_#### -a -- --cycles-device OPTIX' +``` + +Run as `-u abc`, the image's user, mapped to uid 1002, so outputs land owned by infra-ops. + +## Acceptance (2026-09-27, 1356) + +- Cycles sees the card on both OptiX and CUDA: "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation + Edition". The build ships `kernel_sm_120.cubin` plus OptiX PTX. +- Self-test (`/tank/blender/render_test.py`): a subdivided glass monkey, 1920×1080, 1024 samples, + 32 bounces, no denoise. **OptiX 3.54 s against CPU 22.71 s (96 threads).** That is n=1 per + device: a functional check that the GPU is really used, not a benchmark. A trivial default-cube + scene could not separate them (0.59 s against 0.65 s), which is why the self-test scene is heavy. +- Web auth: no credentials gives 401, a wrong password 401, the right one 200. +- Selkies: "Render node 1 encodes H264, AV1, H265 on nvenc"; the Wayland renderer runs GL on GPU 3. + +## Agent control (MCP) + +Being researched with dvalin-smithy-dev: which Blender MCP server, and how it reaches Blender +across hosts. This section fills in once the choice is made. ⚠ Most Blender MCP add-ons expose +"run arbitrary Python inside Blender" on a TCP port. Treat that port like a shell: bind it to +the container or the LAN only, never publish it wider. diff --git a/stacks/blender/compose.yaml b/stacks/blender/compose.yaml new file mode 100644 index 0000000..742a38d --- /dev/null +++ b/stacks/blender/compose.yaml @@ -0,0 +1,50 @@ +# blender: Blender 5.2 LTS on fv-ml1 GPU 3, driven by agents (MCP) with a browser desktop for +# watching. Prime, 2026-09-27: "go ahead with gpu 3, both". +# +# ⚠ ON DEMAND ONLY. GPU 3 is the fleet's deliberately empty card, the reserve for a full-size +# vLLM seat (flash-next needs 93 of 96 GiB, and vLLM sizes KV against TOTAL VRAM). Blender +# borrows it: `restart: "no"`, so a reboot never brings it back, and it is stopped whenever a big +# seat needs the card. Start: `docker compose up -d`. Stop: `docker compose down`. +# +# Image: linuxserver/blender (Selkies web desktop, official Blender build with sm_120 CUDA + +# OptiX kernels). Its NVIDIA mode needs host driver >= 580 (fv-ml1: 580.65.06) and +# /dev/nvidia-modeset. HTTPS only (Selkies' WebCodecs need a secure context), self-signed cert. +# Basic auth from .env (CUSTOM_USER / PASSWORD; vault fv-ml1/blender-web). The desktop grants +# a shell inside the container, so never expose it beyond the LAN. +# +# Paths: /config (home: prefs, add-ons) → /opt/docker/data/blender (restic via /opt/docker). +# /work (projects, assets, renders) → /tank/blender (big, NOT backed up; renders are +# regenerable, and anything precious is copied out). +# +# .env (tunables): IMAGE, HTTPS_PORT, CUSTOM_USER, PASSWORD. + +name: blender + +services: + blender: + image: ${IMAGE:?set IMAGE} + container_name: blender + restart: "no" + runtime: nvidia + environment: + PUID: "1002" # infra-ops, so agents can read and write /work over ssh + PGID: "1003" + TZ: America/Los_Angeles + CUSTOM_USER: ${CUSTOM_USER:?set CUSTOM_USER} + PASSWORD: ${PASSWORD:?set PASSWORD} + NVIDIA_VISIBLE_DEVICES: "3" # GPU 3 only: never touch the serving cards + NVIDIA_DRIVER_CAPABILITIES: all # compute (Cycles), graphics (viewport), video (NVENC stream) + devices: + - /dev/nvidia-modeset:/dev/nvidia-modeset + volumes: + - /opt/docker/data/blender:/config + - /tank/blender:/work + ports: + - "${HTTPS_PORT:-3001}:3001" + shm_size: "1gb" + labels: + - homepage.group=AI - Studios + - homepage.name=Blender + - homepage.icon=si-blender + - homepage.description=Blender 5.2 on fv-ml1 GPU 3 (on demand) + - homepage.href=https://10.251.50.54:${HTTPS_PORT:-3001}