Files
esh-pfi-infrastructure/stacks/blender/README.md
T

6.0 KiB
Raw Blame History

blender

Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand. Agents drive it; there is also a browser desktop to watch it or take over. Prime, 2026-09-27: "go ahead with gpu 3, both". He does not use Blender himself, so the agent side (MCP) is the primary interface.

Desktop https://10.251.50.54:3001 (self-signed cert). Basic auth: user blender, password secret get fv-ml1/blender-web-password.
Image lscr.io/linuxserver/blender:5.2.2-ls241@sha256:9216c77a… (Selkies 2.0 Wayland desktop, NVENC stream).
GPU GPU 3 only (NVIDIA_VISIBLE_DEVICES=3). About 270 MiB is held while the desktop runs.
Files /work → /tank/blender (projects, assets, renders; not backed up). /config → /opt/docker/data/blender (prefs, add-ons; restic). Files are owned by infra-ops (uid 1002), so agents can scp in and out.

⚠ On demand: GPU 3 is borrowed

GPU 3 is the fleet's reserve card for a full-size vLLM seat (servers/fv-ml1/README.md). Blender uses it only while in use:

ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose up -d'   # start
ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose down'    # stop, card back to 0

restart: "no", so a reboot never brings it back. When a big seat moves onto GPU 3, Blender stays down.

Headless rendering (no desktop needed)

ssh infra-ops@10.251.50.54 'docker exec -u abc blender blender -b /work/<file>.blend -E CYCLES -o /work/out/frame_#### -a -- --cycles-device OPTIX'

Run as -u abc, the image's user, mapped to uid 1002, so outputs land owned by infra-ops.

Acceptance (2026-09-27, 1356)

  • Cycles sees the card on both OptiX and CUDA: "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition". The build ships kernel_sm_120.cubin plus OptiX PTX.
  • Self-test (/tank/blender/render_test.py): a subdivided glass monkey, 1920×1080, 1024 samples, 32 bounces, no denoise. OptiX 3.54 s against CPU 22.71 s (96 threads). That is n=1 per device: a functional check that the GPU is really used, not a benchmark. A trivial default-cube scene could not separate them (0.59 s against 0.65 s), which is why the self-test scene is heavy.
  • Web auth: no credentials gives 401, a wrong password 401, the right one 200.
  • Selkies: "Render node 1 encodes H264, AV1, H265 on nvenc"; the Wayland renderer runs GL on GPU 3.

Agent control (MCP): scripts/blender-mcp

The chosen server is mcp-for-blender (MIT, one maintainer, ~29k stars; researched by dvalin-smithy-dev 2026-09-27, thread 01M3JA61FTFW2ZPSD20MHF7RPJ, full note in dvalin-smithy research/blender-agent-drive-2026-09-27.md). Two halves:

  • The add-on is inside the running GUI Blender and serves a socket that executes arbitrary Python with no authentication. It is vendored at upstream commit 41a18432 (conf/scripts/addons/blender_mcp.py, MIT licence alongside) and started by conf/scripts/startup/fleet_mcp.py. The add-on only serves from a GUI Blender, never from blender -b, which is one reason the desktop exists.
  • The MCP server: mcp-for-blender==2.1.1, frozen in conf/mcp-requirements.txt and installed in /work/.mcp-venv inside the container (scripts/blender-mcp setup). Agents reach it as stdio over ssh infra-ops@fv-ml1 docker exec -i. It runs in the container for two reasons:
    1. No port is published. The socket stays on the container's localhost, so access means ssh + docker on fv-ml1.
    2. Viewport screenshots need a shared filesystem. Blender writes the image and the server reads it back. With the server on nh3-dev it failed ("Screenshot file was not created").

Settings: DISABLE_TELEMETRY=true and BLENDER_MCP_SAFE_MODE=1. Safe mode puts upstream's AST allowlist in front of execute_blender_code: bpy, bmesh, mathutils and pure stdlib only, and no os, open, eval or network. It guards against prompt injection from third-party asset text; it is not a sandbox.

For an agent

scripts/blender-mcp up          # GPU 3 is borrowed: start only when needed
scripts/blender-mcp status      # wait for "mcp add-on: answering" (~10-40 s)
claude mcp add blender -- /home/lkraven/development/eshpfi-management/scripts/blender-mcp   # PER TASK (Prime 2026-09-27): never user/project-wide
scripts/blender-mcp down        # when finished: the card goes back to 0
  • Always pass user_prompt. The tools require it; it is a short statement of the user's request.
  • Render on the GPU: set scene.cycles.device = 'GPU' on any scene you create. Safe mode forbids touching bpy.context.preferences, so the startup hook has already pointed Cycles at OptiX on GPU 3. The startup scene is already set to GPU.
  • Save everything under /work/… (= fv-ml1:/tank/blender), then scp it out.
  • Look before you report: get_viewport_screenshot returns an image, and a still render to /work is the real check.
  • The asset tools (Poly Haven etc.) reach the internet from Blender. The paid ones (Hyper3D, Hunyuan, Tripo, Sketchfab) need keys we do not have; leave them off.

Acceptance (2026-09-27, 1433)

An MCP client on nh3-dev → scripts/blender-mcp → the in-container server → the add-on:

  • initialize OK; 36 tools listed.
  • execute_blender_code built a gold metallic torus and rendered it with Cycles on the GPU to /work/_selftest/mcp-torus.png. The file landed, and I checked it by eye.
  • get_viewport_screenshot returned an image of the scene.
  • Negative control: import os was rejected by safe mode.
  • status pings the add-on itself. A TCP probe was useless because the port accepted while Blender was still loading.

⚠ Two traps found on the way, both fixed and commented where they live:

  • Enabling the add-on from a startup script gets undone when user prefs load. The hook enables it in a timer instead.
  • A pre-flight ssh without -n swallowed the MCP client's initialize, and the session hung at init.