Files
esh-pfi-infrastructure/stacks/blender/README.md
T
vh d0f68a3b18 feat(blender): blender-run one-shot headless wrapper + FLEETTOOLS entry
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).

Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
2026-09-28 08:39:43 -07:00

6.5 KiB
Raw Blame History

blender

Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand. Agents drive it; there is also a browser desktop to watch it or take over. Prime, 2026-09-27: "go ahead with gpu 3, both". He does not use Blender himself, so the agent side (MCP) is the primary interface.

Desktop https://10.251.50.54:3001 (self-signed cert). Basic auth: user blender, password secret get fv-ml1/blender-web-password.
Image lscr.io/linuxserver/blender:5.2.2-ls241@sha256:9216c77a… (Selkies 2.0 Wayland desktop, NVENC stream).
GPU GPU 3 only (NVIDIA_VISIBLE_DEVICES=3). About 270 MiB is held while the desktop runs.
Files /work → /tank/blender (projects, assets, renders; not backed up). /config → /opt/docker/data/blender (prefs, add-ons; restic). Files are owned by infra-ops (uid 1002), so agents can scp in and out.

⚠ On demand: GPU 3 is borrowed

GPU 3 is the fleet's reserve card for a full-size vLLM seat (servers/fv-ml1/README.md). Blender uses it only while in use:

ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose up -d'   # start
ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose down'    # stop, card back to 0

restart: "no", so a reboot never brings it back. When a big seat moves onto GPU 3, Blender stays down.

Headless rendering: scripts/blender-run (for scripts and CLI callers)

A one-shot docker run --rm of this image with Blender as the entrypoint. It needs no desktop and does not collide with the GUI container's up/down. --job DIR stages a local dir to /tank/blender/jobs/<name>/ and copies results back. Engines tested headless on 2026-09-28: Cycles GPU and CPU, EEVEE (EGL), Workbench. STL import is built in; there is no STEP importer. Foot-guns and budget are in docs/fleettools/blender.md. First consumer: draupnir.

Headless rendering inside the running GUI container

ssh infra-ops@10.251.50.54 'docker exec -u abc blender blender -b /work/<file>.blend -E CYCLES -o /work/out/frame_#### -a -- --cycles-device OPTIX'

Run as -u abc, the image's user, mapped to uid 1002, so outputs land owned by infra-ops.

Acceptance (2026-09-27, 1356)

  • Cycles sees the card on both OptiX and CUDA: "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition". The build ships kernel_sm_120.cubin plus OptiX PTX.
  • Self-test (/tank/blender/render_test.py): a subdivided glass monkey, 1920×1080, 1024 samples, 32 bounces, no denoise. OptiX 3.54 s against CPU 22.71 s (96 threads). That is n=1 per device: a functional check that the GPU is really used, not a benchmark. A trivial default-cube scene could not separate them (0.59 s against 0.65 s), which is why the self-test scene is heavy.
  • Web auth: no credentials gives 401, a wrong password 401, the right one 200.
  • Selkies: "Render node 1 encodes H264, AV1, H265 on nvenc"; the Wayland renderer runs GL on GPU 3.

Agent control (MCP): scripts/blender-mcp

The chosen server is mcp-for-blender (MIT, one maintainer, ~29k stars; researched by dvalin-smithy-dev 2026-09-27, thread 01M3JA61FTFW2ZPSD20MHF7RPJ, full note in dvalin-smithy research/blender-agent-drive-2026-09-27.md). Two halves:

  • The add-on is inside the running GUI Blender and serves a socket that executes arbitrary Python with no authentication. It is vendored at upstream commit 41a18432 (conf/scripts/addons/blender_mcp.py, MIT licence alongside) and started by conf/scripts/startup/fleet_mcp.py. The add-on only serves from a GUI Blender, never from blender -b, which is one reason the desktop exists.
  • The MCP server: mcp-for-blender==2.1.1, frozen in conf/mcp-requirements.txt and installed in /work/.mcp-venv inside the container (scripts/blender-mcp setup). Agents reach it as stdio over ssh infra-ops@fv-ml1 docker exec -i. It runs in the container for two reasons:
    1. No port is published. The socket stays on the container's localhost, so access means ssh + docker on fv-ml1.
    2. Viewport screenshots need a shared filesystem. Blender writes the image and the server reads it back. With the server on nh3-dev it failed ("Screenshot file was not created").

Settings: DISABLE_TELEMETRY=true and BLENDER_MCP_SAFE_MODE=1. Safe mode puts upstream's AST allowlist in front of execute_blender_code: bpy, bmesh, mathutils and pure stdlib only, and no os, open, eval or network. It guards against prompt injection from third-party asset text; it is not a sandbox.

For an agent

scripts/blender-mcp up          # GPU 3 is borrowed: start only when needed
scripts/blender-mcp status      # wait for "mcp add-on: answering" (~10-40 s)
claude mcp add blender -- /home/lkraven/development/eshpfi-management/scripts/blender-mcp   # PER TASK (Prime 2026-09-27): never user/project-wide
scripts/blender-mcp down        # when finished: the card goes back to 0
  • Always pass user_prompt. The tools require it; it is a short statement of the user's request.
  • Render on the GPU: set scene.cycles.device = 'GPU' on any scene you create. Safe mode forbids touching bpy.context.preferences, so the startup hook has already pointed Cycles at OptiX on GPU 3. The startup scene is already set to GPU.
  • Save everything under /work/… (= fv-ml1:/tank/blender), then scp it out.
  • Look before you report: get_viewport_screenshot returns an image, and a still render to /work is the real check.
  • The asset tools (Poly Haven etc.) reach the internet from Blender. The paid ones (Hyper3D, Hunyuan, Tripo, Sketchfab) need keys we do not have; leave them off.

Acceptance (2026-09-27, 1433)

An MCP client on nh3-dev → scripts/blender-mcp → the in-container server → the add-on:

  • initialize OK; 36 tools listed.
  • execute_blender_code built a gold metallic torus and rendered it with Cycles on the GPU to /work/_selftest/mcp-torus.png. The file landed, and I checked it by eye.
  • get_viewport_screenshot returned an image of the scene.
  • Negative control: import os was rejected by safe mode.
  • status pings the add-on itself. A TCP probe was useless because the port accepted while Blender was still loading.

⚠ Two traps found on the way, both fixed and commented where they live:

  • Enabling the add-on from a startup script gets undone when user prefs load. The hook enables it in a timer instead.
  • A pre-flight ssh without -n swallowed the MCP client's initialize, and the session hung at init.