Files
esh-pfi-infrastructure/docs/fleettools/blender.md
T
vh 46276611a7 fix(blender-run): unique remote job dir per local path, mirrored with --delete
draupnir review: the remote job dir was keyed on the basename alone, so two
local dirs with the same name shared one remote dir, and a rerun inherited
stale files. The dir is now <basename>-<8 hex of sha256(abs path)>, and it is
mirrored with rsync --delete, confined to that one directory. There is no rm on
a computed path. Verified: a file deleted locally is gone from the rerun's
remote dir. The /defaults stderr line is documented as harmless noise.
2026-09-28 08:43:55 -07:00

3.4 KiB

Blender: headless and agent-driven 3D on fv-ml1 GPU 3

Blender 5.2.2 LTS (bundled Python 3.13), runs on fv-ml1 GPU 3 (RTX PRO 6000 Blackwell, 96 GB). Stack and full notes: /home/lkraven/development/eshpfi-management/stacks/blender/README.md.

⚠ GPU 3 is borrowed. It is the fleet's reserve card for a full-size vLLM seat. Blender runs only while in use, and this access ends if a big seat moves onto the card.

Two ways in

You are… Use Shape
a script or CLI caller (renders, conversions) scripts/blender-run one-shot docker run --rm: no desktop, gone when Blender exits
an agent building scenes interactively scripts/blender-mcp (MCP, per task) a GUI Blender plus the mcp-for-blender add-on; up / status / down

Both scripts are in /home/lkraven/development/eshpfi-management/scripts/ and run from nh3-dev, reaching fv-ml1 as infra-ops@10.251.50.54 over ssh. No HTTP API and no openapi.json.

blender-run (headless)

scripts/blender-run --job /path/to/jobdir -- --python render.py -- out.png
scripts/blender-run -- --python-expr 'import bpy; print(bpy.app.version_string)'
  • Always adds -b --factory-startup --python-exit-code 1. Arguments go after --.
  • Files: fv-ml1 does not mount /mnt/smithy. --job DIR mirrors DIR to fv-ml1:/tank/blender/jobs/<basename>-<hash of DIR's absolute path>/, runs with that as the working directory, and copies new or changed files back into DIR. Nothing is ever deleted locally, and a rerun starts from an exact mirror, never from leftovers. Use relative input paths inside the job. Do not run the same DIR twice at once.
  • Engines headless (tested 2026-09-28): Cycles on GPU (OptiX/CUDA), Cycles on CPU, EEVEE (EGL, no display needed, ~9 s with the shader compile on first use), Workbench. In 5.2 the EEVEE id is BLENDER_EEVEE (BLENDER_EEVEE_NEXT is gone). For Cycles GPU, set prefs.compute_device_type = 'OPTIX' and enable the OPTIX devices, then scene.cycles.device = 'GPU'. Factory startup defaults to CPU.
  • Import: STL is built in (bpy.ops.wm.stl_import). No STEP importer is installed or built in.
  • Budget: each run is capped at 64 GB RAM and 48 CPUs, with up to the whole 96 GB of VRAM. Keep to about 2 concurrent renders. It is not on irv-ml1, so irv-ml1's working-set budget does not apply.

⚠ Foot-guns (all measured):

  • A --python script that raises exits 0 unless --python-exit-code is set. blender-run sets it.
  • render.filepath must be absolute. Blender does not resolve a relative output path against the working directory ("cannot save 'out.png'"). Importers and Python file I/O do.
  • Harmless stderr noise: "HIPEW initialization failed" (the AMD backend probing), and "Failed to create secure directory (/defaults): Operation not permitted" (the image's runtime-dir setup running as a non-root user; reported by draupnir 2026-09-28).
  • The first OptiX render in a process includes about 1-2 s of kernel load.

blender-mcp (agents)

Register it per task, never user- or project-wide (Prime, 2026-09-27): claude mcp add blender -- /home/lkraven/development/eshpfi-management/scripts/blender-mcp. Then run scripts/blender-mcp up, wait for status to say answering, work, and run down when finished. Safe mode is on (no os/open/network in agent code), so the startup hook has already pointed Cycles at OptiX. Save to /work/…. Details and traps are in the stack README.