Files
esh-pfi-infrastructure/docs/fleettools/blender.md
T
vh d0f68a3b18 feat(blender): blender-run one-shot headless wrapper + FLEETTOOLS entry
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).

Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
2026-09-28 08:39:43 -07:00

3.1 KiB

Blender: headless and agent-driven 3D on fv-ml1 GPU 3

Blender 5.2.2 LTS (bundled Python 3.13), runs on fv-ml1 GPU 3 (RTX PRO 6000 Blackwell, 96 GB). Stack and full notes: /home/lkraven/development/eshpfi-management/stacks/blender/README.md.

⚠ GPU 3 is borrowed. It is the fleet's reserve card for a full-size vLLM seat. Blender runs only while in use, and this access ends if a big seat moves onto the card.

Two ways in

You are… Use Shape
a script or CLI caller (renders, conversions) scripts/blender-run one-shot docker run --rm: no desktop, gone when Blender exits
an agent building scenes interactively scripts/blender-mcp (MCP, per task) a GUI Blender plus the mcp-for-blender add-on; up / status / down

Both scripts are in /home/lkraven/development/eshpfi-management/scripts/ and run from nh3-dev, reaching fv-ml1 as infra-ops@10.251.50.54 over ssh. No HTTP API and no openapi.json.

blender-run (headless)

scripts/blender-run --job /path/to/jobdir -- --python render.py -- out.png
scripts/blender-run -- --python-expr 'import bpy; print(bpy.app.version_string)'
  • Always adds -b --factory-startup --python-exit-code 1. Arguments go after --.
  • Files: fv-ml1 does not mount /mnt/smithy. --job DIR copies DIR to fv-ml1:/tank/blender/jobs/<name>/, runs with that as the working directory, and copies new or changed files back into DIR. Nothing is deleted on either side. Use relative input paths inside the job.
  • Engines headless (tested 2026-09-28): Cycles on GPU (OptiX/CUDA), Cycles on CPU, EEVEE (EGL, no display needed, ~9 s with the shader compile on first use), Workbench. In 5.2 the EEVEE id is BLENDER_EEVEE (BLENDER_EEVEE_NEXT is gone). For Cycles GPU, set prefs.compute_device_type = 'OPTIX' and enable the OPTIX devices, then scene.cycles.device = 'GPU'. Factory startup defaults to CPU.
  • Import: STL is built in (bpy.ops.wm.stl_import). No STEP importer is installed or built in.
  • Budget: each run is capped at 64 GB RAM and 48 CPUs, with up to the whole 96 GB of VRAM. Keep to about 2 concurrent renders. It is not on irv-ml1, so irv-ml1's working-set budget does not apply.

⚠ Foot-guns (all measured):

  • A --python script that raises exits 0 unless --python-exit-code is set. blender-run sets it.
  • render.filepath must be absolute. Blender does not resolve a relative output path against the working directory ("cannot save 'out.png'"). Importers and Python file I/O do.
  • "HIPEW initialization failed" on stderr is harmless: it is the AMD backend probing.
  • The first OptiX render in a process includes about 1-2 s of kernel load.

blender-mcp (agents)

Register it per task, never user- or project-wide (Prime, 2026-09-27): claude mcp add blender -- /home/lkraven/development/eshpfi-management/scripts/blender-mcp. Then run scripts/blender-mcp up, wait for status to say answering, work, and run down when finished. Safe mode is on (no os/open/network in agent code), so the startup hook has already pointed Cycles at OptiX. Save to /work/…. Details and traps are in the stack README.