Files
esh-pfi-infrastructure/stacks/blender/README.md
T
vh d0f68a3b18 feat(blender): blender-run one-shot headless wrapper + FLEETTOOLS entry
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).

Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
2026-09-28 08:39:43 -07:00

115 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# blender
**Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand.** Agents drive it; there is also a browser
desktop to watch it or take over. Prime, 2026-09-27: "go ahead with gpu 3, both". He does not
use Blender himself, so the agent side (MCP) is the primary interface.
| | |
|---|---|
| **Desktop** | `https://10.251.50.54:3001` (self-signed cert). Basic auth: user `blender`, password `secret get fv-ml1/blender-web-password`. |
| **Image** | `lscr.io/linuxserver/blender:5.2.2-ls241@sha256:9216c77a…` (Selkies 2.0 Wayland desktop, NVENC stream). |
| **GPU** | GPU 3 only (`NVIDIA_VISIBLE_DEVICES=3`). About 270 MiB is held while the desktop runs. |
| **Files** | `/work` → `/tank/blender` (projects, assets, renders; **not backed up**). `/config` → `/opt/docker/data/blender` (prefs, add-ons; restic). Files are owned by infra-ops (uid 1002), so agents can `scp` in and out. |
## ⚠ On demand: GPU 3 is borrowed
GPU 3 is the fleet's reserve card for a full-size vLLM seat (`servers/fv-ml1/README.md`).
Blender uses it only while in use:
```bash
ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose up -d' # start
ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/blender && docker compose down' # stop, card back to 0
```
`restart: "no"`, so a reboot never brings it back. **When a big seat moves onto GPU 3, Blender
stays down.**
## Headless rendering: `scripts/blender-run` (for scripts and CLI callers)
A one-shot `docker run --rm` of this image with Blender as the entrypoint. It needs no desktop and
does not collide with the GUI container's up/down. `--job DIR` stages a local dir to
`/tank/blender/jobs/<name>/` and copies results back. Engines tested headless on 2026-09-28:
Cycles GPU and CPU, EEVEE (EGL), Workbench. STL import is built in; there is **no STEP importer**.
Foot-guns and budget are in `docs/fleettools/blender.md`. First consumer: draupnir.
## Headless rendering inside the running GUI container
```bash
ssh infra-ops@10.251.50.54 'docker exec -u abc blender blender -b /work/<file>.blend -E CYCLES -o /work/out/frame_#### -a -- --cycles-device OPTIX'
```
Run as `-u abc`, the image's user, mapped to uid 1002, so outputs land owned by infra-ops.
## Acceptance (2026-09-27, 1356)
- Cycles sees the card on both OptiX and CUDA: "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation
Edition". The build ships `kernel_sm_120.cubin` plus OptiX PTX.
- Self-test (`/tank/blender/render_test.py`): a subdivided glass monkey, 1920×1080, 1024 samples,
32 bounces, no denoise. **OptiX 3.54 s against CPU 22.71 s (96 threads).** That is n=1 per
device: a functional check that the GPU is really used, not a benchmark. A trivial default-cube
scene could not separate them (0.59 s against 0.65 s), which is why the self-test scene is heavy.
- Web auth: no credentials gives 401, a wrong password 401, the right one 200.
- Selkies: "Render node 1 encodes H264, AV1, H265 on nvenc"; the Wayland renderer runs GL on GPU 3.
## Agent control (MCP): `scripts/blender-mcp`
**The chosen server is [mcp-for-blender](https://github.com/ahujasid/mcp-for-blender)** (MIT, one
maintainer, ~29k stars; researched by dvalin-smithy-dev 2026-09-27, thread
`01M3JA61FTFW2ZPSD20MHF7RPJ`, full note in dvalin-smithy
`research/blender-agent-drive-2026-09-27.md`). Two halves:
- **The add-on** is inside the running GUI Blender and serves a socket that executes arbitrary
Python with **no authentication**. It is vendored at upstream commit `41a18432`
(`conf/scripts/addons/blender_mcp.py`, MIT licence alongside) and started by
`conf/scripts/startup/fleet_mcp.py`. The add-on only serves from a GUI Blender, never from
`blender -b`, which is one reason the desktop exists.
- **The MCP server**: `mcp-for-blender==2.1.1`, frozen in `conf/mcp-requirements.txt` and installed in
`/work/.mcp-venv` **inside the container** (`scripts/blender-mcp setup`). Agents reach it as
stdio over `ssh infra-ops@fv-ml1 docker exec -i`. It runs in the container for two reasons:
1. **No port is published.** The socket stays on the container's localhost, so access means
ssh + docker on fv-ml1.
2. **Viewport screenshots need a shared filesystem.** Blender writes the image and the server
reads it back. With the server on nh3-dev it failed ("Screenshot file was not created").
Settings: `DISABLE_TELEMETRY=true` and `BLENDER_MCP_SAFE_MODE=1`. Safe mode puts upstream's AST
allowlist in front of `execute_blender_code`: bpy, bmesh, mathutils and pure stdlib only, and no
os, open, eval or network. It guards against prompt injection from third-party asset text; it is
not a sandbox.
### For an agent
```bash
scripts/blender-mcp up # GPU 3 is borrowed: start only when needed
scripts/blender-mcp status # wait for "mcp add-on: answering" (~10-40 s)
claude mcp add blender -- /home/lkraven/development/eshpfi-management/scripts/blender-mcp # PER TASK (Prime 2026-09-27): never user/project-wide
scripts/blender-mcp down # when finished: the card goes back to 0
```
- **Always pass `user_prompt`.** The tools require it; it is a short statement of the user's
request.
- **Render on the GPU:** set `scene.cycles.device = 'GPU'` on any scene you create. Safe mode
forbids touching `bpy.context.preferences`, so the startup hook has already pointed Cycles at
OptiX on GPU 3. The startup scene is already set to GPU.
- **Save everything under `/work/…`** (= `fv-ml1:/tank/blender`), then `scp` it out.
- **Look before you report:** `get_viewport_screenshot` returns an image, and a still render to
`/work` is the real check.
- The asset tools (Poly Haven etc.) reach the internet from Blender. The paid ones (Hyper3D,
Hunyuan, Tripo, Sketchfab) need keys we do not have; leave them off.
### Acceptance (2026-09-27, 1433)
An MCP client on nh3-dev → `scripts/blender-mcp` → the in-container server → the add-on:
- `initialize` OK; **36 tools** listed.
- `execute_blender_code` built a gold metallic torus and rendered it with Cycles on the GPU to
`/work/_selftest/mcp-torus.png`. The file landed, and I checked it by eye.
- `get_viewport_screenshot` returned an image of the scene.
- **Negative control:** `import os` was rejected by safe mode.
- `status` pings the add-on itself. A TCP probe was useless because the port accepted while
Blender was still loading.
⚠ Two traps found on the way, both fixed and commented where they live:
- Enabling the add-on from a startup script gets undone when user prefs load. The hook enables it
in a timer instead.
- A pre-flight `ssh` without `-n` swallowed the MCP client's `initialize`, and the session hung at
init.