Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with its pinned submodules, dread-dev's proven Python lock with the torch family swapped to cu129, and a driver adapted from dread-dev's run_mia.py that seeds every mesh (fix_random + trimesh's module RNG) and writes weights_effective into the npz. Weights stay in fv-ml1's shared HF cache at pinned revisions, mounted read-only. scripts/mia-run mirrors blender-run: --job DIR is shipped to fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back. Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical across rotated mesh order; GPU-vs-CPU distances the same size as sampling noise, with an unseeded GPU run as the positive control.
6.6 KiB
mia — Make-It-Animatable v2 auto-rigger (fv-ml1 GPU 3, one-shot)
What: Make-It-Animatable v2 auto-rigs a
humanoid mesh (GLB in): it predicts a 52-bone Mixamo-style skeleton and skin weights. Dread Naught
(dread-dev) uses it to rig AI-generated characters. On 2026-10-01 it beat Blender bone-heat rigging
on the four game characters (Booth dreadnaught-rigging).
Shape: on demand only, like blender-run (Prime, 2026-10-01). Each scripts/mia-run call is
one docker run --rm: load the models, rig every mesh passed, write the job's out/, exit, hand
GPU 3 back. Nothing stays up. There is no compose service; this dir is the image's build context.
scripts/mia-run --job DIR -- goblin=in/goblin.glb ogre=in/ogre.glb # seeded (seed 0)
scripts/mia-run --job DIR --seed 7 -- goblin=in/goblin.glb
scripts/mia-run --job DIR --unseeded -- goblin=in/goblin.glb # upstream behaviour
Inputs are relative to DIR. Per mesh, DIR/out/ gets <name>_pred.npz (with
weights_effective), <name>.fbx, <name>.glb, <name>_rest.glb, the three _apose-hint
variants, and <name>_run.json (timings, peak GPU memory, seed, provenance). The job mirroring
is blender-run's: DIR → fv-ml1:/tank/mia/jobs/<basename>-<hash>/, results copied back.
What is pinned
| piece | pin | where |
|---|---|---|
| code | jasongzy/Make-It-Animatable branch v2 @ bbd8b158 (MIT); submodule util/Hunyuan3D_21 @ b6911977 |
cloned into the image, .git dropped |
| MIA weights | HF jasongzy/Make-It-Animatable @ ca0daf6c (Apache-2.0): output/best/v2/{bw_joints,joints_coarse,pose}.pth, data/Standard Run.fbx, data/examples/ |
fv-ml1 HF cache /tank/aimodels/huggingface, mounted read-only at /hf |
| shape VAE | HF tencent/Hunyuan3D-2.1 @ 0b946776: hunyuan3d-vae-v2-1/{config.yaml,model.fp16.ckpt} (Tencent Hunyuan community licence) |
same |
| Python | requirements.lock: dread-dev's CPU venv verbatim, torch family swapped to cu129 (torch 2.8.0, torchvision 0.23.0, torch-cluster 1.6.3, pytorch3d 0.7.8); bpy 4.3.0, trimesh 5.1.0, numpy 1.26.4 |
image (/opt/mia/freeze.txt records the install) |
| driver | mia_driver.py (adapted from dread-dev's tools/mia/run_mia.py + finalize.py) |
image, entrypoint |
Every weight file's sha256 was checked against the HF API's LFS hash after download (2026-10-01).
/opt/mia/provenance in the image and provenance in each _run.json carry the three pins.
Deviations from upstream (all deliberate)
- The skeleton template is a substitute.
data/Mixamo/bones.fbx→data/Standard Run.fbx. The official file is in the gated HF datasetjasongzy/Mixamo, whose terms were NOT accepted on Prime's account; accepting them is his call. The template sets bone roll only, never joints or weights (dread-dev). Swapping in the official file later changes FBX bone roll, not the npz. - Seeded by default. Before every mesh the driver calls MIA's
fix_random(seed)(also sets cudnn deterministic) and resetstrimesh.util._RANDOM_DEFAULT, whichfix_randomnever reaches (trimesh ≥ 4 samples surface points from its own module RNG). Per-mesh reset also makes a mesh's result independent of the other meshes in the call. - No animation baked (
animation_file=None; the UI default is "Standard Run.fbx"). - app_v2.py is unpatched. dread-dev's CPU patch only upcasts to fp32 on CPU; the GPU runs upstream's fp16 autocast path.
Acceptance (2026-10-01) — instruments and raw numbers in acceptance/
Harness: fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, otherwise idle at 2 MiB), image
local/mia:0.1.0, four meshes (dreadnaught/art/models/{goblin,villager,captain,ogre}.glb,
2.7k–3.4k verts), three seeded calls with the mesh order rotated, plus one unseeded call.
| mesh | GPU pipeline s, median of 3 (spread) | model part s | A-pose-hint export s | CPU pipeline s (dread-dev, median of 3) |
|---|---|---|---|---|
| goblin | 4.53 (4.43–4.57) | 1.88 | 1.27 | 160.8 |
| villager | 3.97 (3.93–3.99) | 1.61 | 1.24 | 116.9 |
| captain | 3.90 (3.79–3.94) | 1.65 | 1.24 | 121.6 |
| ogre | 4.72 (4.45–4.80) | 2.08 | 1.35 | 124.7 |
- Model load: 12.7–13.0 s per call (4 calls), paid once per call. End to end, a 4-mesh call takes ~50 s including ssh, rsync and container start; a 1-mesh call ~31 s.
- Peak VRAM: 3,394 MiB on GPU 3 as nvidia-smi saw it (100 ms polling over all four calls, includes the CUDA context); torch's peak reserved was 2,712–2,732 MiB per mesh.
- Repeatable: the three seeded calls are bit-identical for every mesh (max abs diff 0.0 on joints, tails, weights, weights_effective, pose), despite the rotated mesh order.
- Matches CPU within the noise. Distances use dread-dev's
noise_floor.pymetrics (joint distance / height, weight mass moved, dominant bone changed). GPU-vs-CPU averages 0.83–1.27× the CPU-vs-CPU distance, per mesh and metric. Positive control: one unseeded GPU call against a seeded one, i.e. pure sampling noise, lands at 0.80–1.83× — same size, so the instrument sees sampling noise and the GPU path adds nothing visible on top of it. 24 of 72 GPU-vs-CPU pair values exceed the CPU-vs-CPU maximum, but that maximum comes from only 3 pairs per mesh, and the sampling-only control overshoots it the same way. - Sensitivity floor: this comparison cannot resolve a systematic GPU-vs-CPU offset (fp16 vs fp32) smaller than the sampling noise, ~0.2% of mesh height for joints and ~2% weight mass moved. A seeded CPU run would isolate it; not done (would cost ~8 min of nh3-dev CPU).
Gotchas
init_blocks()readsdata/examples/log.csveven headless; the image symlinksdata/examplesinto the HF snapshot. Without it the driver dies before the first mesh.- uv needs
UV_LINK_MODE=copyin the build (reflink clone fails on overlayfs, "os error 22"). - The image is 10.5 GB on fv-ml1's zroot, which sat at 85% after the build. Check free space before building a new tag, and remove the old tag after.
- GPU 3 is borrowed from the vLLM reserve and shared on demand with Blender and Scriberr. MIA needs ~3.4 GB, so it coexists with both; if a full-size seat claims the card, this tool has to move.
Rebuild
# on nh3-dev
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g root /opt/docker/src/mia-<ver>'
rsync -a stacks/mia/{Dockerfile,requirements.lock,mia_driver.py} infra-ops@10.251.50.54:/opt/docker/src/mia-<ver>/
ssh infra-ops@10.251.50.54 'cd /opt/docker/src/mia-<ver> && docker build -t local/mia:<ver> .'
# then point fv-ml1:/opt/docker/compose/mia/.env IMAGE= at the new tag