Files
esh-pfi-infrastructure/stacks/mia/README.md
T
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00

6.6 KiB
Raw Blame History

mia — Make-It-Animatable v2 auto-rigger (fv-ml1 GPU 3, one-shot)

What: Make-It-Animatable v2 auto-rigs a humanoid mesh (GLB in): it predicts a 52-bone Mixamo-style skeleton and skin weights. Dread Naught (dread-dev) uses it to rig AI-generated characters. On 2026-10-01 it beat Blender bone-heat rigging on the four game characters (Booth dreadnaught-rigging).

Shape: on demand only, like blender-run (Prime, 2026-10-01). Each scripts/mia-run call is one docker run --rm: load the models, rig every mesh passed, write the job's out/, exit, hand GPU 3 back. Nothing stays up. There is no compose service; this dir is the image's build context.

scripts/mia-run --job DIR -- goblin=in/goblin.glb ogre=in/ogre.glb     # seeded (seed 0)
scripts/mia-run --job DIR --seed 7 -- goblin=in/goblin.glb
scripts/mia-run --job DIR --unseeded -- goblin=in/goblin.glb           # upstream behaviour

Inputs are relative to DIR. Per mesh, DIR/out/ gets <name>_pred.npz (with weights_effective), <name>.fbx, <name>.glb, <name>_rest.glb, the three _apose-hint variants, and <name>_run.json (timings, peak GPU memory, seed, provenance). The job mirroring is blender-run's: DIR → fv-ml1:/tank/mia/jobs/<basename>-<hash>/, results copied back.

What is pinned

piece pin where
code jasongzy/Make-It-Animatable branch v2 @ bbd8b158 (MIT); submodule util/Hunyuan3D_21 @ b6911977 cloned into the image, .git dropped
MIA weights HF jasongzy/Make-It-Animatable @ ca0daf6c (Apache-2.0): output/best/v2/{bw_joints,joints_coarse,pose}.pth, data/Standard Run.fbx, data/examples/ fv-ml1 HF cache /tank/aimodels/huggingface, mounted read-only at /hf
shape VAE HF tencent/Hunyuan3D-2.1 @ 0b946776: hunyuan3d-vae-v2-1/{config.yaml,model.fp16.ckpt} (Tencent Hunyuan community licence) same
Python requirements.lock: dread-dev's CPU venv verbatim, torch family swapped to cu129 (torch 2.8.0, torchvision 0.23.0, torch-cluster 1.6.3, pytorch3d 0.7.8); bpy 4.3.0, trimesh 5.1.0, numpy 1.26.4 image (/opt/mia/freeze.txt records the install)
driver mia_driver.py (adapted from dread-dev's tools/mia/run_mia.py + finalize.py) image, entrypoint

Every weight file's sha256 was checked against the HF API's LFS hash after download (2026-10-01). /opt/mia/provenance in the image and provenance in each _run.json carry the three pins.

Deviations from upstream (all deliberate)

  • The skeleton template is a substitute. data/Mixamo/bones.fbx → data/Standard Run.fbx. The official file is in the gated HF dataset jasongzy/Mixamo, whose terms were NOT accepted on Prime's account; accepting them is his call. The template sets bone roll only, never joints or weights (dread-dev). Swapping in the official file later changes FBX bone roll, not the npz.
  • Seeded by default. Before every mesh the driver calls MIA's fix_random(seed) (also sets cudnn deterministic) and resets trimesh.util._RANDOM_DEFAULT, which fix_random never reaches (trimesh ≥ 4 samples surface points from its own module RNG). Per-mesh reset also makes a mesh's result independent of the other meshes in the call.
  • No animation baked (animation_file=None; the UI default is "Standard Run.fbx").
  • app_v2.py is unpatched. dread-dev's CPU patch only upcasts to fp32 on CPU; the GPU runs upstream's fp16 autocast path.

Acceptance (2026-10-01) — instruments and raw numbers in acceptance/

Harness: fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, otherwise idle at 2 MiB), image local/mia:0.1.0, four meshes (dreadnaught/art/models/{goblin,villager,captain,ogre}.glb, 2.7k–3.4k verts), three seeded calls with the mesh order rotated, plus one unseeded call.

mesh GPU pipeline s, median of 3 (spread) model part s A-pose-hint export s CPU pipeline s (dread-dev, median of 3)
goblin 4.53 (4.43–4.57) 1.88 1.27 160.8
villager 3.97 (3.93–3.99) 1.61 1.24 116.9
captain 3.90 (3.79–3.94) 1.65 1.24 121.6
ogre 4.72 (4.45–4.80) 2.08 1.35 124.7
  • Model load: 12.7–13.0 s per call (4 calls), paid once per call. End to end, a 4-mesh call takes ~50 s including ssh, rsync and container start; a 1-mesh call ~31 s.
  • Peak VRAM: 3,394 MiB on GPU 3 as nvidia-smi saw it (100 ms polling over all four calls, includes the CUDA context); torch's peak reserved was 2,712–2,732 MiB per mesh.
  • Repeatable: the three seeded calls are bit-identical for every mesh (max abs diff 0.0 on joints, tails, weights, weights_effective, pose), despite the rotated mesh order.
  • Matches CPU within the noise. Distances use dread-dev's noise_floor.py metrics (joint distance / height, weight mass moved, dominant bone changed). GPU-vs-CPU averages 0.83–1.27× the CPU-vs-CPU distance, per mesh and metric. Positive control: one unseeded GPU call against a seeded one, i.e. pure sampling noise, lands at 0.80–1.83× — same size, so the instrument sees sampling noise and the GPU path adds nothing visible on top of it. 24 of 72 GPU-vs-CPU pair values exceed the CPU-vs-CPU maximum, but that maximum comes from only 3 pairs per mesh, and the sampling-only control overshoots it the same way.
  • Sensitivity floor: this comparison cannot resolve a systematic GPU-vs-CPU offset (fp16 vs fp32) smaller than the sampling noise, ~0.2% of mesh height for joints and ~2% weight mass moved. A seeded CPU run would isolate it; not done (would cost ~8 min of nh3-dev CPU).

Gotchas

  • init_blocks() reads data/examples/log.csv even headless; the image symlinks data/examples into the HF snapshot. Without it the driver dies before the first mesh.
  • uv needs UV_LINK_MODE=copy in the build (reflink clone fails on overlayfs, "os error 22").
  • The image is 10.5 GB on fv-ml1's zroot, which sat at 85% after the build. Check free space before building a new tag, and remove the old tag after.
  • GPU 3 is borrowed from the vLLM reserve and shared on demand with Blender and Scriberr. MIA needs ~3.4 GB, so it coexists with both; if a full-size seat claims the card, this tool has to move.

Rebuild

# on nh3-dev
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g root /opt/docker/src/mia-<ver>'
rsync -a stacks/mia/{Dockerfile,requirements.lock,mia_driver.py} infra-ops@10.251.50.54:/opt/docker/src/mia-<ver>/
ssh infra-ops@10.251.50.54 'cd /opt/docker/src/mia-<ver> && docker build -t local/mia:<ver> .'
# then point fv-ml1:/opt/docker/compose/mia/.env IMAGE= at the new tag