Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with its pinned submodules, dread-dev's proven Python lock with the torch family swapped to cu129, and a driver adapted from dread-dev's run_mia.py that seeds every mesh (fix_random + trimesh's module RNG) and writes weights_effective into the npz. Weights stay in fv-ml1's shared HF cache at pinned revisions, mounted read-only. scripts/mia-run mirrors blender-run: --job DIR is shipped to fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back. Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical across rotated mesh order; GPU-vs-CPU distances the same size as sampling noise, with an unseeded GPU run as the positive control.
99 lines
6.6 KiB
Markdown
99 lines
6.6 KiB
Markdown
# mia — Make-It-Animatable v2 auto-rigger (fv-ml1 GPU 3, one-shot)
|
||
|
||
**What:** [Make-It-Animatable](https://github.com/jasongzy/Make-It-Animatable) v2 auto-rigs a
|
||
humanoid mesh (GLB in): it predicts a 52-bone Mixamo-style skeleton and skin weights. Dread Naught
|
||
(dread-dev) uses it to rig AI-generated characters. On 2026-10-01 it beat Blender bone-heat rigging
|
||
on the four game characters (Booth `dreadnaught-rigging`).
|
||
|
||
**Shape: on demand only, like blender-run (Prime, 2026-10-01).** Each `scripts/mia-run` call is
|
||
one `docker run --rm`: load the models, rig every mesh passed, write the job's `out/`, exit, hand
|
||
GPU 3 back. Nothing stays up. There is no compose service; this dir is the image's build context.
|
||
|
||
```bash
|
||
scripts/mia-run --job DIR -- goblin=in/goblin.glb ogre=in/ogre.glb # seeded (seed 0)
|
||
scripts/mia-run --job DIR --seed 7 -- goblin=in/goblin.glb
|
||
scripts/mia-run --job DIR --unseeded -- goblin=in/goblin.glb # upstream behaviour
|
||
```
|
||
|
||
Inputs are relative to DIR. Per mesh, `DIR/out/` gets `<name>_pred.npz` (with
|
||
`weights_effective`), `<name>.fbx`, `<name>.glb`, `<name>_rest.glb`, the three `_apose-hint`
|
||
variants, and `<name>_run.json` (timings, peak GPU memory, seed, provenance). The job mirroring
|
||
is blender-run's: DIR → `fv-ml1:/tank/mia/jobs/<basename>-<hash>/`, results copied back.
|
||
|
||
## What is pinned
|
||
|
||
| piece | pin | where |
|
||
|---|---|---|
|
||
| code | `jasongzy/Make-It-Animatable` branch v2 @ `bbd8b158` (MIT); submodule `util/Hunyuan3D_21` @ `b6911977` | cloned into the image, `.git` dropped |
|
||
| MIA weights | HF `jasongzy/Make-It-Animatable` @ `ca0daf6c` (Apache-2.0): `output/best/v2/{bw_joints,joints_coarse,pose}.pth`, `data/Standard Run.fbx`, `data/examples/` | fv-ml1 HF cache `/tank/aimodels/huggingface`, mounted read-only at `/hf` |
|
||
| shape VAE | HF `tencent/Hunyuan3D-2.1` @ `0b946776`: `hunyuan3d-vae-v2-1/{config.yaml,model.fp16.ckpt}` (Tencent Hunyuan community licence) | same |
|
||
| Python | `requirements.lock`: dread-dev's CPU venv verbatim, torch family swapped to cu129 (torch 2.8.0, torchvision 0.23.0, torch-cluster 1.6.3, pytorch3d 0.7.8); `bpy` 4.3.0, `trimesh` 5.1.0, `numpy` 1.26.4 | image (`/opt/mia/freeze.txt` records the install) |
|
||
| driver | `mia_driver.py` (adapted from dread-dev's `tools/mia/run_mia.py` + `finalize.py`) | image, entrypoint |
|
||
|
||
Every weight file's sha256 was checked against the HF API's LFS hash after download (2026-10-01).
|
||
`/opt/mia/provenance` in the image and `provenance` in each `_run.json` carry the three pins.
|
||
|
||
## Deviations from upstream (all deliberate)
|
||
|
||
- **The skeleton template is a substitute.** `data/Mixamo/bones.fbx` → `data/Standard Run.fbx`.
|
||
The official file is in the **gated** HF dataset `jasongzy/Mixamo`, whose terms were NOT accepted
|
||
on Prime's account; accepting them is his call. The template sets bone roll only, never joints
|
||
or weights (dread-dev). Swapping in the official file later changes FBX bone roll, not the npz.
|
||
- **Seeded by default.** Before every mesh the driver calls MIA's `fix_random(seed)` (also sets
|
||
cudnn deterministic) and resets `trimesh.util._RANDOM_DEFAULT`, which `fix_random` never reaches
|
||
(trimesh ≥ 4 samples surface points from its own module RNG). Per-mesh reset also makes a mesh's
|
||
result independent of the other meshes in the call.
|
||
- **No animation baked** (`animation_file=None`; the UI default is "Standard Run.fbx").
|
||
- **app_v2.py is unpatched.** dread-dev's CPU patch only upcasts to fp32 on CPU; the GPU runs
|
||
upstream's fp16 autocast path.
|
||
|
||
## Acceptance (2026-10-01) — instruments and raw numbers in `acceptance/`
|
||
|
||
Harness: fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, otherwise idle at 2 MiB), image
|
||
`local/mia:0.1.0`, four meshes (`dreadnaught/art/models/{goblin,villager,captain,ogre}.glb`,
|
||
2.7k–3.4k verts), three seeded calls with the mesh order rotated, plus one unseeded call.
|
||
|
||
| mesh | GPU pipeline s, median of 3 (spread) | model part s | A-pose-hint export s | CPU pipeline s (dread-dev, median of 3) |
|
||
|---|---|---|---|---|
|
||
| goblin | **4.53** (4.43–4.57) | 1.88 | 1.27 | 160.8 |
|
||
| villager | **3.97** (3.93–3.99) | 1.61 | 1.24 | 116.9 |
|
||
| captain | **3.90** (3.79–3.94) | 1.65 | 1.24 | 121.6 |
|
||
| ogre | **4.72** (4.45–4.80) | 2.08 | 1.35 | 124.7 |
|
||
|
||
- **Model load:** 12.7–13.0 s per call (4 calls), paid once per call. End to end, a 4-mesh call
|
||
takes ~50 s including ssh, rsync and container start; a 1-mesh call ~31 s.
|
||
- **Peak VRAM:** 3,394 MiB on GPU 3 as nvidia-smi saw it (100 ms polling over all four calls,
|
||
includes the CUDA context); torch's peak reserved was 2,712–2,732 MiB per mesh.
|
||
- **Repeatable:** the three seeded calls are **bit-identical** for every mesh (max abs diff 0.0 on
|
||
joints, tails, weights, weights_effective, pose), despite the rotated mesh order.
|
||
- **Matches CPU within the noise.** Distances use dread-dev's `noise_floor.py` metrics (joint
|
||
distance / height, weight mass moved, dominant bone changed). GPU-vs-CPU averages 0.83–1.27× the
|
||
CPU-vs-CPU distance, per mesh and metric. Positive control: one unseeded GPU call against a
|
||
seeded one, i.e. pure sampling noise, lands at 0.80–1.83× — same size, so the instrument sees
|
||
sampling noise and the GPU path adds nothing visible on top of it. 24 of 72 GPU-vs-CPU pair
|
||
values exceed the CPU-vs-CPU maximum, but that maximum comes from only 3 pairs per mesh, and the
|
||
sampling-only control overshoots it the same way.
|
||
- **Sensitivity floor:** this comparison cannot resolve a systematic GPU-vs-CPU offset (fp16 vs
|
||
fp32) smaller than the sampling noise, ~0.2% of mesh height for joints and ~2% weight mass moved.
|
||
A seeded CPU run would isolate it; not done (would cost ~8 min of nh3-dev CPU).
|
||
|
||
## Gotchas
|
||
|
||
- `init_blocks()` reads `data/examples/log.csv` even headless; the image symlinks `data/examples`
|
||
into the HF snapshot. Without it the driver dies before the first mesh.
|
||
- uv needs `UV_LINK_MODE=copy` in the build (reflink clone fails on overlayfs, "os error 22").
|
||
- The image is 10.5 GB on fv-ml1's zroot, which sat at 85% after the build. Check free space before
|
||
building a new tag, and remove the old tag after.
|
||
- GPU 3 is borrowed from the vLLM reserve and shared on demand with Blender and Scriberr. MIA needs
|
||
~3.4 GB, so it coexists with both; if a full-size seat claims the card, this tool has to move.
|
||
|
||
## Rebuild
|
||
|
||
```bash
|
||
# on nh3-dev
|
||
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g root /opt/docker/src/mia-<ver>'
|
||
rsync -a stacks/mia/{Dockerfile,requirements.lock,mia_driver.py} infra-ops@10.251.50.54:/opt/docker/src/mia-<ver>/
|
||
ssh infra-ops@10.251.50.54 'cd /opt/docker/src/mia-<ver> && docker build -t local/mia:<ver> .'
|
||
# then point fv-ml1:/opt/docker/compose/mia/.env IMAGE= at the new tag
|
||
```
|