Files
esh-pfi-infrastructure/stacks/mia/README.md
T
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00

99 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# mia — Make-It-Animatable v2 auto-rigger (fv-ml1 GPU 3, one-shot)
**What:** [Make-It-Animatable](https://github.com/jasongzy/Make-It-Animatable) v2 auto-rigs a
humanoid mesh (GLB in): it predicts a 52-bone Mixamo-style skeleton and skin weights. Dread Naught
(dread-dev) uses it to rig AI-generated characters. On 2026-10-01 it beat Blender bone-heat rigging
on the four game characters (Booth `dreadnaught-rigging`).
**Shape: on demand only, like blender-run (Prime, 2026-10-01).** Each `scripts/mia-run` call is
one `docker run --rm`: load the models, rig every mesh passed, write the job's `out/`, exit, hand
GPU 3 back. Nothing stays up. There is no compose service; this dir is the image's build context.
```bash
scripts/mia-run --job DIR -- goblin=in/goblin.glb ogre=in/ogre.glb # seeded (seed 0)
scripts/mia-run --job DIR --seed 7 -- goblin=in/goblin.glb
scripts/mia-run --job DIR --unseeded -- goblin=in/goblin.glb # upstream behaviour
```
Inputs are relative to DIR. Per mesh, `DIR/out/` gets `<name>_pred.npz` (with
`weights_effective`), `<name>.fbx`, `<name>.glb`, `<name>_rest.glb`, the three `_apose-hint`
variants, and `<name>_run.json` (timings, peak GPU memory, seed, provenance). The job mirroring
is blender-run's: DIR → `fv-ml1:/tank/mia/jobs/<basename>-<hash>/`, results copied back.
## What is pinned
| piece | pin | where |
|---|---|---|
| code | `jasongzy/Make-It-Animatable` branch v2 @ `bbd8b158` (MIT); submodule `util/Hunyuan3D_21` @ `b6911977` | cloned into the image, `.git` dropped |
| MIA weights | HF `jasongzy/Make-It-Animatable` @ `ca0daf6c` (Apache-2.0): `output/best/v2/{bw_joints,joints_coarse,pose}.pth`, `data/Standard Run.fbx`, `data/examples/` | fv-ml1 HF cache `/tank/aimodels/huggingface`, mounted read-only at `/hf` |
| shape VAE | HF `tencent/Hunyuan3D-2.1` @ `0b946776`: `hunyuan3d-vae-v2-1/{config.yaml,model.fp16.ckpt}` (Tencent Hunyuan community licence) | same |
| Python | `requirements.lock`: dread-dev's CPU venv verbatim, torch family swapped to cu129 (torch 2.8.0, torchvision 0.23.0, torch-cluster 1.6.3, pytorch3d 0.7.8); `bpy` 4.3.0, `trimesh` 5.1.0, `numpy` 1.26.4 | image (`/opt/mia/freeze.txt` records the install) |
| driver | `mia_driver.py` (adapted from dread-dev's `tools/mia/run_mia.py` + `finalize.py`) | image, entrypoint |
Every weight file's sha256 was checked against the HF API's LFS hash after download (2026-10-01).
`/opt/mia/provenance` in the image and `provenance` in each `_run.json` carry the three pins.
## Deviations from upstream (all deliberate)
- **The skeleton template is a substitute.** `data/Mixamo/bones.fbx` → `data/Standard Run.fbx`.
The official file is in the **gated** HF dataset `jasongzy/Mixamo`, whose terms were NOT accepted
on Prime's account; accepting them is his call. The template sets bone roll only, never joints
or weights (dread-dev). Swapping in the official file later changes FBX bone roll, not the npz.
- **Seeded by default.** Before every mesh the driver calls MIA's `fix_random(seed)` (also sets
cudnn deterministic) and resets `trimesh.util._RANDOM_DEFAULT`, which `fix_random` never reaches
(trimesh ≥ 4 samples surface points from its own module RNG). Per-mesh reset also makes a mesh's
result independent of the other meshes in the call.
- **No animation baked** (`animation_file=None`; the UI default is "Standard Run.fbx").
- **app_v2.py is unpatched.** dread-dev's CPU patch only upcasts to fp32 on CPU; the GPU runs
upstream's fp16 autocast path.
## Acceptance (2026-10-01) — instruments and raw numbers in `acceptance/`
Harness: fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, otherwise idle at 2 MiB), image
`local/mia:0.1.0`, four meshes (`dreadnaught/art/models/{goblin,villager,captain,ogre}.glb`,
2.7k–3.4k verts), three seeded calls with the mesh order rotated, plus one unseeded call.
| mesh | GPU pipeline s, median of 3 (spread) | model part s | A-pose-hint export s | CPU pipeline s (dread-dev, median of 3) |
|---|---|---|---|---|
| goblin | **4.53** (4.43–4.57) | 1.88 | 1.27 | 160.8 |
| villager | **3.97** (3.93–3.99) | 1.61 | 1.24 | 116.9 |
| captain | **3.90** (3.79–3.94) | 1.65 | 1.24 | 121.6 |
| ogre | **4.72** (4.45–4.80) | 2.08 | 1.35 | 124.7 |
- **Model load:** 12.7–13.0 s per call (4 calls), paid once per call. End to end, a 4-mesh call
takes ~50 s including ssh, rsync and container start; a 1-mesh call ~31 s.
- **Peak VRAM:** 3,394 MiB on GPU 3 as nvidia-smi saw it (100 ms polling over all four calls,
includes the CUDA context); torch's peak reserved was 2,712–2,732 MiB per mesh.
- **Repeatable:** the three seeded calls are **bit-identical** for every mesh (max abs diff 0.0 on
joints, tails, weights, weights_effective, pose), despite the rotated mesh order.
- **Matches CPU within the noise.** Distances use dread-dev's `noise_floor.py` metrics (joint
distance / height, weight mass moved, dominant bone changed). GPU-vs-CPU averages 0.83–1.27× the
CPU-vs-CPU distance, per mesh and metric. Positive control: one unseeded GPU call against a
seeded one, i.e. pure sampling noise, lands at 0.80–1.83× — same size, so the instrument sees
sampling noise and the GPU path adds nothing visible on top of it. 24 of 72 GPU-vs-CPU pair
values exceed the CPU-vs-CPU maximum, but that maximum comes from only 3 pairs per mesh, and the
sampling-only control overshoots it the same way.
- **Sensitivity floor:** this comparison cannot resolve a systematic GPU-vs-CPU offset (fp16 vs
fp32) smaller than the sampling noise, ~0.2% of mesh height for joints and ~2% weight mass moved.
A seeded CPU run would isolate it; not done (would cost ~8 min of nh3-dev CPU).
## Gotchas
- `init_blocks()` reads `data/examples/log.csv` even headless; the image symlinks `data/examples`
into the HF snapshot. Without it the driver dies before the first mesh.
- uv needs `UV_LINK_MODE=copy` in the build (reflink clone fails on overlayfs, "os error 22").
- The image is 10.5 GB on fv-ml1's zroot, which sat at 85% after the build. Check free space before
building a new tag, and remove the old tag after.
- GPU 3 is borrowed from the vLLM reserve and shared on demand with Blender and Scriberr. MIA needs
~3.4 GB, so it coexists with both; if a full-size seat claims the card, this tool has to move.
## Rebuild
```bash
# on nh3-dev
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g root /opt/docker/src/mia-<ver>'
rsync -a stacks/mia/{Dockerfile,requirements.lock,mia_driver.py} infra-ops@10.251.50.54:/opt/docker/src/mia-<ver>/
ssh infra-ops@10.251.50.54 'cd /opt/docker/src/mia-<ver> && docker build -t local/mia:<ver> .'
# then point fv-ml1:/opt/docker/compose/mia/.env IMAGE= at the new tag
```