feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)

Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
This commit is contained in:
vh
2026-10-01 09:45:22 -07:00
parent e5536784e0
commit f44280e6b2
12 changed files with 3121 additions and 1 deletions
+6
View File
@@ -82,6 +82,12 @@ live contract for its API. Fetch it rather than trusting a transcription.
*When:* 3D rendering, STL → image, scene building.
*Detail:* `/home/lkraven/development/eshpfi-management/docs/fleettools/blender.md`
- **MIA auto-rig** — Make-It-Animatable v2 on fv-ml1 GPU 3, one-shot: humanoid GLB in, 52-bone
Mixamo-style skeleton + skin weights out (npz, FBX, GLB), ~4.5 s a mesh plus ~13 s model load.
`scripts/mia-run --job DIR -- <name>=<in.glb> ...`. Seeded by default.
*When:* rigging a character mesh for animation.
*Detail:* `/home/lkraven/development/eshpfi-management/stacks/mia/README.md`
## Working on the fleet itself
- **elway** — SSH playbook runner for **CHANGING** things.
+2 -1
View File
@@ -150,7 +150,7 @@ _As of 2026-10-01 ~0446 PT._
- **GPU 0:** cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
- **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
- **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1.
- **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, Scriberr (0 idle, ~5.5 GB per job), and MIA auto-rig (one-shot `scripts/mia-run`, peak 3.4 GB, added 2026-10-01 for dread-dev). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1.
### intern-decision (replaced SemIf on 2026-09-30)
@@ -290,6 +290,7 @@ _As of 2026-10-01 ~0446 PT._
## Recent decisions
- `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md`
- `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it.
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
+85
View File
@@ -0,0 +1,85 @@
#!/usr/bin/env bash
# mia-run — one-shot Make-It-Animatable v2 auto-rig on fv-ml1 GPU 3 (dread-dev / Dread Naught;
# Prime's request, 2026-10-01). Same shape as blender-run: a container starts, rigs, exits, and
# hands the GPU back. NOT a seat; nothing stays up between calls.
#
# scripts/mia-run --job DIR [--seed N | --unseeded] -- <name>=<input.glb> [<name>=<input.glb> ...]
# scripts/mia-run --job ~/development/dreadnaught/rig-jobs/j1 -- goblin=in/goblin.glb ogre=in/ogre.glb
#
# Inputs are paths RELATIVE TO DIR and must be inside it: DIR is the only thing shipped to fv-ml1.
# Per mesh, DIR/out/ receives:
# <name>_pred.npz predictions in the input's coordinates, WITH `weights_effective`
# (the finalize.py formula), so no finalize step is needed afterwards
# <name>.fbx .glb _rest.glb MIA's own exports (UI defaults, static rig)
# <name>_apose-hint.fbx .glb _rest.glb same predictions, Blender stage with the A-pose hint
# <name>_run.json stage timings, model-load time, peak GPU memory, seed, provenance
# The model load (~3 x 1 GB checkpoints) is paid once per call, so pass every mesh in one call.
#
# SEEDED BY DEFAULT (seed 0): before every mesh the driver calls MIA's fix_random(seed) and resets
# trimesh's module RNG, which fix_random never reaches (trimesh >= 4 samples surfaces from
# trimesh.util._RANDOM_DEFAULT). Without that, runs of the same mesh differ: joints by 0.1-0.27% of
# height on dread-dev's CPU noise floor. --unseeded restores upstream behaviour.
#
# Files: fv-ml1 does NOT mount /mnt/smithy. DIR is mirrored to
# fv-ml1:/tank/mia/jobs/<basename>-<hash of DIR's absolute path>/ (rsync --delete, so a rerun never
# inherits leftovers), the driver runs WITH THAT AS ITS WORKING DIRECTORY, and new or changed files
# are copied back into DIR afterwards (nothing is ever deleted locally). Running the SAME DIR twice
# concurrently shares one remote dir, so don't. Remote job dirs are never swept.
#
# Image: local/mia (stacks/mia; IMAGE= in fv-ml1:/opt/docker/compose/mia/.env). Weights come from
# the read-only /tank/aimodels/huggingface mount at pinned revisions; the run never touches the
# network (HF_HUB_OFFLINE=1). The skeleton template is a SUBSTITUTE ("Standard Run.fbx"): the
# official one is in a gated HF dataset whose terms were not accepted (see stacks/mia/README.md).
#
# Budget: GPU 3 (96 GB, borrowed from the vLLM reserve, shared on demand with Blender and Scriberr;
# this ends if a full-size seat moves in). Capped at 32 GB RAM / 16 CPUs so a runaway mesh cannot
# starve the inference seats on the same host.
set -euo pipefail
HOST=${MIA_SSH_HOST:-infra-ops@10.251.50.54}
ENV_FILE=/opt/docker/compose/mia/.env
JOB=""
DRIVER_OPTS=()
while [ $# -gt 0 ]; do
case "$1" in
--job) JOB=${2:?--job needs a directory}; shift 2 ;;
--seed) DRIVER_OPTS+=(--seed "${2:?--seed needs an integer}"); shift 2 ;;
--unseeded) DRIVER_OPTS+=(--unseeded); shift ;;
--) shift; break ;;
-h|--help) sed -n 2,40p "$0"; exit 0 ;;
*) echo "mia-run: unknown option $1 (meshes go after --)" >&2; exit 2 ;;
esac
done
[ -n "$JOB" ] || { echo "mia-run: --job DIR is required (inputs are read from it, out/ is written into it)" >&2; exit 2; }
[ $# -gt 0 ] || { echo "mia-run: no meshes; pass <name>=<input.glb> after --" >&2; exit 2; }
[ -d "$JOB" ] || { echo "mia-run: --job $JOB is not a directory" >&2; exit 2; }
ABS=$(realpath "$JOB")
BASE=$(basename "$ABS")
[[ $BASE =~ ^[A-Za-z0-9._-]+$ ]] || { echo "mia-run: job dir name '$BASE' must be [A-Za-z0-9._-]" >&2; exit 2; }
for m in "$@"; do
[[ $m == *=* ]] || { echo "mia-run: '$m' is not <name>=<input.glb>" >&2; exit 2; }
f=$(realpath -m "$ABS/${m#*=}")
[[ $f == "$ABS"/* && -f $f ]] || { echo "mia-run: input '${m#*=}' is not a file inside $ABS" >&2; exit 2; }
done
# Keyed on basename + a hash of the ABSOLUTE local path, as blender-run: two dirs that share a
# basename never share a remote dir. rsync removes only inside that one remote dir; there is no rm
# on a computed path.
NAME=$BASE-$(printf '%s' "$ABS" | sha256sum | cut -c1-8)
ssh -n -o BatchMode=yes "$HOST" "mkdir -p /tank/mia/jobs/$NAME"
rsync -a --delete "$JOB"/ "$HOST:/tank/mia/jobs/$NAME/"
# Arguments travel as one shell-quoted string: ssh flattens argv into a remote command line.
ARGS=$(printf '%q ' "${DRIVER_OPTS[@]}" "$@")
set +e
ssh -n -o BatchMode=yes "$HOST" "IMG=\$(grep '^IMAGE=' $ENV_FILE | cut -d= -f2) && \
exec docker run --rm --name mia-run-\$\$ --hostname fv-ml1-mia \
--runtime nvidia -e NVIDIA_VISIBLE_DEVICES=3 -e NVIDIA_DRIVER_CAPABILITIES=compute,utility \
--user 1002:1003 --memory 32g --cpus 16 \
--mount type=bind,src=/tank/aimodels/huggingface,dst=/hf,readonly \
--mount type=bind,src=/tank/mia/jobs/$NAME,dst=/job \
-w /job \"\$IMG\" $ARGS"
RC=$?
set -e
rsync -a --update "$HOST:/tank/mia/jobs/$NAME/" "$JOB"/
exit $RC
+2
View File
@@ -0,0 +1,2 @@
# Image scripts/mia-run starts (built from this dir on fv-ml1 under /opt/docker/src/mia-<ver>/).
IMAGE=local/mia:0.1.0
+69
View File
@@ -0,0 +1,69 @@
# Make-It-Animatable (MIA) v2: one-shot ML auto-rigger for humanoid meshes (GLB in, 52-bone
# Mixamo-style skeleton + skin weights out). Run by scripts/mia-run on fv-ml1 GPU 3, for dread-dev
# (Dread Naught), Prime's request 2026-10-01. NOT a seat: each run is a `docker run --rm` that
# rigs, writes the job's out/, exits and hands the GPU back.
#
# Code: github.com/jasongzy/Make-It-Animatable, branch v2 @ MIA_COMMIT (MIT), cloned with the
# submodules that commit pins (util/Hunyuan3D_21 @ b691197, util/auto_rig_pro, util/3dgs-render-
# blender-addon). Python: requirements.lock (dread-dev's proven CPU venv, torch family swapped to
# cu129). Upstream app_v2.py is UNPATCHED: dread-dev's CPU patch only upcasts on CPU.
#
# Weights are NOT baked in. /tank/aimodels/huggingface is bind-mounted read-only at /hf, and the two
# pinned snapshots are reached through the symlinks made below:
# repo/output -> jasongzy/Make-It-Animatable @ MIA_WEIGHTS_REV (output/best/v2/*.pth, Apache-2.0)
# repo/data/Standard Run.fbx, repo/data/examples -> the same snapshot (init_blocks() reads
# data/examples/log.csv even headless; without it the driver dies before the first mesh)
# /opt/mia/hy3dgen/tencent/Hunyuan3D-2.1 -> tencent/Hunyuan3D-2.1 @ HY3D_REV (hunyuan3d-vae-v2-1)
# repo/data/Mixamo/bones.fbx -> "Standard Run.fbx": the official template sits in the GATED HF
# dataset jasongzy/Mixamo, whose terms were not accepted on Prime's account. It sets bone roll
# only, never joints or weights (dread-dev, 2026-10-01). Accepting those terms is Prime's call.
FROM python:3.11-slim-bookworm
ARG MIA_COMMIT=bbd8b158d88879c310ad130f9b25056935d221e9
ARG MIA_WEIGHTS_REV=ca0daf6cb164f939e77bf32667513fc7558d5f98
ARG HY3D_REV=0b94677654c57bb9a6b6845cd7b704ccf551d327
ENV DEBIAN_FRONTEND=noninteractive \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
HF_HOME=/hf \
HF_HUB_OFFLINE=1 \
GRADIO_ANALYTICS_ENABLED=False \
MPLCONFIGDIR=/tmp/mpl \
HY3DGEN_MODELS=/opt/mia/hy3dgen \
HOME=/tmp
# git for the clone; the X11/GL libs are what the bpy 4.3 wheel links against (import fails without).
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates libgl1 libegl1 libglib2.0-0 libx11-6 libxrender1 libxxf86vm1 libxfixes3 \
libxi6 libxkbcommon0 libsm6 libice6 libgomp1 \
&& rm -rf /var/lib/apt/lists/*
# UV_LINK_MODE=copy: uv's default clone (reflink) fails on the build's overlay fs, "os error 22".
COPY requirements.lock /opt/mia/requirements.lock
RUN pip install -q uv \
&& UV_LINK_MODE=copy uv pip install --system --no-cache --index-strategy unsafe-best-match \
--extra-index-url https://download.pytorch.org/whl/cu129 \
--extra-index-url https://download.blender.org/pypi/ \
--extra-index-url https://miropsota.github.io/torch_packages_builder \
--find-links https://data.pyg.org/whl/torch-2.8.0+cu129.html \
-r /opt/mia/requirements.lock \
&& uv pip freeze --system > /opt/mia/freeze.txt
RUN git clone --single-branch --branch v2 https://github.com/jasongzy/Make-It-Animatable /opt/mia/repo \
&& cd /opt/mia/repo && git checkout -q "$MIA_COMMIT" && git submodule update --init --recursive -q \
&& git -C util/Hunyuan3D_21 rev-parse HEAD > /opt/mia/hunyuan3d_21.commit \
&& rm -rf .git util/*/.git output \
&& W=/hf/hub/models--jasongzy--Make-It-Animatable/snapshots/$MIA_WEIGHTS_REV \
&& ln -s "$W/output" output \
&& ln -s "$W/data/Standard Run.fbx" "data/Standard Run.fbx" \
&& ln -s "$W/data/examples" data/examples \
&& mkdir -p data/Mixamo && ln -s "../Standard Run.fbx" data/Mixamo/bones.fbx \
&& mkdir -p /opt/mia/hy3dgen/tencent \
&& ln -s /hf/hub/models--tencent--Hunyuan3D-2.1/snapshots/$HY3D_REV /opt/mia/hy3dgen/tencent/Hunyuan3D-2.1 \
&& printf 'MIA_COMMIT=%s\nMIA_WEIGHTS_REV=%s\nHY3D_REV=%s\n' "$MIA_COMMIT" "$MIA_WEIGHTS_REV" "$HY3D_REV" > /opt/mia/provenance
COPY mia_driver.py /opt/mia/mia_driver.py
WORKDIR /job
ENTRYPOINT ["python", "/opt/mia/mia_driver.py"]
+98
View File
@@ -0,0 +1,98 @@
# mia — Make-It-Animatable v2 auto-rigger (fv-ml1 GPU 3, one-shot)
**What:** [Make-It-Animatable](https://github.com/jasongzy/Make-It-Animatable) v2 auto-rigs a
humanoid mesh (GLB in): it predicts a 52-bone Mixamo-style skeleton and skin weights. Dread Naught
(dread-dev) uses it to rig AI-generated characters. On 2026-10-01 it beat Blender bone-heat rigging
on the four game characters (Booth `dreadnaught-rigging`).
**Shape: on demand only, like blender-run (Prime, 2026-10-01).** Each `scripts/mia-run` call is
one `docker run --rm`: load the models, rig every mesh passed, write the job's `out/`, exit, hand
GPU 3 back. Nothing stays up. There is no compose service; this dir is the image's build context.
```bash
scripts/mia-run --job DIR -- goblin=in/goblin.glb ogre=in/ogre.glb # seeded (seed 0)
scripts/mia-run --job DIR --seed 7 -- goblin=in/goblin.glb
scripts/mia-run --job DIR --unseeded -- goblin=in/goblin.glb # upstream behaviour
```
Inputs are relative to DIR. Per mesh, `DIR/out/` gets `<name>_pred.npz` (with
`weights_effective`), `<name>.fbx`, `<name>.glb`, `<name>_rest.glb`, the three `_apose-hint`
variants, and `<name>_run.json` (timings, peak GPU memory, seed, provenance). The job mirroring
is blender-run's: DIR → `fv-ml1:/tank/mia/jobs/<basename>-<hash>/`, results copied back.
## What is pinned
| piece | pin | where |
|---|---|---|
| code | `jasongzy/Make-It-Animatable` branch v2 @ `bbd8b158` (MIT); submodule `util/Hunyuan3D_21` @ `b6911977` | cloned into the image, `.git` dropped |
| MIA weights | HF `jasongzy/Make-It-Animatable` @ `ca0daf6c` (Apache-2.0): `output/best/v2/{bw_joints,joints_coarse,pose}.pth`, `data/Standard Run.fbx`, `data/examples/` | fv-ml1 HF cache `/tank/aimodels/huggingface`, mounted read-only at `/hf` |
| shape VAE | HF `tencent/Hunyuan3D-2.1` @ `0b946776`: `hunyuan3d-vae-v2-1/{config.yaml,model.fp16.ckpt}` (Tencent Hunyuan community licence) | same |
| Python | `requirements.lock`: dread-dev's CPU venv verbatim, torch family swapped to cu129 (torch 2.8.0, torchvision 0.23.0, torch-cluster 1.6.3, pytorch3d 0.7.8); `bpy` 4.3.0, `trimesh` 5.1.0, `numpy` 1.26.4 | image (`/opt/mia/freeze.txt` records the install) |
| driver | `mia_driver.py` (adapted from dread-dev's `tools/mia/run_mia.py` + `finalize.py`) | image, entrypoint |
Every weight file's sha256 was checked against the HF API's LFS hash after download (2026-10-01).
`/opt/mia/provenance` in the image and `provenance` in each `_run.json` carry the three pins.
## Deviations from upstream (all deliberate)
- **The skeleton template is a substitute.** `data/Mixamo/bones.fbx` → `data/Standard Run.fbx`.
The official file is in the **gated** HF dataset `jasongzy/Mixamo`, whose terms were NOT accepted
on Prime's account; accepting them is his call. The template sets bone roll only, never joints
or weights (dread-dev). Swapping in the official file later changes FBX bone roll, not the npz.
- **Seeded by default.** Before every mesh the driver calls MIA's `fix_random(seed)` (also sets
cudnn deterministic) and resets `trimesh.util._RANDOM_DEFAULT`, which `fix_random` never reaches
(trimesh ≥ 4 samples surface points from its own module RNG). Per-mesh reset also makes a mesh's
result independent of the other meshes in the call.
- **No animation baked** (`animation_file=None`; the UI default is "Standard Run.fbx").
- **app_v2.py is unpatched.** dread-dev's CPU patch only upcasts to fp32 on CPU; the GPU runs
upstream's fp16 autocast path.
## Acceptance (2026-10-01) — instruments and raw numbers in `acceptance/`
Harness: fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, otherwise idle at 2 MiB), image
`local/mia:0.1.0`, four meshes (`dreadnaught/art/models/{goblin,villager,captain,ogre}.glb`,
2.7k–3.4k verts), three seeded calls with the mesh order rotated, plus one unseeded call.
| mesh | GPU pipeline s, median of 3 (spread) | model part s | A-pose-hint export s | CPU pipeline s (dread-dev, median of 3) |
|---|---|---|---|---|
| goblin | **4.53** (4.43–4.57) | 1.88 | 1.27 | 160.8 |
| villager | **3.97** (3.93–3.99) | 1.61 | 1.24 | 116.9 |
| captain | **3.90** (3.79–3.94) | 1.65 | 1.24 | 121.6 |
| ogre | **4.72** (4.45–4.80) | 2.08 | 1.35 | 124.7 |
- **Model load:** 12.7–13.0 s per call (4 calls), paid once per call. End to end, a 4-mesh call
takes ~50 s including ssh, rsync and container start; a 1-mesh call ~31 s.
- **Peak VRAM:** 3,394 MiB on GPU 3 as nvidia-smi saw it (100 ms polling over all four calls,
includes the CUDA context); torch's peak reserved was 2,712–2,732 MiB per mesh.
- **Repeatable:** the three seeded calls are **bit-identical** for every mesh (max abs diff 0.0 on
joints, tails, weights, weights_effective, pose), despite the rotated mesh order.
- **Matches CPU within the noise.** Distances use dread-dev's `noise_floor.py` metrics (joint
distance / height, weight mass moved, dominant bone changed). GPU-vs-CPU averages 0.83–1.27× the
CPU-vs-CPU distance, per mesh and metric. Positive control: one unseeded GPU call against a
seeded one, i.e. pure sampling noise, lands at 0.80–1.83× — same size, so the instrument sees
sampling noise and the GPU path adds nothing visible on top of it. 24 of 72 GPU-vs-CPU pair
values exceed the CPU-vs-CPU maximum, but that maximum comes from only 3 pairs per mesh, and the
sampling-only control overshoots it the same way.
- **Sensitivity floor:** this comparison cannot resolve a systematic GPU-vs-CPU offset (fp16 vs
fp32) smaller than the sampling noise, ~0.2% of mesh height for joints and ~2% weight mass moved.
A seeded CPU run would isolate it; not done (would cost ~8 min of nh3-dev CPU).
## Gotchas
- `init_blocks()` reads `data/examples/log.csv` even headless; the image symlinks `data/examples`
into the HF snapshot. Without it the driver dies before the first mesh.
- uv needs `UV_LINK_MODE=copy` in the build (reflink clone fails on overlayfs, "os error 22").
- The image is 10.5 GB on fv-ml1's zroot, which sat at 85% after the build. Check free space before
building a new tag, and remove the old tag after.
- GPU 3 is borrowed from the vLLM reserve and shared on demand with Blender and Scriberr. MIA needs
~3.4 GB, so it coexists with both; if a full-size seat claims the card, this tool has to move.
## Rebuild
```bash
# on nh3-dev
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g root /opt/docker/src/mia-<ver>'
rsync -a stacks/mia/{Dockerfile,requirements.lock,mia_driver.py} infra-ops@10.251.50.54:/opt/docker/src/mia-<ver>/
ssh infra-ops@10.251.50.54 'cd /opt/docker/src/mia-<ver> && docker build -t local/mia:<ver> .'
# then point fv-ml1:/opt/docker/compose/mia/.env IMAGE= at the new tag
```
+394
View File
@@ -0,0 +1,394 @@
{
"goblin": {
"gpu_pipeline_s": [
4.569874900858849,
4.53466919856146,
4.427527715917677
],
"gpu_model_s": [
1.9738918957300484,
1.881500325864181,
1.8626845581457019
],
"apose_hint_s": [
1.190177327953279,
1.4380995847750455,
1.2671978930011392
],
"cpu_pipeline_s": [
119.3426258880063,
708.4532103899983,
160.76836042600917
],
"peak_reserved_mib": [
2712.0,
2732.0,
2732.0,
2712.0
],
"seeded_maxdiff": {
"joints_head": 0.0,
"joints_tail": 0.0,
"weights": 0.0,
"weights_effective": 0.0,
"pose_to_rest": 0.0
},
"cpu_cpu": [
{
"jh_med": 0.002404872328042984,
"jh_max": 0.009904514066874981,
"jt_max": 0.010788610205054283,
"w_moved_mean": 0.01768580637872219,
"w_moved_p95": 0.06185286119580269,
"dom_changed": 0.02601522842639594
},
{
"jh_med": 0.0025397357530891895,
"jh_max": 0.009374394081532955,
"jt_max": 0.010417704470455647,
"w_moved_mean": 0.02369445003569126,
"w_moved_p95": 0.07524719834327698,
"dom_changed": 0.03140862944162436
},
{
"jh_med": 0.0021242855582386255,
"jh_max": 0.005268011707812548,
"jt_max": 0.0065661026164889336,
"w_moved_mean": 0.01971343532204628,
"w_moved_p95": 0.06007026135921478,
"dom_changed": 0.022208121827411168
}
],
"gpu_cpu": [
{
"jh_med": 0.002763624768704176,
"jh_max": 0.01039000041782856,
"jt_max": 0.011829185299575329,
"w_moved_mean": 0.025863967835903168,
"w_moved_p95": 0.083327516913414,
"dom_changed": 0.031725888324873094
},
{
"jh_med": 0.0024375617504119873,
"jh_max": 0.006176925264298916,
"jt_max": 0.006694309413433075,
"w_moved_mean": 0.023365769535303116,
"w_moved_p95": 0.07489541918039322,
"dom_changed": 0.02728426395939086
},
{
"jh_med": 0.0030529131181538105,
"jh_max": 0.00709641445428133,
"jt_max": 0.007498544175177813,
"w_moved_mean": 0.020390696823596954,
"w_moved_p95": 0.06874266266822815,
"dom_changed": 0.02950507614213198
}
],
"unseeded_vs_seeded": [
{
"jh_med": 0.0029595012310892344,
"jh_max": 0.007351752370595932,
"jt_max": 0.010119971819221973,
"w_moved_mean": 0.022764530032873154,
"w_moved_p95": 0.06456947326660156,
"dom_changed": 0.030139593908629442
}
],
"same_geometry": true
},
"villager": {
"gpu_pipeline_s": [
3.9700973057188094,
3.9323174457531422,
3.9949949418660253
],
"gpu_model_s": [
1.6040664857719094,
1.7117067980580032,
1.6055225860327482
],
"apose_hint_s": [
1.277365405112505,
1.0884231911040843,
1.2362516748253256
],
"cpu_pipeline_s": [
114.40334213199094,
116.87012534309179,
118.42003130697412
],
"peak_reserved_mib": [
2732.0,
2712.0,
2732.0,
2732.0
],
"seeded_maxdiff": {
"joints_head": 0.0,
"joints_tail": 0.0,
"weights": 0.0,
"weights_effective": 0.0,
"pose_to_rest": 0.0
},
"cpu_cpu": [
{
"jh_med": 0.0013478414621204138,
"jh_max": 0.003135952167212963,
"jt_max": 0.004719415679574013,
"w_moved_mean": 0.015827687457203865,
"w_moved_p95": 0.047268014401197433,
"dom_changed": 0.0183884988298228
},
{
"jh_med": 0.0016320310533046722,
"jh_max": 0.004904848523437977,
"jt_max": 0.0048230430111289024,
"w_moved_mean": 0.014390517957508564,
"w_moved_p95": 0.04615660756826401,
"dom_changed": 0.01671681711802073
},
{
"jh_med": 0.0015891867224127054,
"jh_max": 0.0036788983270525932,
"jt_max": 0.005405386444181204,
"w_moved_mean": 0.01590096764266491,
"w_moved_p95": 0.050222158432006836,
"dom_changed": 0.0183884988298228
}
],
"gpu_cpu": [
{
"jh_med": 0.001549383858218789,
"jh_max": 0.002919533057138324,
"jt_max": 0.007295706775039434,
"w_moved_mean": 0.015061692334711552,
"w_moved_p95": 0.04662124067544937,
"dom_changed": 0.01805416248746239
},
{
"jh_med": 0.0012639821507036686,
"jh_max": 0.004058246500790119,
"jt_max": 0.005211320705711842,
"w_moved_mean": 0.016790034249424934,
"w_moved_p95": 0.050929658114910126,
"dom_changed": 0.01905717151454363
},
{
"jh_med": 0.0015581324696540833,
"jh_max": 0.0054196459241211414,
"jt_max": 0.005114810075610876,
"w_moved_mean": 0.01441498938947916,
"w_moved_p95": 0.046087272465229034,
"dom_changed": 0.014042126379137413
}
],
"unseeded_vs_seeded": [
{
"jh_med": 0.0020237602293491364,
"jh_max": 0.00419682078063488,
"jt_max": 0.005229325499385595,
"w_moved_mean": 0.015298573300242424,
"w_moved_p95": 0.044845279306173325,
"dom_changed": 0.01905717151454363
}
],
"same_geometry": true
},
"captain": {
"gpu_pipeline_s": [
3.902754598064348,
3.792409762972966,
3.9422666197642684
],
"gpu_model_s": [
1.637031597085297,
1.6480189089197665,
1.741625044029206
],
"apose_hint_s": [
1.3538039769046009,
1.244770233053714,
1.139452091883868
],
"cpu_pipeline_s": [
125.6441692209919,
121.16770973894745,
121.56508291099453
],
"peak_reserved_mib": [
2732.0,
2732.0,
2712.0,
2732.0
],
"seeded_maxdiff": {
"joints_head": 0.0,
"joints_tail": 0.0,
"weights": 0.0,
"weights_effective": 0.0,
"pose_to_rest": 0.0
},
"cpu_cpu": [
{
"jh_med": 0.0017565949819982052,
"jh_max": 0.00641083437949419,
"jt_max": 0.0061119236052036285,
"w_moved_mean": 0.01863820292055607,
"w_moved_p95": 0.06350292265415192,
"dom_changed": 0.021580102414045354
},
{
"jh_med": 0.0020125731825828552,
"jh_max": 0.003998186439275742,
"jt_max": 0.006467505358159542,
"w_moved_mean": 0.017755920067429543,
"w_moved_p95": 0.05389416217803955,
"dom_changed": 0.019385515727871252
},
{
"jh_med": 0.0018191682174801826,
"jh_max": 0.0042957765981554985,
"jt_max": 0.0039120386354625225,
"w_moved_mean": 0.02013971656560898,
"w_moved_p95": 0.057384975254535675,
"dom_changed": 0.019385515727871252
}
],
"gpu_cpu": [
{
"jh_med": 0.001694112434051931,
"jh_max": 0.0057935346849262714,
"jt_max": 0.004407331347465515,
"w_moved_mean": 0.01909082941710949,
"w_moved_p95": 0.06022503599524498,
"dom_changed": 0.016093635698610095
},
{
"jh_med": 0.0016945323441177607,
"jh_max": 0.005641191732138395,
"jt_max": 0.006225286982953548,
"w_moved_mean": 0.018217289820313454,
"w_moved_p95": 0.0654626339673996,
"dom_changed": 0.019019751280175568
},
{
"jh_med": 0.0017832419835031033,
"jh_max": 0.003895824309438467,
"jt_max": 0.00681948009878397,
"w_moved_mean": 0.018107891082763672,
"w_moved_p95": 0.05214930698275566,
"dom_changed": 0.023043160204828092
}
],
"unseeded_vs_seeded": [
{
"jh_med": 0.001529946457594633,
"jh_max": 0.003934196196496487,
"jt_max": 0.005103697534650564,
"w_moved_mean": 0.022645380347967148,
"w_moved_p95": 0.07639797031879425,
"dom_changed": 0.024140453547915143
}
],
"same_geometry": true
},
"ogre": {
"gpu_pipeline_s": [
4.71703689894639,
4.802805577637628,
4.453627999639139
],
"gpu_model_s": [
2.076565559953451,
2.098935986869037,
1.9066526568494737
],
"apose_hint_s": [
1.3545660751406103,
1.3703913709614426,
1.3345670951530337
],
"cpu_pipeline_s": [
123.53674800007138,
124.70388515893137,
336.82813009101665
],
"peak_reserved_mib": [
2732.0,
2732.0,
2732.0,
2732.0
],
"seeded_maxdiff": {
"joints_head": 0.0,
"joints_tail": 0.0,
"weights": 0.0,
"weights_effective": 0.0,
"pose_to_rest": 0.0
},
"cpu_cpu": [
{
"jh_med": 0.002341177314519882,
"jh_max": 0.004295716993510723,
"jt_max": 0.010711956769227982,
"w_moved_mean": 0.019047463312745094,
"w_moved_p95": 0.05714641511440277,
"dom_changed": 0.0243612596553773
},
{
"jh_med": 0.0027327858842909336,
"jh_max": 0.007534473203122616,
"jt_max": 0.011530852876603603,
"w_moved_mean": 0.021840794011950493,
"w_moved_p95": 0.06544578820466995,
"dom_changed": 0.024955436720142603
},
{
"jh_med": 0.0020891886670142412,
"jh_max": 0.007218110840767622,
"jt_max": 0.007911363616585732,
"w_moved_mean": 0.021831553429365158,
"w_moved_p95": 0.06872576475143433,
"dom_changed": 0.023767082590612002
}
],
"gpu_cpu": [
{
"jh_med": 0.002296903170645237,
"jh_max": 0.005948834586888552,
"jt_max": 0.007127378601580858,
"w_moved_mean": 0.023190144449472427,
"w_moved_p95": 0.06846748292446136,
"dom_changed": 0.029411764705882353
},
{
"jh_med": 0.0035880799405276775,
"jh_max": 0.0077993483282625675,
"jt_max": 0.009203977882862091,
"w_moved_mean": 0.02328675426542759,
"w_moved_p95": 0.07298485934734344,
"dom_changed": 0.030303030303030304
},
{
"jh_med": 0.0032017994672060013,
"jh_max": 0.006755590904504061,
"jt_max": 0.008556965738534927,
"w_moved_mean": 0.020125865936279297,
"w_moved_p95": 0.05771474540233612,
"dom_changed": 0.026737967914438502
}
],
"unseeded_vs_seeded": [
{
"jh_med": 0.0030766851268708706,
"jh_max": 0.01164184045046568,
"jt_max": 0.013770738616585732,
"w_moved_mean": 0.023339038714766502,
"w_moved_p95": 0.06661652028560638,
"dom_changed": 0.035056446821152706
}
],
"same_geometry": true
}
}
+55
View File
@@ -0,0 +1,55 @@
"""MIA GPU acceptance analysis: timings (median of 3 seeded runs), determinism, and GPU-vs-CPU
agreement measured with dread-dev's noise_floor.py metrics against the CPU run-to-run floor."""
import itertools, json, statistics as st, sys, numpy as np
A, C = sys.argv[1], sys.argv[2]
MESHES = ["goblin", "villager", "captain", "ogre"]
def load(p):
d = np.load(p)
if "weights_effective" not in d:
raise SystemExit(f"{p}: no weights_effective")
return d
def metrics(Ar, Br):
H = float(np.ptp(Ar["verts"][:, 1]))
dj = np.linalg.norm(Ar["joints_head"] - Br["joints_head"], axis=1) / H
dt = np.linalg.norm(Ar["joints_tail"] - Br["joints_tail"], axis=1) / H
WA, WB = Ar["weights_effective"], Br["weights_effective"]
moved = 0.5 * np.abs(WA - WB).sum(1)
return dict(jh_med=float(np.median(dj)), jh_max=float(dj.max()), jt_max=float(dt.max()),
w_moved_mean=float(moved.mean()), w_moved_p95=float(np.percentile(moved, 95)),
dom_changed=float((WA.argmax(1) != WB.argmax(1)).mean()))
def rng(ms, k):
v = [m[k] for m in ms]
return f"{min(v):.4f}..{max(v):.4f}"
KEYS = ["jh_med", "jh_max", "jt_max", "w_moved_mean", "w_moved_p95", "dom_changed"]
out = {}
for n in MESHES:
g = {r: load(f"{A}/{r}/out/{n}_pred.npz") for r in ("r1", "r2", "r3", "u1")}
c = [load(f"{C}/{n}_pred.npz"), load(f"{C}/repeats/{n}_rep2_pred.npz"), load(f"{C}/repeats/{n}_rep3_pred.npz")]
runs = {r: json.load(open(f"{A}/{r}/out/{n}_run.json")) for r in ("r1", "r2", "r3", "u1")}
same_geom = all(np.array_equal(g["r1"]["verts"], x["verts"]) and np.array_equal(g["r1"]["faces"], x["faces"]) for x in c)
# determinism across the 3 seeded runs (different mesh orders): exact?
det = {k: max(float(np.abs(g["r1"][k] - g[r][k]).max()) for r in ("r2", "r3"))
for k in ("joints_head", "joints_tail", "weights", "weights_effective", "pose_to_rest")}
cpu_cpu = [metrics(a, b) for a, b in itertools.combinations(c, 2)] # the CPU noise floor (null)
gpu_cpu = [metrics(g["r1"], x) for x in c] # the claim under test
unseeded = [metrics(g["u1"], g[r]) for r in ("r1",)] # positive control: must move
t = [runs[r]["pipeline_total_wall_s"] for r in ("r1", "r2", "r3")]
tm = [runs[r]["model_total_wall_s"] for r in ("r1", "r2", "r3")]
ta = [runs[r]["timings"]["vis_blender_apose_hint"]["wall_s"] for r in ("r1", "r2", "r3")]
pk = [runs[r]["gpu_peak_reserved_mib"] for r in ("r1", "r2", "r3", "u1")]
cpu_t = [json.load(open(p))["pipeline_total_wall_s"] for p in (f"{C}/{n}_run.json", f"{C}/repeats/{n}_rep2_run.json", f"{C}/repeats/{n}_rep3_run.json")]
print(f"\n== {n}: verts {runs['r1']['n_verts']}, same geometry as CPU: {same_geom}")
print(f" GPU pipeline s: {[round(x,2) for x in t]} median {st.median(t):.2f} | model part median {st.median(tm):.2f} | A-pose-hint export median {st.median(ta):.2f} | CPU pipeline median {st.median(cpu_t):.1f}")
print(f" peak torch reserved MiB: {pk}")
print(f" seeded r1/r2/r3 max abs diff: " + ", ".join(f"{k} {v:.2e}" for k, v in det.items()))
for k in KEYS:
print(f" {k:13s} CPU-vs-CPU {rng(cpu_cpu,k)} | GPU-vs-CPU {rng(gpu_cpu,k)} | unseeded-vs-seeded GPU {rng(unseeded,k)}")
out[n] = dict(gpu_pipeline_s=t, gpu_model_s=tm, apose_hint_s=ta, cpu_pipeline_s=cpu_t, peak_reserved_mib=pk,
seeded_maxdiff=det, cpu_cpu=cpu_cpu, gpu_cpu=gpu_cpu, unseeded_vs_seeded=unseeded, same_geometry=same_geom)
init = [json.load(open(f"{A}/{r}/out/goblin_run.json"))["init_wall_s"] for r in ("r1", "r2", "r3", "u1")]
print("\nmodel load (init) s per container:", [round(x, 1) for x in init])
json.dump(out, open(f"{A}/acceptance.json", "w"), indent=1)
+17
View File
@@ -0,0 +1,17 @@
#!/usr/bin/env bash
# MIA GPU acceptance: 3 seeded runs (mesh order rotated) + 1 unseeded run, GPU 3 memory polled throughout.
set -uo pipefail
# Run from a scratch dir holding r1/ r2/ r3/ u1/, each with in/{goblin,villager,captain,ogre}.glb;
# then: python3 compare.py <that dir> <dread-dev CPU out/ dir>.
A=$(dirname "$(realpath "$0")")
cd ~/development/eshpfi-management
H=infra-ops@10.251.50.54
POLLPID=$(ssh -n $H 'nohup nvidia-smi -i 3 --query-gpu=timestamp,memory.used --format=csv,noheader,nounits -lms 100 > /tank/mia/vram-accept.csv 2>/dev/null & echo $!')
echo "poller pid $POLLPID"
run() { local r=$1; shift; scripts/mia-run --job "$A/$r" "$@" > "$A/$r.log" 2>&1; local rc=$?; echo "$r rc=$rc"; }
run r1 -- goblin=in/goblin.glb villager=in/villager.glb captain=in/captain.glb ogre=in/ogre.glb
run r2 -- villager=in/villager.glb captain=in/captain.glb ogre=in/ogre.glb goblin=in/goblin.glb
run r3 -- captain=in/captain.glb ogre=in/ogre.glb goblin=in/goblin.glb villager=in/villager.glb
run u1 --unseeded -- goblin=in/goblin.glb villager=in/villager.glb captain=in/captain.glb ogre=in/ogre.glb
ssh -n $H "kill $POLLPID"
rsync -a $H:/tank/mia/vram-accept.csv "$A/"
File diff suppressed because it is too large Load Diff
+262
View File
@@ -0,0 +1,262 @@
"""Fleet driver for Make-It-Animatable v2 (app_v2.py): one container run, any number of meshes.
Usage (inside the local/mia image; scripts/mia-run is the caller-facing wrapper):
python mia_driver.py [--seed N | --unseeded] <name>=<input.glb> [<name>=<input.glb> ...]
Inputs are paths relative to the working directory (the job dir). Per mesh it writes to out/:
<name>_pred.npz predictions in the INPUT file's coordinates, plus `weights_effective`
<name>.fbx / .glb MIA's FBX (Blender export) and its FBX2glTF preview
<name>_rest.glb Blender glTF export of the rigged rest pose
<name>_apose-hint.* the same predictions, Blender stage re-run with "Input Rest Pose = A-pose"
<name>_run.json stage timings, peak GPU memory, seed, provenance
Adapted from dread-dev's tools/mia/run_mia.py + finalize.py (dreadnaught repo, 2026-10-01). The
pipeline calls, the npz keys and the artifact set are theirs, unchanged. Differences:
- Paths: repo at /opt/mia/repo; MIA's scratch (the input copy and the work dir it writes beside
it) lives in a temp dir that dies with the container, so only out/ comes back to the caller.
- SEEDED BY DEFAULT. Before EVERY mesh: util.utils.fix_random(seed) (python/numpy/torch seeds,
cudnn deterministic), and trimesh's module RNG reset to default_rng(seed). trimesh >= 4 samples
surface points from trimesh.util._RANDOM_DEFAULT, which fix_random never reaches; unseeded runs
moved joints 0.1-0.27% of height (dread-dev's CPU noise floor). Resetting per mesh also makes a
mesh's result independent of which meshes ran before it in the same process.
--unseeded restores upstream behaviour (fix_random once at model load, trimesh unseeded).
- weights_effective is written straight into the npz (finalize.py's formula: weights > 1e-3 kept,
rows renormalised; what the FBX vertex groups and the GLB WEIGHTS_0 actually carry).
- Each stage is followed by torch.cuda.synchronize() so the per-stage wall times are honest.
"""
import argparse
import json
import os
import re
import shutil
import sys
import tempfile
import time
REPO = "/opt/mia/repo"
JOB = os.getcwd()
OUT_DIR = os.path.join(JOB, "out")
NAME_RE = re.compile(r"^[A-Za-z0-9._-]+$")
def parse_args():
ap = argparse.ArgumentParser(description=__doc__.split("\n")[0])
g = ap.add_mutually_exclusive_group()
g.add_argument("--seed", type=int, default=0, help="seed for every mesh (default 0)")
g.add_argument("--unseeded", action="store_true", help="upstream behaviour: trimesh sampling unseeded")
ap.add_argument("meshes", nargs="+", metavar="NAME=INPUT.glb")
a = ap.parse_args()
jobs = []
for m in a.meshes:
if "=" not in m:
ap.error(f"{m!r}: expected NAME=INPUT.glb")
name, path = m.split("=", 1)
if not NAME_RE.match(name):
ap.error(f"{name!r}: NAME must be [A-Za-z0-9._-]+")
src = os.path.realpath(os.path.join(JOB, path))
if not src.startswith(JOB + os.sep):
ap.error(f"{path!r}: inputs must be inside the job dir")
if not os.path.isfile(src):
ap.error(f"{path!r}: no such file in the job dir")
jobs.append((name, src))
if len({n for n, _ in jobs}) != len(jobs):
ap.error("NAMEs must be unique")
return a, jobs
args, JOBS = parse_args()
SEED = None if args.unseeded else args.seed
os.environ.setdefault("GRADIO_ANALYTICS_ENABLED", "False")
os.environ["HF_HUB_OFFLINE"] = "1" # all weights are on the read-only /hf mount; never touch the network
sys.path.insert(0, REPO)
os.chdir(REPO)
import numpy as np # noqa: E402
import torch # noqa: E402
import trimesh.util # noqa: E402
t0 = time.perf_counter()
import app_v2 as A # noqa: E402
from util.utils import fix_random # noqa: E402
A.init_models()
A.init_blocks() # only defines the Gradio component globals the stage functions return as dict keys; no server
t_init = time.perf_counter() - t0
CUDA = torch.cuda.is_available()
print(f"[mia] init (imports + 3 models + blocks): {t_init:.1f}s, device {'cuda:' + torch.cuda.get_device_name(0) if CUDA else 'cpu'}, seed {SEED}", flush=True)
STAGES = ["prepare_input", "preprocess", "infer", "vis", "vis_blender", "finish"]
BONE_NAMES = [None] * len(A.BONES_IDX_DICT)
for k, v in A.BONES_IDX_DICT.items():
BONE_NAMES[v] = k
PARENTS = list(A.KINEMATIC_TREE.parent_indices)
PROVENANCE = dict(line.strip().split("=", 1) for line in open("/opt/mia/provenance") if "=" in line)
def sync():
if CUDA:
torch.cuda.synchronize()
def to_np(x):
if isinstance(x, torch.Tensor):
x = x.detach().cpu().numpy()
return np.asarray(x)
def run_one(name: str, src: str, scratch: str):
if SEED is not None:
fix_random(SEED)
trimesh.util._RANDOM_DEFAULT = np.random.default_rng(SEED)
if CUDA:
torch.cuda.reset_peak_memory_stats()
inp = os.path.join(scratch, f"{name}.glb")
shutil.copyfile(src, inp) # MIA writes its work dir next to the input file, so run on a copy
work = os.path.join(scratch, name)
db = A.DB()
gen = A._pipeline(
input_path=inp,
is_gs=False,
opacity_threshold=0.01,
no_fingers=False, # UI default
rest_pose_type="No", # UI default
ignore_pose_parts=[], # UI default
input_normal=True, # UI default (fixed)
bw_fix=True, # UI default: weight post-processing on
bw_vis_bone="LeftArm", # UI default (visualisation only)
restore_global=False, # UI default: outputs in MIA's normalised frame
reset_to_rest=True, # UI default: apply predicted T-pose as the rest pose
animation_file=None, # deviation from UI default ("Standard Run.fbx"): static rig, no animation baked
retarget=True,
inplace=True,
db=db,
)
timings = {}
snap = {}
for stage in STAGES:
w0 = time.perf_counter()
next(gen)
sync()
timings[stage] = {"wall_s": time.perf_counter() - w0}
print(f"[mia] {name}: {stage} {timings[stage]['wall_s']:.2f}s", flush=True)
if stage == "prepare_input":
snap["verts_input"] = to_np(db.verts)[0].copy()
snap["faces"] = np.asarray(db.faces).copy()
elif stage == "infer":
snap["bw_raw"] = to_np(db.bw)[0].copy()
try:
next(gen)
except StopIteration:
pass
# --- predictions in the INPUT file's coordinates ---
T = db.global_transform # pytorch3d Transform3d (row-vector): input -> MIA-normalised frame
Tinv = T.inverse()
T_mat = T.get_matrix().transpose(-1, -2)[0].cpu().numpy() # column-vector 4x4
Tinv_mat = Tinv.get_matrix().transpose(-1, -2)[0].cpu().numpy()
joints_n = np.asarray(db.joints, dtype=np.float32)
tails_n = np.asarray(db.joints_tail, dtype=np.float32)
dev = Tinv.device
joints_in = Tinv.transform_points(torch.from_numpy(joints_n)[None].to(dev))[0].cpu().numpy()
tails_in = Tinv.transform_points(torch.from_numpy(tails_n)[None].to(dev))[0].cpu().numpy()
verts_n = np.asarray(db.verts, dtype=np.float32)
verts_back = Tinv.transform_points(torch.from_numpy(verts_n)[None].to(dev))[0].cpu().numpy()
roundtrip_err = float(np.abs(verts_back - snap["verts_input"]).max())
pose_n = np.asarray(db.pose, dtype=np.float32)
pose_in = np.einsum("ij,kjl,lm->kim", Tinv_mat, pose_n, T_mat).astype(np.float32)
W = np.asarray(db.bw, dtype=np.float32)
We = np.where(W > 1e-3, W, 0).astype(np.float32) # finalize.py: MIA's set_weights threshold
We /= np.maximum(We.sum(1, keepdims=True), 1e-12)
np.savez_compressed(
os.path.join(OUT_DIR, f"{name}_pred.npz"),
verts=snap["verts_input"].astype(np.float32),
faces=snap["faces"].astype(np.int64),
weights=W,
weights_raw=snap["bw_raw"].astype(np.float32),
weights_effective=We,
bone_names=np.array(BONE_NAMES),
parents=np.array(PARENTS, dtype=np.int64),
joints_head=joints_in.astype(np.float32),
joints_tail=tails_in.astype(np.float32),
pose_to_rest=pose_in,
input_to_mia=T_mat.astype(np.float32),
joints_head_mia=joints_n,
joints_tail_mia=tails_n,
verts_mia=verts_n,
)
# --- MIA's own artifacts (default settings) ---
copies = {
f"{name}.fbx": db.anim_path, # MIA's primary output (Blender FBX export)
f"{name}.glb": db.anim_vis_path, # MIA's FBX2glTF conversion of the FBX (the app's "GLB preview")
f"{name}_rest.glb": db.rest_vis_path, # Blender glTF export of the rigged rest-pose model
}
missing = []
for dst, srcp in copies.items():
if srcp and os.path.isfile(srcp):
shutil.copyfile(srcp, os.path.join(OUT_DIR, dst))
else:
missing.append(dst)
# --- variant: same predictions, Blender stage re-run with the UI's "Input Rest Pose = A-pose" hint ---
w0 = time.perf_counter()
db.anim_path = os.path.join(work, f"{name}_apose-hint.fbx")
db.anim_vis_path = os.path.join(work, f"{name}_apose-hint.glb")
db.rest_vis_path = os.path.join(work, f"{name}_apose-hint_rest.glb")
A.vis_blender(
reset_to_rest=True,
remove_fingers=False,
rest_pose_type="A-pose",
ignore_pose_parts=[],
animation_file=None,
retarget=True,
inplace=True,
restore_global=False,
db=db,
)
timings["vis_blender_apose_hint"] = {"wall_s": time.perf_counter() - w0}
for suffix in ("_apose-hint.fbx", "_apose-hint.glb", "_apose-hint_rest.glb"):
srcp = os.path.join(work, f"{name}{suffix}")
if os.path.isfile(srcp):
shutil.copyfile(srcp, os.path.join(OUT_DIR, f"{name}{suffix}"))
else:
missing.append(f"{name}{suffix}")
meta = {
"name": name,
"source": os.path.relpath(src, JOB),
"n_verts": int(snap["verts_input"].shape[0]),
"n_faces": int(snap["faces"].shape[0]),
"seed": SEED,
"device": torch.cuda.get_device_name(0) if CUDA else "cpu",
"timings": timings,
"model_total_wall_s": sum(timings[s]["wall_s"] for s in ("preprocess", "infer")),
"pipeline_total_wall_s": sum(timings[s]["wall_s"] for s in STAGES),
"init_wall_s": t_init,
"gpu_peak_allocated_mib": torch.cuda.max_memory_allocated() / 2**20 if CUDA else None,
"gpu_peak_reserved_mib": torch.cuda.max_memory_reserved() / 2**20 if CUDA else None,
"roundtrip_err_input_to_mia_and_back": roundtrip_err,
"missing_artifacts": missing,
"provenance": PROVENANCE,
}
with open(os.path.join(OUT_DIR, f"{name}_run.json"), "w") as f:
json.dump(meta, f, indent=2)
print(f"[mia] {name}: done, pipeline {meta['pipeline_total_wall_s']:.1f}s, peak reserved "
f"{meta['gpu_peak_reserved_mib'] or 0:.0f} MiB, roundtrip err {roundtrip_err:.2e}"
+ (f", MISSING {missing}" if missing else ""), flush=True)
return meta
if __name__ == "__main__":
os.makedirs(OUT_DIR, exist_ok=True)
failed = 0
with tempfile.TemporaryDirectory(prefix="mia-") as scratch:
for name, src in JOBS:
meta = run_one(name, src, scratch)
failed += bool(meta["missing_artifacts"])
sys.exit(1 if failed else 0)
+113
View File
@@ -0,0 +1,113 @@
# MIA v2 Python lock: dread-dev's CPU venv (nh3-dev, 2026-10-01) verbatim, except the four
# torch-family lines, swapped from +cpu to the cu129 builds upstream's requirements.txt names.
accelerate==1.9.0
annotated-doc==0.0.5
annotated-types==0.8.0
antlr4-python3-runtime==4.9.3
anyio==4.15.1
attrs==26.1.0
bpy==4.3.0
brotli==1.2.0
certifi==2026.7.22
charset-normalizer==3.5.2
click==8.5.0
colorlog==6.12.0
contourpy==1.3.3
cycler==0.12.1
cython==3.3.0
diffusers==0.39.0
einops==0.8.2
embreex==4.4.0
fastapi==0.142.2
filelock==3.32.3
fonttools==4.66.1
fsspec==2026.7.0
gradio==6.17.3
gradio-client==2.5.0
groovy==0.1.2
h11==0.16.0
hf-gradio==0.4.1
hf-xet==1.6.0
httpcore==1.0.9
httpx==0.28.1
huggingface-hub==0.36.2
idna==3.20
imageio==2.38.0
importlib-metadata==9.0.1
iopath==0.1.10
jinja2==3.1.6
jsonschema==4.26.0
jsonschema-specifications==2025.9.1
kiwisolver==1.5.1
lazy-loader==0.6
lxml==6.1.3
manifold3d==3.5.4
mapbox-earcut==2.1.0
markdown-it-py==4.2.0
markupsafe==3.0.3
matplotlib==3.11.2
mdurl==0.1.2
mpmath==1.3.0
networkx==3.6.1
numpy==1.26.4
omegaconf==2.3.1
opencv-contrib-python==4.11.0.86
opentelemetry-api==1.45.0
orjson==3.12.0
packaging==26.3
pandas==3.0.6
pillow==12.3.0
plyfile==1.1.3
portalocker==4.4.0
potpourri3d==1.4.0
psutil==7.2.2
pycollada==0.9.3
pydantic==2.13.5
pydantic-core==2.46.5
pydub==0.25.1
pygments==2.21.0
pymcubes==0.1.6
pymeshlab==2023.12.post3
pyparsing==3.3.3
python-dateutil==2.9.0.post0
python-multipart==0.0.32
pytorch3d==0.7.8+pt2.8.0cu129
pytz==2026.4
pyyaml==6.0.3
referencing==0.37.0
regex==2026.9.29
requests==2.34.2
rich==15.0.0
rpds-py==2026.6.3
rtree==1.4.1
safehttpx==0.1.7
safetensors==0.8.0
scikit-image==0.26.0
scipy==1.17.1
semantic-version==2.10.0
shapely==2.1.2
shellingham==1.5.4
six==1.17.0
spaces==0.51.3
starlette==1.7.0
svg-path==7.1
sympy==1.14.0
tifffile==2026.3.3
timm==1.0.30
tokenizers==0.20.3
tomlkit==0.14.0
torch==2.8.0+cu129
torch-cluster==1.6.3+pt28cu129
torchvision==0.23.0+cu129
tqdm==4.70.1
transformers==4.46.0
trimesh==5.1.0
typer==0.27.2
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.8.0
uvicorn==0.54.0
vhacdx==0.1.0
xxhash==4.0.1
zipp==4.1.0
zstandard==0.25.0