feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)

Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
This commit is contained in:
vh
2026-10-01 09:45:22 -07:00
parent e5536784e0
commit f44280e6b2
12 changed files with 3121 additions and 1 deletions
+85
View File
@@ -0,0 +1,85 @@
#!/usr/bin/env bash
# mia-run — one-shot Make-It-Animatable v2 auto-rig on fv-ml1 GPU 3 (dread-dev / Dread Naught;
# Prime's request, 2026-10-01). Same shape as blender-run: a container starts, rigs, exits, and
# hands the GPU back. NOT a seat; nothing stays up between calls.
#
# scripts/mia-run --job DIR [--seed N | --unseeded] -- <name>=<input.glb> [<name>=<input.glb> ...]
# scripts/mia-run --job ~/development/dreadnaught/rig-jobs/j1 -- goblin=in/goblin.glb ogre=in/ogre.glb
#
# Inputs are paths RELATIVE TO DIR and must be inside it: DIR is the only thing shipped to fv-ml1.
# Per mesh, DIR/out/ receives:
# <name>_pred.npz predictions in the input's coordinates, WITH `weights_effective`
# (the finalize.py formula), so no finalize step is needed afterwards
# <name>.fbx .glb _rest.glb MIA's own exports (UI defaults, static rig)
# <name>_apose-hint.fbx .glb _rest.glb same predictions, Blender stage with the A-pose hint
# <name>_run.json stage timings, model-load time, peak GPU memory, seed, provenance
# The model load (~3 x 1 GB checkpoints) is paid once per call, so pass every mesh in one call.
#
# SEEDED BY DEFAULT (seed 0): before every mesh the driver calls MIA's fix_random(seed) and resets
# trimesh's module RNG, which fix_random never reaches (trimesh >= 4 samples surfaces from
# trimesh.util._RANDOM_DEFAULT). Without that, runs of the same mesh differ: joints by 0.1-0.27% of
# height on dread-dev's CPU noise floor. --unseeded restores upstream behaviour.
#
# Files: fv-ml1 does NOT mount /mnt/smithy. DIR is mirrored to
# fv-ml1:/tank/mia/jobs/<basename>-<hash of DIR's absolute path>/ (rsync --delete, so a rerun never
# inherits leftovers), the driver runs WITH THAT AS ITS WORKING DIRECTORY, and new or changed files
# are copied back into DIR afterwards (nothing is ever deleted locally). Running the SAME DIR twice
# concurrently shares one remote dir, so don't. Remote job dirs are never swept.
#
# Image: local/mia (stacks/mia; IMAGE= in fv-ml1:/opt/docker/compose/mia/.env). Weights come from
# the read-only /tank/aimodels/huggingface mount at pinned revisions; the run never touches the
# network (HF_HUB_OFFLINE=1). The skeleton template is a SUBSTITUTE ("Standard Run.fbx"): the
# official one is in a gated HF dataset whose terms were not accepted (see stacks/mia/README.md).
#
# Budget: GPU 3 (96 GB, borrowed from the vLLM reserve, shared on demand with Blender and Scriberr;
# this ends if a full-size seat moves in). Capped at 32 GB RAM / 16 CPUs so a runaway mesh cannot
# starve the inference seats on the same host.
set -euo pipefail
HOST=${MIA_SSH_HOST:-infra-ops@10.251.50.54}
ENV_FILE=/opt/docker/compose/mia/.env
JOB=""
DRIVER_OPTS=()
while [ $# -gt 0 ]; do
case "$1" in
--job) JOB=${2:?--job needs a directory}; shift 2 ;;
--seed) DRIVER_OPTS+=(--seed "${2:?--seed needs an integer}"); shift 2 ;;
--unseeded) DRIVER_OPTS+=(--unseeded); shift ;;
--) shift; break ;;
-h|--help) sed -n 2,40p "$0"; exit 0 ;;
*) echo "mia-run: unknown option $1 (meshes go after --)" >&2; exit 2 ;;
esac
done
[ -n "$JOB" ] || { echo "mia-run: --job DIR is required (inputs are read from it, out/ is written into it)" >&2; exit 2; }
[ $# -gt 0 ] || { echo "mia-run: no meshes; pass <name>=<input.glb> after --" >&2; exit 2; }
[ -d "$JOB" ] || { echo "mia-run: --job $JOB is not a directory" >&2; exit 2; }
ABS=$(realpath "$JOB")
BASE=$(basename "$ABS")
[[ $BASE =~ ^[A-Za-z0-9._-]+$ ]] || { echo "mia-run: job dir name '$BASE' must be [A-Za-z0-9._-]" >&2; exit 2; }
for m in "$@"; do
[[ $m == *=* ]] || { echo "mia-run: '$m' is not <name>=<input.glb>" >&2; exit 2; }
f=$(realpath -m "$ABS/${m#*=}")
[[ $f == "$ABS"/* && -f $f ]] || { echo "mia-run: input '${m#*=}' is not a file inside $ABS" >&2; exit 2; }
done
# Keyed on basename + a hash of the ABSOLUTE local path, as blender-run: two dirs that share a
# basename never share a remote dir. rsync removes only inside that one remote dir; there is no rm
# on a computed path.
NAME=$BASE-$(printf '%s' "$ABS" | sha256sum | cut -c1-8)
ssh -n -o BatchMode=yes "$HOST" "mkdir -p /tank/mia/jobs/$NAME"
rsync -a --delete "$JOB"/ "$HOST:/tank/mia/jobs/$NAME/"
# Arguments travel as one shell-quoted string: ssh flattens argv into a remote command line.
ARGS=$(printf '%q ' "${DRIVER_OPTS[@]}" "$@")
set +e
ssh -n -o BatchMode=yes "$HOST" "IMG=\$(grep '^IMAGE=' $ENV_FILE | cut -d= -f2) && \
exec docker run --rm --name mia-run-\$\$ --hostname fv-ml1-mia \
--runtime nvidia -e NVIDIA_VISIBLE_DEVICES=3 -e NVIDIA_DRIVER_CAPABILITIES=compute,utility \
--user 1002:1003 --memory 32g --cpus 16 \
--mount type=bind,src=/tank/aimodels/huggingface,dst=/hf,readonly \
--mount type=bind,src=/tank/mia/jobs/$NAME,dst=/job \
-w /job \"\$IMG\" $ARGS"
RC=$?
set -e
rsync -a --update "$HOST:/tank/mia/jobs/$NAME/" "$JOB"/
exit $RC