dae77ee118
Rehome run 3c from ana-ml2 to pfi-gx10 unchanged: same corpus, base,
recipe and hyperparameters, different host. Slower (~13.3 h vs ~2.5 h)
and correct — an Anaheim breaker trip costs a 40-minute drive each way
and 13 hosts down, three of them SureFire client machines, while the
GX10 is a ~240 W appliance at NH3 that can take nothing else down.
Verified rather than assumed, because ana-ml2 ran transformers 5.15.1
on x86-64 and this box runs 5.16.1 on aarch64 — the silent
backend-delta class that has already voided conclusions here:
- both 49 GB base shards sha256-match ana-ml2's (size equality is a
weaker claim and was already true)
- a full encode was run into a throwaway dir and the encoded corpus
compared byte-for-byte: 197,360,233 B, sha256 c08bb1fe2ecb0be3,
identical. Every aggregate matched too. That verified artifact is
what the run will train on — it is seeded into run-03c/encode-cache
- the harness's own suite: 122 passed on aarch64
- the config generator asserts key-by-key that no non-path value
differs from run-03c.json
The encode-cache filename differs by design (base_model_path is part of
the key) — an input hash, not an output hash. Documented so it is not
misread as drift, or "fixed" by faking /tank on this box.
Corpus is copied to local NVMe; the box mounts no NFS. nh3-nas is now on
the same subnet, which makes mounting it tempting and still wrong under
a 13 h unattended run.
The launcher refuses on a live pidfile rather than a pgrep: `pgrep -f
erp_sft_harness` invoked over ssh matches the invoking shell's own argv.
That self-match cost a shell during staging.
Not launched. 13.3 h is the operator's call.
74 lines
2.8 KiB
Bash
Executable File
74 lines
2.8 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Launch ERP-seat SFT run 3c on pfi-gx10 (NVIDIA GB10, aarch64, sm_121).
|
|
#
|
|
# Run this ON pfi-gx10 as infra-ops. It detaches the job from the invoking
|
|
# shell and logs to the box, so a reaped SSH session cannot take the run with
|
|
# it -- the failure mode that lost the first probe launch on 2026-09-01.
|
|
#
|
|
# Expected: 604 optimizer steps at ~79.4 s/it => ~13.3 h.
|
|
# Checkpoints every 50 steps, ~852 MB each (~10 GB total).
|
|
set -euo pipefail
|
|
|
|
ROOT=/home/infra-ops/erp-tune
|
|
HARNESS=$ROOT/eitri-smithy
|
|
VENV=/home/infra-ops/ml/.venv/bin/python
|
|
CONFIG=$ROOT/run-03c-gx10.json
|
|
LOG=$ROOT/run-03c.log
|
|
|
|
# --- Preconditions, asserted rather than assumed -----------------------------
|
|
|
|
# A stuck orphan holding unified memory while PyTorch reports zero allocated
|
|
# already doomed three relaunches on this box and got blamed on the new run
|
|
# each time. Assert the GPU is clear.
|
|
apps=$(nvidia-smi --query-compute-apps=pid --format=csv,noheader | tr -d '[:space:]')
|
|
if [ -n "$apps" ]; then
|
|
echo "REFUSING: GPU is not clear -- compute apps still resident:" >&2
|
|
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv >&2
|
|
exit 1
|
|
fi
|
|
|
|
# Deliberately NOT `pgrep -f erp_sft_harness`: run this over ssh and the
|
|
# pattern appears in the invoking shell's own argv, so the guard matches
|
|
# itself and refuses every launch. Same self-match that makes `pkill -f`
|
|
# unsafe over ssh. The pidfile is exact and cannot self-match; the GPU
|
|
# assertion above catches an orphan under any name.
|
|
if [ -f "$ROOT/run-03c.pid" ] && kill -0 "$(cat "$ROOT/run-03c.pid")" 2>/dev/null; then
|
|
echo "REFUSING: run-03c.pid names a live process $(cat "$ROOT/run-03c.pid"):" >&2
|
|
ps -p "$(cat "$ROOT/run-03c.pid")" -o pid,etime,cmd >&2
|
|
exit 1
|
|
fi
|
|
|
|
if [ -e "$LOG" ]; then
|
|
echo "REFUSING: $LOG exists. Move it aside first so two runs cannot share a log." >&2
|
|
exit 1
|
|
fi
|
|
|
|
for p in "$HARNESS/erp_sft_harness/__main__.py" "$VENV" "$CONFIG"; do
|
|
[ -e "$p" ] || { echo "REFUSING: missing $p" >&2; exit 1; }
|
|
done
|
|
|
|
# Free space for checkpoints: 12 x 852 MB + final adapter, with headroom.
|
|
avail=$(df --output=avail -BG "$ROOT" | tail -1 | tr -dc '0-9')
|
|
if [ "$avail" -lt 40 ]; then
|
|
echo "REFUSING: only ${avail}G free under $ROOT; want >=40G for checkpoints." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- Launch ------------------------------------------------------------------
|
|
|
|
cd "$HARNESS"
|
|
{
|
|
echo "# launched $(date -Is) on $(hostname) by ${USER}"
|
|
echo "# harness $(git rev-parse --short HEAD) config $CONFIG"
|
|
} > "$LOG"
|
|
|
|
setsid nohup "$VENV" -m erp_sft_harness --config "$CONFIG" >> "$LOG" 2>&1 < /dev/null &
|
|
pid=$!
|
|
echo "$pid" > "$ROOT/run-03c.pid"
|
|
|
|
echo "launched pid $pid -> $LOG"
|
|
echo
|
|
echo "watch: tail -f $LOG | tr '\\r' '\\n'"
|
|
echo "steps: grep -ao '[0-9]*/604 \[[^]]*\]' $LOG | tail -1"
|
|
echo "stop: kill \$(cat $ROOT/run-03c.pid) # by PID -- never pkill -f over ssh"
|