feat(erp-tune): run 6 on pfi-gx10 — jenerallee78 ARA-abliterated base (index 33c59654) pulled + byte-verified, run-5 recipe byte-held, launched under operator-2026-09-08-rnd-run6

- scripts/erp-tune-gx10/pull-verify-jenerallee78.sh + base-pin-jenerallee78-shards.txt:
  revision-pinned root-shard pull, 32/32 sha256+size vs brokkr-smithy pins, index
  set-equal to stock, STOCK tokenizer set installed over the repo's (which bakes in
  a 256-token truncation); repo originals kept as *.repo
- scripts/erp-tune-gx10/run-06-gx10.json + launch-run-06.sh: run-05 config with the
  base swapped, recipe-r6, survivors-r5 verbatim, stock template path
- docs/runbooks/gx10-run-06.md: pull/verify record, free-check result (encode
  reproduces run 5 exactly), hf download --include gotcha, gate naming
  (erp-seat-base-ara / erp-tune-v6)
This commit is contained in:
2026-09-08 04:24:38 -07:00
parent 55631e28bc
commit 3fec668bf2
6 changed files with 299 additions and 1 deletions
+13 -1
View File
@@ -11,9 +11,14 @@ The harness itself (`eitri-smithy`) is not vendored here; it lives on the box at
| `run-05-gx10.json` | `/home/infra-ops/erp-tune/run-05-gx10.json` | 5 |
| `launch-run-05.sh` | `/home/infra-ops/erp-tune/launch-run-05.sh` | 5 |
| `build_r5_survivors.py` | `/home/infra-ops/erp-tune/build_r5_survivors.py` | 5 |
| `run-06-gx10.json` | `/home/infra-ops/erp-tune/run-06-gx10.json` | 6 |
| `launch-run-06.sh` | `/home/infra-ops/erp-tune/launch-run-06.sh` | 6 |
| `pull-verify-jenerallee78.sh` | `/home/infra-ops/erp-tune/pull-verify-jenerallee78.sh` | 6 (base pull + byte verify) |
| `base-pin-jenerallee78-shards.txt` | `/home/infra-ops/erp-tune/base-pin-jenerallee78-shards.txt` | 6 (32 shard pins, from brokkr-smithy `base-pin-jenerallee78.json`) |
Runbooks: [`docs/runbooks/gx10-run-03c.md`](../../docs/runbooks/gx10-run-03c.md),
[`docs/runbooks/gx10-run-05.md`](../../docs/runbooks/gx10-run-05.md).
[`docs/runbooks/gx10-run-05.md`](../../docs/runbooks/gx10-run-05.md),
[`docs/runbooks/gx10-run-06.md`](../../docs/runbooks/gx10-run-06.md).
**Run 3c** — the LoRA that died on ana-ml2 at step 24 when an Anaheim breaker
tripped, rehomed here unchanged (eight path keys rehomed to local NVMe, two host/
@@ -25,6 +30,13 @@ value differs, verified key-by-key).
everything else held; kvasir byte-identical (survivors reused from run 4).
`run-05-gx10.json` is run 4's config with recipe/survivors/override swapped.
**Run 6** — the run-5 recipe byte-held on a different BASE: jenerallee78's ARA
abliteration of Gemma-4-26B-A4B-it @ `0631379a` (index sha `33c59654…`), the
line's first abliterated base (runs 35 were settled as stock). Corpus,
survivors (`survivors-r5.jsonl`), mask, template and hyperparameters unchanged.
The landed dir carries the STOCK tokenizer set (the repo's `tokenizer.json`
bakes in a 256-token truncation); repo originals kept beside as `*.repo`.
> **Run 4 is not vendored here.** It ran on the box (config `run-04-gx10.json`,
> gated STILL-COUPLED) but its canonical copies were never committed; run 5's
> `build_r5_survivors.py` derives from `survivors-r4.jsonl` on the box, so run 4
@@ -0,0 +1,32 @@
model-00001-of-00032.safetensors cb38d992e7292af270c76c5ad89d582b9be170bc2ddb15d3320ebe0505d05977 1990394256
model-00002-of-00032.safetensors d5e92288b94df9c607bf31c8bdef80a16ea6b59807ca81644943cec769bd3fa0 1628192850
model-00003-of-00032.safetensors 2302ffb7482cacab78b12ced015696e00042fbeb0634595c3389904295fd85a6 1628192850
model-00004-of-00032.safetensors f3f41bb1e7d81587dbb60b2f8c3062dd70cb56e83abfd89b6f6be242d78043e7 1628192850
model-00005-of-00032.safetensors 13564f050bd4878736fed1da0d39bc4e09283247c83f2ef1f7657b9094f849e5 1628192850
model-00006-of-00032.safetensors 6bd2f3dc341f5afacb567c7fce6aeac839b8dee3eba862d2a4fae31b07085b89 1628192850
model-00007-of-00032.safetensors dbdb67891bae9b4d0cc964956aa3a44f286528754c418128b4670788c909dc2f 1657029578
model-00008-of-00032.safetensors c12f29da42d1308b9e2aaf487d8b0d7c4e21968c81c0a4558a673dc1c750c540 1628192850
model-00009-of-00032.safetensors e003f08afeb765e620f78c3384001844bbcb0cb05aa5b52fc674faccab0d9c14 1628192850
model-00010-of-00032.safetensors a9b6311add23b28a38cdb52e04895be345fd7c02799895933b6e10a5ecb98466 1628192850
model-00011-of-00032.safetensors aa6a373c5b367ff93f5849c53c3de1f27f2163d3e63f2bff2ee3aaa296d739c2 1628192842
model-00012-of-00032.safetensors 0d2c639e0225c14f0eb82bab59e3a375109e4cf6b40169208032ca54dea0209e 1628192866
model-00013-of-00032.safetensors 4b476fe09d8e52a08d85f354bce18e03aebb7569d1d343c27490d878aa831d79 1657029602
model-00014-of-00032.safetensors cf2a6cb084e498b3576a262c79fa21f40000d529702f99ab17fd08ca66ffd768 1628192866
model-00015-of-00032.safetensors 2eb5cf8a58d24419e8d206b9f4a6c87900ea1c098557dec9cd5bd27e04bf246f 1628192866
model-00016-of-00032.safetensors 1023217e914a724069656925e1957fc32cab0ff98eb2f1b6b1d978da71b79baf 1628192866
model-00017-of-00032.safetensors ac28d02bdd63d2ef8178d30993339fa65656c36e09503424522461a5fa9fc62b 1628192866
model-00018-of-00032.safetensors eba4791821709bcb6bf9462d52ddf193752ab71e197da8e74a79324808ab10ee 1628192866
model-00019-of-00032.safetensors 3f7eff449d7d59eaf447cb85f0e8948950c820c7598c78e64cd7c3e544733da6 1657029602
model-00020-of-00032.safetensors 1d5bbd51267175bde103092ba2cdc4609e90eaab21edbcc7dbb2f3826544c702 1628192866
model-00021-of-00032.safetensors ed40710cd36d74200663e0dfe18db659a764b8e2510897a5b1c6d1f269eaeb31 1628192866
model-00022-of-00032.safetensors 0e341e75c659828897be00a0236fbe141480e1ed1b06694ac53fdefb5201bad0 1628192866
model-00023-of-00032.safetensors 4be775b953a330eb6d2ccacfd003bcc976e33005e87a7998537b16f10ed1a255 1628192866
model-00024-of-00032.safetensors bf2156ec7cc0389873198f4fb88622365f0c69dee5d3a0bf003489ac9c6172c4 1628192866
model-00025-of-00032.safetensors 9bc722adedc5b9042f4e976b8ff657f1a5e2e7c2b5689ba03ff8a5f3f2ac6ad0 1657029602
model-00026-of-00032.safetensors 86096e378a7cd9254ab95dd17557c3960d50d9ef230febf40cd7cb3fd8c76b68 1628192866
model-00027-of-00032.safetensors 9dc27c40b43d42459cdf222102580e90f4a431116692d8b765bef56f9c2c70b9 1628192866
model-00028-of-00032.safetensors aabccf617bfc00f86581e70646bac5aa7b9be564a34e263b8c7080623f7b4933 1628192866
model-00029-of-00032.safetensors 5e3e4090b6c5fec39ca7a694899386910fc1a31f8c11e2de122099414b144cd0 1628192866
model-00030-of-00032.safetensors a424e02c63531f8b4efc62e390e3c5db8680509c04772b4425c63e09db5cf184 1628192866
model-00031-of-00032.safetensors ce11bf78b3f19cfd0814779f14aa7aab0dfbfa5eae7392823d9ffa6ada43ff7d 1997452570
model-00032-of-00032.safetensors f4ed47cc36b78bc3b4720b96a362a6d2e1eae60a57d33d1937b27b7468970778 291222376
+72
View File
@@ -0,0 +1,72 @@
#!/usr/bin/env bash
# Launch ERP-seat SFT run 6 on pfi-gx10 (NVIDIA GB10, aarch64, sm_121).
#
# Run this ON pfi-gx10 as infra-ops. It detaches the job from the invoking
# shell and logs to the box, so a reaped SSH session cannot take the run with
# it -- the failure mode that lost the first probe launch on 2026-09-01.
#
# Run 6 = run 5 recipe UNCHANGED on the jenerallee78 ARA-abliterated base (the single variable).
# 8,212 survivors -> 524 optimizer steps, encode must match run 5 exactly.
# Checkpoints every 50 steps.
set -euo pipefail
ROOT=/home/infra-ops/erp-tune
HARNESS=$ROOT/eitri-smithy
VENV=/home/infra-ops/ml/.venv/bin/python
CONFIG=$ROOT/run-06-gx10.json
LOG=$ROOT/run-06.log
# --- Preconditions, asserted rather than assumed -----------------------------
# A stuck orphan holding unified memory while PyTorch reports zero allocated
# already doomed three relaunches on this box and got blamed on the new run
# each time. Assert the GPU is clear.
apps=$(nvidia-smi --query-compute-apps=pid --format=csv,noheader | tr -d '[:space:]')
if [ -n "$apps" ]; then
echo "REFUSING: GPU is not clear -- compute apps still resident:" >&2
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv >&2
exit 1
fi
# Deliberately NOT `pgrep -f erp_sft_harness`: run this over ssh and the
# pattern appears in the invoking shell's own argv, so the guard matches
# itself and refuses every launch. The pidfile is exact and cannot self-match;
# the GPU assertion above catches an orphan under any name.
if [ -f "$ROOT/run-06.pid" ] && kill -0 "$(cat "$ROOT/run-06.pid")" 2>/dev/null; then
echo "REFUSING: run-06.pid names a live process $(cat "$ROOT/run-06.pid"):" >&2
ps -p "$(cat "$ROOT/run-06.pid")" -o pid,etime,cmd >&2
exit 1
fi
if [ -e "$LOG" ]; then
echo "REFUSING: $LOG exists. Move it aside first so two runs cannot share a log." >&2
exit 1
fi
for p in "$HARNESS/erp_sft_harness/__main__.py" "$VENV" "$CONFIG"; do
[ -e "$p" ] || { echo "REFUSING: missing $p" >&2; exit 1; }
done
# Free space for checkpoints, with headroom.
avail=$(df --output=avail -BG "$ROOT" | tail -1 | tr -dc '0-9')
if [ "$avail" -lt 40 ]; then
echo "REFUSING: only ${avail}G free under $ROOT; want >=40G for checkpoints." >&2
exit 1
fi
# --- Launch ------------------------------------------------------------------
cd "$HARNESS"
{
echo "# launched $(date -Is) on $(hostname) by ${USER}"
echo "# harness $(git rev-parse --short HEAD) config $CONFIG"
} > "$LOG"
setsid nohup "$VENV" -m erp_sft_harness --config "$CONFIG" >> "$LOG" 2>&1 < /dev/null &
pid=$!
echo "$pid" > "$ROOT/run-06.pid"
echo "launched pid $pid -> $LOG"
echo
echo "watch: tail -f $LOG | tr '\\r' '\\n'"
echo "stop: kill \$(cat $ROOT/run-06.pid) # by PID -- never pkill -f over ssh"
@@ -0,0 +1,54 @@
#!/usr/bin/env bash
# Pull + verify jenerallee78/gemma-4-26B-A4B-it-ara-abliterated @ 0631379a onto pfi-gx10.
# Root shards + small files only; no GGUFs, no mlx-4bit. Verifies bytes against the
# brokkr-smithy pins (base-pin-jenerallee78.json) and installs the STOCK tokenizer set.
# Expects ~/erp-tune/base-pin-jenerallee78-shards.txt (file sha256 bytes per line).
set -uo pipefail
export PATH="$HOME/.local/bin:$PATH"
REPO=jenerallee78/gemma-4-26B-A4B-it-ara-abliterated
REV=0631379a3d859e0059bc8d9b21ab5b654dfc272c
DEST=$HOME/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a
STOCK=$HOME/models/gemma4-26b-a4b-it-bf16
PINS=$HOME/erp-tune/base-pin-jenerallee78-shards.txt
echo "== start $(date -u +%FT%TZ) on $(hostname)"
mkdir -p "$DEST"
echo "== download"
uv run --quiet --with 'huggingface_hub[hf_transfer]' hf download "$REPO" --revision "$REV" \
--local-dir "$DEST" \
--include 'model-*-of-00032.safetensors' --include 'config.json' --include 'generation_config.json' \
--include 'model.safetensors.index.json' --include 'chat_template.jinja' --include 'ara_config.json' --include 'README.md' \
--include 'tokenizer.json' --include 'tokenizer_config.json'
rc=$?
echo "== download rc=$rc $(date -u +%FT%TZ)"
[ $rc -eq 0 ] || { echo "DOWNLOAD FAILED rc=$rc"; exit 2; }
cd "$DEST"
echo "== shard sha256 vs pins"
fail=0
while read -r f oid bytes; do
[ -n "$f" ] || continue
sz=$(stat -c %s "$f" 2>/dev/null || echo MISSING)
got=$(sha256sum "$f" 2>/dev/null | cut -d' ' -f1)
if [ "$sz" = "$bytes" ] && [ "$got" = "$oid" ]; then echo "OK $f"; else echo "FAIL $f size=$sz want=$bytes sha=$got want=$oid"; fail=$((fail+1)); fi
done < "$PINS"
echo "== shard result: fail=$fail"
echo "== index + config checks"
python3 - "$STOCK" <<'PY'
import json,sys,hashlib
stock=sys.argv[1]
d=json.load(open('model.safetensors.index.json'));s=json.load(open(f'{stock}/model.safetensors.index.json'))
names=set(d['weight_map']);snames=set(s['weight_map'])
print('index weight_map:',len(names),'stock:',len(snames),'set_equal:',names==snames)
print('index total_size:',d['metadata'].get('total_size'),'stock:',s['metadata'].get('total_size'),'equal:',d['metadata'].get('total_size')==s['metadata'].get('total_size'))
c=json.load(open('config.json'))
print('config architectures:',c.get('architectures'),'dtype:',c.get('dtype') or c.get('torch_dtype'))
print('INDEX_SHA256', hashlib.sha256(open('model.safetensors.index.json','rb').read()).hexdigest())
PY
echo "== repo tokenizer set as shipped (kept aside as *.repo)"
sha256sum tokenizer.json tokenizer_config.json chat_template.jinja
python3 -c "import json;print('repo tokenizer.json truncation:',json.load(open('tokenizer.json')).get('truncation'))"
for f in tokenizer.json tokenizer_config.json chat_template.jinja; do mv -n "$f" "$f.repo"; cp "$STOCK/$f" "$f"; done
echo "== STOCK tokenizer set installed (sha256):"
sha256sum tokenizer.json tokenizer_config.json chat_template.jinja
python3 -c "import json;print('installed tokenizer.json truncation:',json.load(open('tokenizer.json')).get('truncation'))"
echo "== listing"; ls -la "$DEST"; du -sh "$DEST"
echo "== done $(date -u +%FT%TZ) shard_fail=$fail"
+47
View File
@@ -0,0 +1,47 @@
{
"output_dir": "/home/infra-ops/erp-tune/run-06",
"roots_dir": "/home/infra-ops/erp-tune/datasets/derived",
"base_model_path": "/home/infra-ops/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a",
"base_model_revision": "jenerallee78/gemma-4-26B-A4B-it-ara-abliterated @ 0631379a3d859e0059bc8d9b21ab5b654dfc272c (ARA abliteration of stock google/gemma-4-26B-A4B-it; 32 bf16 root shards sha256-verified against brokkr-smithy base-pin-jenerallee78.json; index sha256 33c59654e658a30fa29cdc87ccd6a752bfa0bb3e32cd56f95ff1eb82075e593a). THE SINGLE VARIABLE vs run 5: base only. Run-5 recipe, survivors, mask, template, hyperparameters all UNCHANGED. Tokenizer set = STOCK (tokenizer.json cc8d3a0c / tokenizer_config.json 9f4fec4b / chat_template.jinja ae53464b) copied over the repo's, whose shipped tokenizer.json carries a baked-in max_length=256 truncation; repo originals kept beside as *.repo. Runs 3/3c/4/5 were settled 2026-09-08 as STOCK base (index 907826a6), so this is the line's first abliterated base.",
"recipe": "/home/infra-ops/erp-tune/recipe-r6/recipe-erp-seat-sft-r6.json",
"survivors": "/home/infra-ops/erp-tune/recipe-r5/survivors-r5.jsonl",
"chat_template_path": "/home/infra-ops/models/gemma4-26b-a4b-it-bf16/chat_template.jinja",
"impersonation_mask_path": "/home/infra-ops/erp-tune/recipe-r3/lossmask-r3.jsonl",
"lora_rank": 64,
"lora_alpha": 128,
"lora_dropout": 0.0,
"max_seq_len": 16384,
"epochs": 1,
"seed": 20260824,
"per_device_batch_size": 2,
"gradient_accumulation_steps": 8,
"learning_rate": 0.0002,
"warmup_ratio": 0.1,
"lr_scheduler_type": "cosine",
"weight_decay": 0.01,
"load_in_4bit": false,
"gradient_checkpointing": true,
"loss_chunk_tokens": 1024,
"training_eligibility_override": "operator-2026-09-08-rnd-run6",
"overridden_blockers": [
"contamination-scan-not-implemented",
"stage-2-csam-detector-inert"
],
"substitute_controls": [
"pre-training holdout, run-1 (8,404 samples, work/card/session split)",
"pre-training holdout, govreport/holdout-v1 (416 reports, sha256-ranked, never_trained_on)",
"pre-training holdout, qmsum/holdout-v1 (5 transcripts, sha256-ranked, never_trained_on)",
"stage-A lexical quarantine, RP (829 records held unread)",
"stage-A lexical quarantine, run-5 slot (133 records held unread, /mnt/smithy/datasets/quarantine/r47-run5-longdep-screen/)",
"SCROLLS-membership disclosure on both slot sources (avoidance, NOT a scan): govreport + qmsum are SCROLLS/ZeroSCROLLS members, in no hoard/default-benchmarks.yaml entry and used by no R47 instrument",
"kvasir is HELD, not re-cut: the 1,613 kvasir survivors are reused verbatim from survivors-r4.jsonl (which cut run-3's seed-20260824 prefix at 3,347,622 ctx). survivors-r5.jsonl = survivors-r4 minus airoboros plus the govreport + qmsum roots whole; sha256 a25169a6258cd4abb0cb494a176a921c0e98eb73d65c53d033b6ee18293a43ae.",
"window_count belt-and-suspenders (SFT-RECIPE-run5-SCOPE.md 7.1): every govreport + qmsum row renders <= 14,000 tokens (max 9,385 / 13,700) and the harness never packs across samples, so window_count MUST be 1 on every slot row; a chunked_into_2 or single_window_truncated on either new root in truncation-report.json is a BUILD DEFECT and the run is killed before training.",
"SINGLE VARIABLE vs run 5: the BASE. Stock google/gemma-4-26B-A4B-it OUT, jenerallee78 ARA abliteration @ 0631379a IN. Corpus (survivors-r5 rows verbatim), impersonation loss-mask (lossmask-r3), stock chat template ae53464b, lr 2e-04, max_seq_len 16384, rank 64, alpha 128, dropout 0.0, cosine, warmup 0.1, wd 0.01, batch 2 x accum 8, 1 epoch, seed 20260824 all UNCHANGED from run 5.",
"FREE CHECK (brokkr-smithy, 2026-09-08): the [encode] pass must reproduce run 5 EXACTLY -- 8,370 records, ctx 18,598,779 tok, loss 9,935,076 tok -- because corpus, tokenizer and template are identical; any difference means the wrong tokenizer/template loaded and the run is killed before [train].",
"HOST: pfi-gx10 (GB10, aarch64, sm_121, 121 GB unified). Base shards sha256-verified against the revision-pinned HF LFS oids after landing; harness eitri-smithy 0a6bd2e; corpus COPIED, box mounts no NFS. Grant: operator directive to infra-ops in-session 2026-09-08 (\"unload the gx10 and commence training on the gx10. window is open now.\").",
"SURVIVORS: recipe-r6 ships no survivor list of its own (targets byte-identical to r5), so survivors-r5.jsonl (sha256 a25169a6...) is reused verbatim; recipe-erp-seat-sft-r6.json sha256 6499555471181bd8ef7a273340f162e44fe54bf4f070198f050c8be33ee769a8 from brokkr-smithy 4dd7590."
],
"unfittable": "drop",
"holdout_dir": "/home/infra-ops/erp-tune/datasets/holdout",
"save_steps": 50
}