From 3fec668bf24c1a294250d65c68684a0b1facd2df Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Tue, 8 Sep 2026 04:24:38 -0700 Subject: [PATCH] =?UTF-8?q?feat(erp-tune):=20run=206=20on=20pfi-gx10=20?= =?UTF-8?q?=E2=80=94=20jenerallee78=20ARA-abliterated=20base=20(index=2033?= =?UTF-8?q?c59654)=20pulled=20+=20byte-verified,=20run-5=20recipe=20byte-h?= =?UTF-8?q?eld,=20launched=20under=20operator-2026-09-08-rnd-run6?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - scripts/erp-tune-gx10/pull-verify-jenerallee78.sh + base-pin-jenerallee78-shards.txt: revision-pinned root-shard pull, 32/32 sha256+size vs brokkr-smithy pins, index set-equal to stock, STOCK tokenizer set installed over the repo's (which bakes in a 256-token truncation); repo originals kept as *.repo - scripts/erp-tune-gx10/run-06-gx10.json + launch-run-06.sh: run-05 config with the base swapped, recipe-r6, survivors-r5 verbatim, stock template path - docs/runbooks/gx10-run-06.md: pull/verify record, free-check result (encode reproduces run 5 exactly), hf download --include gotcha, gate naming (erp-seat-base-ara / erp-tune-v6) --- docs/runbooks/gx10-run-06.md | 81 +++++++++++++++++++ scripts/erp-tune-gx10/README.md | 14 +++- .../base-pin-jenerallee78-shards.txt | 32 ++++++++ scripts/erp-tune-gx10/launch-run-06.sh | 72 +++++++++++++++++ .../erp-tune-gx10/pull-verify-jenerallee78.sh | 54 +++++++++++++ scripts/erp-tune-gx10/run-06-gx10.json | 47 +++++++++++ 6 files changed, 299 insertions(+), 1 deletion(-) create mode 100644 docs/runbooks/gx10-run-06.md create mode 100644 scripts/erp-tune-gx10/base-pin-jenerallee78-shards.txt create mode 100644 scripts/erp-tune-gx10/launch-run-06.sh create mode 100644 scripts/erp-tune-gx10/pull-verify-jenerallee78.sh create mode 100644 scripts/erp-tune-gx10/run-06-gx10.json diff --git a/docs/runbooks/gx10-run-06.md b/docs/runbooks/gx10-run-06.md new file mode 100644 index 0000000..72e26d3 --- /dev/null +++ b/docs/runbooks/gx10-run-06.md @@ -0,0 +1,81 @@ +# pfi-gx10 — ERP-seat SFT run 6 (abliterated base) + +Launched 2026-09-08 04:17 PDT (11:17:43Z) on pfi-gx10, pid 4100375. Grant: the +operator's direct in-session directive to infra-ops — *"unload the gx10 and +commence training on the gx10. window is open now."* — recorded on both sides as +`operator-2026-09-08-rnd-run6` (brokkr-smithy `TRAINING-ELIGIBILITY-OVERRIDE-run6.md`). + +## What run 6 is + +Run 5's recipe **byte-held** on a different base. The single variable is the +base: `jenerallee78/gemma-4-26B-A4B-it-ara-abliterated` @ +`0631379a3d859e0059bc8d9b21ab5b654dfc272c` (ARA 2-pass abliteration of stock +`google/gemma-4-26B-A4B-it`, layers 13–24, o_proj + down_proj). Runs 3/3c/4/5 were +settled from bytes on 2026-09-08 as having trained on **stock** (index sha +`907826a6…`) despite the `-heretic` name; this is the line's first genuinely +abliterated base. Pick and pins: brokkr-smithy +`research/R47-premium-corpus-gate/ABLITERATED-BASE-HUNT-2026-09-08.md` + +`base-pin-jenerallee78.json`; recipe `recipe-erp-seat-sft-r6.json` (sha +`64995554…`, brokkr-smithy `4dd7590`). + +## Base pull + verify (what `pull-verify-jenerallee78.sh` did) + +Landed at `/home/infra-ops/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a` +— named for the bytes, never for the intent (the lesson of `-heretic-bf16`). + +- Root shards + small files only, revision-pinned; the two root GGUFs, mmproj + and `mlx-4bit/` were not pulled. ~143 MB/s, 32 shards in ~7 min. +- Registry cross-check from nh3-dev first: HF tree API at the pinned revision, + all 32 LFS oids + sizes == pins. +- After landing: every shard's sha256 AND size == pin (32/32); index + `weight_map` set-equal to stock's 1013 names; `total_size` 51,611,872,412 == + stock; `config.json` Gemma4ForConditionalGeneration / bfloat16. +- **Base identity (index sha256): `33c59654e658a30fa29cdc87ccd6a752bfa0bb3e32cd56f95ff1eb82075e593a`.** +- ⚠ **Tokenizer hazard (brokkr, measured):** the repo's `tokenizer.json` ships with + `"truncation": {"max_length": 256}` baked in — vocab identical to stock, but loaded + as shipped it silently cuts every text past 256 tokens and the `window_count` guard + would not notice. The STOCK three were copied over it (repo originals kept as + `*.repo`), re-hashed in the landed dir: + `tokenizer.json cc8d3a0c…` / `tokenizer_config.json 9f4fec4b…` / + `chat_template.jinja ae53464b…` (the July stock template runs 3–5 used; the + repo's is the older April one, `2dfbfc7d…`). +- ⚠ `hf download` gotcha: multiple patterns after one `--include` are parsed as + explicit FILENAMES and the include is silently ignored ("Fetching 8 files"). Use + one `--include` per pattern. Attempt 1 landed 62 MB and failed verify 32/32; + attempt 2 is the recorded one. + +## Config + +`run-06-gx10.json` = `run-05-gx10.json` with `base_model_path` → the landed dir, +`recipe` → `recipe-r6/`, `survivors` → `recipe-r5/survivors-r5.jsonl` verbatim +(r6 ships no survivor list; same bytes, sha `a25169a6…`), `chat_template_path` +→ the stock file (same path as run 5), `output_dir` → `run-06`, override → +`operator-2026-09-08-rnd-run6`. Hyperparameters, mask (`lossmask-r3`), seed all +unchanged. + +## Free check — passed exactly + +Same corpus + same tokenizer + same template ⇒ the encode must reproduce run 5: +`[encode] 8,197 samples -> 8,370 records; ctx 18,598,779 tok, loss 9,935,076 tok`, +`[mix]` shares identical to four places, govreport 496/496 and qmsum 97/97 +`fit_whole`, 0 chunked / 0 truncated. Any difference = wrong tokenizer/template → +kill before `[train]`. Encode-cache filename differs by design +(`base_model_path` is in the key). + +## Launch / watch / stop + + ssh infra-ops@10.100.50.60 '~/erp-tune/launch-run-06.sh' + ssh infra-ops@10.100.50.60 "tr '\r' '\n' < ~/erp-tune/run-06.log | tail" + ssh infra-ops@10.100.50.60 'kill $(cat ~/erp-tune/run-06.pid)' # by PID — never pkill -f over ssh + +The `erp-tune-v5` seat (`vllm-run05.pid`) was stopped to clear the GPU; the +LiteLLM `trial` alias is dark until the next serve. + +## After the adapter lands — gate choreography (brokkr-smithy-dev, cc channel) + +Preregistered before any data: cells TRANSFERRED / COUPLED-HERE / FLAT on **this +base's own floors, never stock's**. Naming is load-bearing for Brokkr's pipelines: +serve the abliterated base as **`erp-seat-base-ara`** (`erp-seat-base` means +stock), the merged arm as **`erp-tune-v6`**. Same stack/flags as run 5 (bf16, +max-model-len 8192, max-num-seqs 8, gpu-util 0.60, gemma4 tool parser, template +`ae53464b`). Base floors → lock → swap cue → tuned arm. Hands-off through both. diff --git a/scripts/erp-tune-gx10/README.md b/scripts/erp-tune-gx10/README.md index ec02989..45f94f0 100644 --- a/scripts/erp-tune-gx10/README.md +++ b/scripts/erp-tune-gx10/README.md @@ -11,9 +11,14 @@ The harness itself (`eitri-smithy`) is not vendored here; it lives on the box at | `run-05-gx10.json` | `/home/infra-ops/erp-tune/run-05-gx10.json` | 5 | | `launch-run-05.sh` | `/home/infra-ops/erp-tune/launch-run-05.sh` | 5 | | `build_r5_survivors.py` | `/home/infra-ops/erp-tune/build_r5_survivors.py` | 5 | +| `run-06-gx10.json` | `/home/infra-ops/erp-tune/run-06-gx10.json` | 6 | +| `launch-run-06.sh` | `/home/infra-ops/erp-tune/launch-run-06.sh` | 6 | +| `pull-verify-jenerallee78.sh` | `/home/infra-ops/erp-tune/pull-verify-jenerallee78.sh` | 6 (base pull + byte verify) | +| `base-pin-jenerallee78-shards.txt` | `/home/infra-ops/erp-tune/base-pin-jenerallee78-shards.txt` | 6 (32 shard pins, from brokkr-smithy `base-pin-jenerallee78.json`) | Runbooks: [`docs/runbooks/gx10-run-03c.md`](../../docs/runbooks/gx10-run-03c.md), -[`docs/runbooks/gx10-run-05.md`](../../docs/runbooks/gx10-run-05.md). +[`docs/runbooks/gx10-run-05.md`](../../docs/runbooks/gx10-run-05.md), +[`docs/runbooks/gx10-run-06.md`](../../docs/runbooks/gx10-run-06.md). **Run 3c** — the LoRA that died on ana-ml2 at step 24 when an Anaheim breaker tripped, rehomed here unchanged (eight path keys rehomed to local NVMe, two host/ @@ -25,6 +30,13 @@ value differs, verified key-by-key). everything else held; kvasir byte-identical (survivors reused from run 4). `run-05-gx10.json` is run 4's config with recipe/survivors/override swapped. +**Run 6** — the run-5 recipe byte-held on a different BASE: jenerallee78's ARA +abliteration of Gemma-4-26B-A4B-it @ `0631379a` (index sha `33c59654…`), the +line's first abliterated base (runs 3–5 were settled as stock). Corpus, +survivors (`survivors-r5.jsonl`), mask, template and hyperparameters unchanged. +The landed dir carries the STOCK tokenizer set (the repo's `tokenizer.json` +bakes in a 256-token truncation); repo originals kept beside as `*.repo`. + > **Run 4 is not vendored here.** It ran on the box (config `run-04-gx10.json`, > gated STILL-COUPLED) but its canonical copies were never committed; run 5's > `build_r5_survivors.py` derives from `survivors-r4.jsonl` on the box, so run 4 diff --git a/scripts/erp-tune-gx10/base-pin-jenerallee78-shards.txt b/scripts/erp-tune-gx10/base-pin-jenerallee78-shards.txt new file mode 100644 index 0000000..3e9c180 --- /dev/null +++ b/scripts/erp-tune-gx10/base-pin-jenerallee78-shards.txt @@ -0,0 +1,32 @@ +model-00001-of-00032.safetensors cb38d992e7292af270c76c5ad89d582b9be170bc2ddb15d3320ebe0505d05977 1990394256 +model-00002-of-00032.safetensors d5e92288b94df9c607bf31c8bdef80a16ea6b59807ca81644943cec769bd3fa0 1628192850 +model-00003-of-00032.safetensors 2302ffb7482cacab78b12ced015696e00042fbeb0634595c3389904295fd85a6 1628192850 +model-00004-of-00032.safetensors f3f41bb1e7d81587dbb60b2f8c3062dd70cb56e83abfd89b6f6be242d78043e7 1628192850 +model-00005-of-00032.safetensors 13564f050bd4878736fed1da0d39bc4e09283247c83f2ef1f7657b9094f849e5 1628192850 +model-00006-of-00032.safetensors 6bd2f3dc341f5afacb567c7fce6aeac839b8dee3eba862d2a4fae31b07085b89 1628192850 +model-00007-of-00032.safetensors dbdb67891bae9b4d0cc964956aa3a44f286528754c418128b4670788c909dc2f 1657029578 +model-00008-of-00032.safetensors c12f29da42d1308b9e2aaf487d8b0d7c4e21968c81c0a4558a673dc1c750c540 1628192850 +model-00009-of-00032.safetensors e003f08afeb765e620f78c3384001844bbcb0cb05aa5b52fc674faccab0d9c14 1628192850 +model-00010-of-00032.safetensors a9b6311add23b28a38cdb52e04895be345fd7c02799895933b6e10a5ecb98466 1628192850 +model-00011-of-00032.safetensors aa6a373c5b367ff93f5849c53c3de1f27f2163d3e63f2bff2ee3aaa296d739c2 1628192842 +model-00012-of-00032.safetensors 0d2c639e0225c14f0eb82bab59e3a375109e4cf6b40169208032ca54dea0209e 1628192866 +model-00013-of-00032.safetensors 4b476fe09d8e52a08d85f354bce18e03aebb7569d1d343c27490d878aa831d79 1657029602 +model-00014-of-00032.safetensors cf2a6cb084e498b3576a262c79fa21f40000d529702f99ab17fd08ca66ffd768 1628192866 +model-00015-of-00032.safetensors 2eb5cf8a58d24419e8d206b9f4a6c87900ea1c098557dec9cd5bd27e04bf246f 1628192866 +model-00016-of-00032.safetensors 1023217e914a724069656925e1957fc32cab0ff98eb2f1b6b1d978da71b79baf 1628192866 +model-00017-of-00032.safetensors ac28d02bdd63d2ef8178d30993339fa65656c36e09503424522461a5fa9fc62b 1628192866 +model-00018-of-00032.safetensors eba4791821709bcb6bf9462d52ddf193752ab71e197da8e74a79324808ab10ee 1628192866 +model-00019-of-00032.safetensors 3f7eff449d7d59eaf447cb85f0e8948950c820c7598c78e64cd7c3e544733da6 1657029602 +model-00020-of-00032.safetensors 1d5bbd51267175bde103092ba2cdc4609e90eaab21edbcc7dbb2f3826544c702 1628192866 +model-00021-of-00032.safetensors ed40710cd36d74200663e0dfe18db659a764b8e2510897a5b1c6d1f269eaeb31 1628192866 +model-00022-of-00032.safetensors 0e341e75c659828897be00a0236fbe141480e1ed1b06694ac53fdefb5201bad0 1628192866 +model-00023-of-00032.safetensors 4be775b953a330eb6d2ccacfd003bcc976e33005e87a7998537b16f10ed1a255 1628192866 +model-00024-of-00032.safetensors bf2156ec7cc0389873198f4fb88622365f0c69dee5d3a0bf003489ac9c6172c4 1628192866 +model-00025-of-00032.safetensors 9bc722adedc5b9042f4e976b8ff657f1a5e2e7c2b5689ba03ff8a5f3f2ac6ad0 1657029602 +model-00026-of-00032.safetensors 86096e378a7cd9254ab95dd17557c3960d50d9ef230febf40cd7cb3fd8c76b68 1628192866 +model-00027-of-00032.safetensors 9dc27c40b43d42459cdf222102580e90f4a431116692d8b765bef56f9c2c70b9 1628192866 +model-00028-of-00032.safetensors aabccf617bfc00f86581e70646bac5aa7b9be564a34e263b8c7080623f7b4933 1628192866 +model-00029-of-00032.safetensors 5e3e4090b6c5fec39ca7a694899386910fc1a31f8c11e2de122099414b144cd0 1628192866 +model-00030-of-00032.safetensors a424e02c63531f8b4efc62e390e3c5db8680509c04772b4425c63e09db5cf184 1628192866 +model-00031-of-00032.safetensors ce11bf78b3f19cfd0814779f14aa7aab0dfbfa5eae7392823d9ffa6ada43ff7d 1997452570 +model-00032-of-00032.safetensors f4ed47cc36b78bc3b4720b96a362a6d2e1eae60a57d33d1937b27b7468970778 291222376 diff --git a/scripts/erp-tune-gx10/launch-run-06.sh b/scripts/erp-tune-gx10/launch-run-06.sh new file mode 100644 index 0000000..978f81e --- /dev/null +++ b/scripts/erp-tune-gx10/launch-run-06.sh @@ -0,0 +1,72 @@ +#!/usr/bin/env bash +# Launch ERP-seat SFT run 6 on pfi-gx10 (NVIDIA GB10, aarch64, sm_121). +# +# Run this ON pfi-gx10 as infra-ops. It detaches the job from the invoking +# shell and logs to the box, so a reaped SSH session cannot take the run with +# it -- the failure mode that lost the first probe launch on 2026-09-01. +# +# Run 6 = run 5 recipe UNCHANGED on the jenerallee78 ARA-abliterated base (the single variable). +# 8,212 survivors -> 524 optimizer steps, encode must match run 5 exactly. +# Checkpoints every 50 steps. +set -euo pipefail + +ROOT=/home/infra-ops/erp-tune +HARNESS=$ROOT/eitri-smithy +VENV=/home/infra-ops/ml/.venv/bin/python +CONFIG=$ROOT/run-06-gx10.json +LOG=$ROOT/run-06.log + +# --- Preconditions, asserted rather than assumed ----------------------------- + +# A stuck orphan holding unified memory while PyTorch reports zero allocated +# already doomed three relaunches on this box and got blamed on the new run +# each time. Assert the GPU is clear. +apps=$(nvidia-smi --query-compute-apps=pid --format=csv,noheader | tr -d '[:space:]') +if [ -n "$apps" ]; then + echo "REFUSING: GPU is not clear -- compute apps still resident:" >&2 + nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv >&2 + exit 1 +fi + +# Deliberately NOT `pgrep -f erp_sft_harness`: run this over ssh and the +# pattern appears in the invoking shell's own argv, so the guard matches +# itself and refuses every launch. The pidfile is exact and cannot self-match; +# the GPU assertion above catches an orphan under any name. +if [ -f "$ROOT/run-06.pid" ] && kill -0 "$(cat "$ROOT/run-06.pid")" 2>/dev/null; then + echo "REFUSING: run-06.pid names a live process $(cat "$ROOT/run-06.pid"):" >&2 + ps -p "$(cat "$ROOT/run-06.pid")" -o pid,etime,cmd >&2 + exit 1 +fi + +if [ -e "$LOG" ]; then + echo "REFUSING: $LOG exists. Move it aside first so two runs cannot share a log." >&2 + exit 1 +fi + +for p in "$HARNESS/erp_sft_harness/__main__.py" "$VENV" "$CONFIG"; do + [ -e "$p" ] || { echo "REFUSING: missing $p" >&2; exit 1; } +done + +# Free space for checkpoints, with headroom. +avail=$(df --output=avail -BG "$ROOT" | tail -1 | tr -dc '0-9') +if [ "$avail" -lt 40 ]; then + echo "REFUSING: only ${avail}G free under $ROOT; want >=40G for checkpoints." >&2 + exit 1 +fi + +# --- Launch ------------------------------------------------------------------ + +cd "$HARNESS" +{ + echo "# launched $(date -Is) on $(hostname) by ${USER}" + echo "# harness $(git rev-parse --short HEAD) config $CONFIG" +} > "$LOG" + +setsid nohup "$VENV" -m erp_sft_harness --config "$CONFIG" >> "$LOG" 2>&1 < /dev/null & +pid=$! +echo "$pid" > "$ROOT/run-06.pid" + +echo "launched pid $pid -> $LOG" +echo +echo "watch: tail -f $LOG | tr '\\r' '\\n'" +echo "stop: kill \$(cat $ROOT/run-06.pid) # by PID -- never pkill -f over ssh" diff --git a/scripts/erp-tune-gx10/pull-verify-jenerallee78.sh b/scripts/erp-tune-gx10/pull-verify-jenerallee78.sh new file mode 100644 index 0000000..c3232c1 --- /dev/null +++ b/scripts/erp-tune-gx10/pull-verify-jenerallee78.sh @@ -0,0 +1,54 @@ +#!/usr/bin/env bash +# Pull + verify jenerallee78/gemma-4-26B-A4B-it-ara-abliterated @ 0631379a onto pfi-gx10. +# Root shards + small files only; no GGUFs, no mlx-4bit. Verifies bytes against the +# brokkr-smithy pins (base-pin-jenerallee78.json) and installs the STOCK tokenizer set. +# Expects ~/erp-tune/base-pin-jenerallee78-shards.txt (file sha256 bytes per line). +set -uo pipefail +export PATH="$HOME/.local/bin:$PATH" +REPO=jenerallee78/gemma-4-26B-A4B-it-ara-abliterated +REV=0631379a3d859e0059bc8d9b21ab5b654dfc272c +DEST=$HOME/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a +STOCK=$HOME/models/gemma4-26b-a4b-it-bf16 +PINS=$HOME/erp-tune/base-pin-jenerallee78-shards.txt +echo "== start $(date -u +%FT%TZ) on $(hostname)" +mkdir -p "$DEST" +echo "== download" +uv run --quiet --with 'huggingface_hub[hf_transfer]' hf download "$REPO" --revision "$REV" \ + --local-dir "$DEST" \ + --include 'model-*-of-00032.safetensors' --include 'config.json' --include 'generation_config.json' \ + --include 'model.safetensors.index.json' --include 'chat_template.jinja' --include 'ara_config.json' --include 'README.md' \ + --include 'tokenizer.json' --include 'tokenizer_config.json' +rc=$? +echo "== download rc=$rc $(date -u +%FT%TZ)" +[ $rc -eq 0 ] || { echo "DOWNLOAD FAILED rc=$rc"; exit 2; } +cd "$DEST" +echo "== shard sha256 vs pins" +fail=0 +while read -r f oid bytes; do + [ -n "$f" ] || continue + sz=$(stat -c %s "$f" 2>/dev/null || echo MISSING) + got=$(sha256sum "$f" 2>/dev/null | cut -d' ' -f1) + if [ "$sz" = "$bytes" ] && [ "$got" = "$oid" ]; then echo "OK $f"; else echo "FAIL $f size=$sz want=$bytes sha=$got want=$oid"; fail=$((fail+1)); fi +done < "$PINS" +echo "== shard result: fail=$fail" +echo "== index + config checks" +python3 - "$STOCK" <<'PY' +import json,sys,hashlib +stock=sys.argv[1] +d=json.load(open('model.safetensors.index.json'));s=json.load(open(f'{stock}/model.safetensors.index.json')) +names=set(d['weight_map']);snames=set(s['weight_map']) +print('index weight_map:',len(names),'stock:',len(snames),'set_equal:',names==snames) +print('index total_size:',d['metadata'].get('total_size'),'stock:',s['metadata'].get('total_size'),'equal:',d['metadata'].get('total_size')==s['metadata'].get('total_size')) +c=json.load(open('config.json')) +print('config architectures:',c.get('architectures'),'dtype:',c.get('dtype') or c.get('torch_dtype')) +print('INDEX_SHA256', hashlib.sha256(open('model.safetensors.index.json','rb').read()).hexdigest()) +PY +echo "== repo tokenizer set as shipped (kept aside as *.repo)" +sha256sum tokenizer.json tokenizer_config.json chat_template.jinja +python3 -c "import json;print('repo tokenizer.json truncation:',json.load(open('tokenizer.json')).get('truncation'))" +for f in tokenizer.json tokenizer_config.json chat_template.jinja; do mv -n "$f" "$f.repo"; cp "$STOCK/$f" "$f"; done +echo "== STOCK tokenizer set installed (sha256):" +sha256sum tokenizer.json tokenizer_config.json chat_template.jinja +python3 -c "import json;print('installed tokenizer.json truncation:',json.load(open('tokenizer.json')).get('truncation'))" +echo "== listing"; ls -la "$DEST"; du -sh "$DEST" +echo "== done $(date -u +%FT%TZ) shard_fail=$fail" diff --git a/scripts/erp-tune-gx10/run-06-gx10.json b/scripts/erp-tune-gx10/run-06-gx10.json new file mode 100644 index 0000000..84cdfd3 --- /dev/null +++ b/scripts/erp-tune-gx10/run-06-gx10.json @@ -0,0 +1,47 @@ +{ + "output_dir": "/home/infra-ops/erp-tune/run-06", + "roots_dir": "/home/infra-ops/erp-tune/datasets/derived", + "base_model_path": "/home/infra-ops/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a", + "base_model_revision": "jenerallee78/gemma-4-26B-A4B-it-ara-abliterated @ 0631379a3d859e0059bc8d9b21ab5b654dfc272c (ARA abliteration of stock google/gemma-4-26B-A4B-it; 32 bf16 root shards sha256-verified against brokkr-smithy base-pin-jenerallee78.json; index sha256 33c59654e658a30fa29cdc87ccd6a752bfa0bb3e32cd56f95ff1eb82075e593a). THE SINGLE VARIABLE vs run 5: base only. Run-5 recipe, survivors, mask, template, hyperparameters all UNCHANGED. Tokenizer set = STOCK (tokenizer.json cc8d3a0c / tokenizer_config.json 9f4fec4b / chat_template.jinja ae53464b) copied over the repo's, whose shipped tokenizer.json carries a baked-in max_length=256 truncation; repo originals kept beside as *.repo. Runs 3/3c/4/5 were settled 2026-09-08 as STOCK base (index 907826a6), so this is the line's first abliterated base.", + "recipe": "/home/infra-ops/erp-tune/recipe-r6/recipe-erp-seat-sft-r6.json", + "survivors": "/home/infra-ops/erp-tune/recipe-r5/survivors-r5.jsonl", + "chat_template_path": "/home/infra-ops/models/gemma4-26b-a4b-it-bf16/chat_template.jinja", + "impersonation_mask_path": "/home/infra-ops/erp-tune/recipe-r3/lossmask-r3.jsonl", + "lora_rank": 64, + "lora_alpha": 128, + "lora_dropout": 0.0, + "max_seq_len": 16384, + "epochs": 1, + "seed": 20260824, + "per_device_batch_size": 2, + "gradient_accumulation_steps": 8, + "learning_rate": 0.0002, + "warmup_ratio": 0.1, + "lr_scheduler_type": "cosine", + "weight_decay": 0.01, + "load_in_4bit": false, + "gradient_checkpointing": true, + "loss_chunk_tokens": 1024, + "training_eligibility_override": "operator-2026-09-08-rnd-run6", + "overridden_blockers": [ + "contamination-scan-not-implemented", + "stage-2-csam-detector-inert" + ], + "substitute_controls": [ + "pre-training holdout, run-1 (8,404 samples, work/card/session split)", + "pre-training holdout, govreport/holdout-v1 (416 reports, sha256-ranked, never_trained_on)", + "pre-training holdout, qmsum/holdout-v1 (5 transcripts, sha256-ranked, never_trained_on)", + "stage-A lexical quarantine, RP (829 records held unread)", + "stage-A lexical quarantine, run-5 slot (133 records held unread, /mnt/smithy/datasets/quarantine/r47-run5-longdep-screen/)", + "SCROLLS-membership disclosure on both slot sources (avoidance, NOT a scan): govreport + qmsum are SCROLLS/ZeroSCROLLS members, in no hoard/default-benchmarks.yaml entry and used by no R47 instrument", + "kvasir is HELD, not re-cut: the 1,613 kvasir survivors are reused verbatim from survivors-r4.jsonl (which cut run-3's seed-20260824 prefix at 3,347,622 ctx). survivors-r5.jsonl = survivors-r4 minus airoboros plus the govreport + qmsum roots whole; sha256 a25169a6258cd4abb0cb494a176a921c0e98eb73d65c53d033b6ee18293a43ae.", + "window_count belt-and-suspenders (SFT-RECIPE-run5-SCOPE.md 7.1): every govreport + qmsum row renders <= 14,000 tokens (max 9,385 / 13,700) and the harness never packs across samples, so window_count MUST be 1 on every slot row; a chunked_into_2 or single_window_truncated on either new root in truncation-report.json is a BUILD DEFECT and the run is killed before training.", + "SINGLE VARIABLE vs run 5: the BASE. Stock google/gemma-4-26B-A4B-it OUT, jenerallee78 ARA abliteration @ 0631379a IN. Corpus (survivors-r5 rows verbatim), impersonation loss-mask (lossmask-r3), stock chat template ae53464b, lr 2e-04, max_seq_len 16384, rank 64, alpha 128, dropout 0.0, cosine, warmup 0.1, wd 0.01, batch 2 x accum 8, 1 epoch, seed 20260824 all UNCHANGED from run 5.", + "FREE CHECK (brokkr-smithy, 2026-09-08): the [encode] pass must reproduce run 5 EXACTLY -- 8,370 records, ctx 18,598,779 tok, loss 9,935,076 tok -- because corpus, tokenizer and template are identical; any difference means the wrong tokenizer/template loaded and the run is killed before [train].", + "HOST: pfi-gx10 (GB10, aarch64, sm_121, 121 GB unified). Base shards sha256-verified against the revision-pinned HF LFS oids after landing; harness eitri-smithy 0a6bd2e; corpus COPIED, box mounts no NFS. Grant: operator directive to infra-ops in-session 2026-09-08 (\"unload the gx10 and commence training on the gx10. window is open now.\").", + "SURVIVORS: recipe-r6 ships no survivor list of its own (targets byte-identical to r5), so survivors-r5.jsonl (sha256 a25169a6...) is reused verbatim; recipe-erp-seat-sft-r6.json sha256 6499555471181bd8ef7a273340f162e44fe54bf4f070198f050c8be33ee769a8 from brokkr-smithy 4dd7590." + ], + "unfittable": "drop", + "holdout_dir": "/home/infra-ops/erp-tune/datasets/holdout", + "save_steps": 50 +}