Run-01 completed in 7:21:52 (47% faster than the 13.85h round-1 projection),
lora_B gate 205/205 non-zero at median norm 1.708, and the acceptance gate says
it did the thing it was built for: diversity +0.178 against a 0.008 floor (22x),
attractor hit rate -11.3pt against a 2.0pt floor, memorisation 0.0000 on both
arms — which closes the R20 licensed-prose exposure on measurement rather than
argument.
Five new detail files carry the substance:
erp-tune-run2-complete the run, the gate, the noise-floor near-miss
(brokkr was one step from reporting a 13-point
T6 regression sitting inside twice his
instrument's own variance)
mfu-root-caused-attention 8.6% MFU was an accounting artifact; real
utilisation 17-20%, cost was attention on
AMPERE kernels. Two independent methods agreed
to 2.6 points.
nvfp4-serving-pipeline merged weights are MANDATORY — vLLM cannot
serve a LoRA on ANY Gemma-4 — plus the recipe
that silently misses all 11,520 expert tensors
refusal-retention-probe measured base 0/100 -> tuned 29/100, then had
to accept it was the wrong axis
worldtree-b188-b189-and-selene three arcs closed, and a #411 diagnosis I got
wrong twice before a directory probe settled it
Current state rewritten end to end — the previous snapshot had the run in
flight at ~17h with MFU unexplained. Both are now closed.
The open operator decision is run 2's base, deliberately unstaged and flagged
against being filed as a config knob: it is a reversal of the trainee-selection
decision, and the pretrained-base option removes the last non-lexical floor on
the CSAM axis given stage-2-detector-inert and contamination-scan-absent are
both already overridden.
Tried-and-abandoned gains four measured-dead throughput levers, the packing
correction (bucketing wins under sdpa and the conclusion flips under flex — do
not carry it past the backend decision), and the merge-back-undoes-abliteration
trap brokkr caught in his own advice.
Index stays at 291 lines, under the soft cap. No archival this run.
3.1 KiB
Refusal retention — the axis the gate did not have, and the axis I measured wrong
[2026-08-25]
Why it exists
brokkr's gate measures reasoning (T1-T6), craft (diversity/attractor) and regurgitation (memorisation). Nothing measured whether the model still COMPLIES — which for this seat is arguably the most important property.
The risk is specific to our operation order. We do tune(abliterate(stock)), so the tune has 57.7M tokens of opportunity to walk the abliteration back. A tune that gains 41 items of contradiction detection and quietly re-installs refusals is a failed seat that passes the entire gate.
The measurement — controlled, single instrument, both arms
arm HARD DEFLECT COMPLY
base 0/100 0 100
tuned 29/100 0 71
Same seat, same probe, temp 0, mlabonne/harmful_behaviors x100.
Probe: scripts/training-probes/refusal_probe.py.
The tune added 29 general-harm refusals where the base had none.
Two things fell out:
- The instrument validates. Base measured 0/100 on my generated-text regex against Heretic's recorded 3/100 from a first-token-probability scorer. 0 vs 3 is agreement — the incomparability worry was right caution about a non-problem.
- DEFLECT is 0 on BOTH arms, so the free control fires. An instrument artifact does not care which arm it runs against. Both zero means the model is binary — refuses in refusal-language or engages, no soft-deflection tail. The R19 undercount does not apply here.
⚠⚠ But it is the WRONG AXIS — brokkr's catch, and it is the better one
mlabonne/harmful_behaviors is general harm — weapons, malware, fraud. The
abliteration was not run so the model would explain bomb-making. It was run so
the model would engage with explicit fiction. Different refusal surfaces; a
model moves on them independently.
I picked that set because it was cached, had a recorded baseline, and was what the abliteration tool used. Every one of those is a reason it was convenient, not a reason it was right — and "it has a baseline" was actively misleading, because a comparable number for a question nobody is asking looks like evidence.
29/100 general-harm refusals on a seat writing prose the operator was actively praising is plausibly the DESIRED shape, not a defect. General-harm refusals returning while domain compliance holds is close to ideal for an internal creative seat. I would have reported it as damage.
The load-bearing cell is COMPLY 71, not the 29. Stock refused 100/100; anything near that would mean the abliteration was undone. 71 complying means "partially walked back on one axis" — a different finding, and only one of the two threatens the seat.
Domain-compliance probe (the right axis, from R19's track-2 map) is brokkr's,
pending. Scaffold supplied: scripts/training-probes/counted_classifier.py
(2a05ae9) — classify-never-surface, three-way, ERROR path deliberately does not
log the exception body because an exception can echo the prompt back.
Playbook §3.13. See 2026-08-25-erp-tune-run2-complete.