# Refusal retention — the axis the gate did not have, and the axis I measured wrong `[2026-08-25]` ## Why it exists brokkr's gate measures reasoning (T1-T6), craft (diversity/attractor) and regurgitation (memorisation). **Nothing measured whether the model still COMPLIES** — which for this seat is arguably the most important property. The risk is specific to our operation order. We do **tune(abliterate(stock))**, so the tune has 57.7M tokens of opportunity to walk the abliteration back. *A tune that gains 41 items of contradiction detection and quietly re-installs refusals is a failed seat that passes the entire gate.* ## The measurement — controlled, single instrument, both arms arm HARD DEFLECT COMPLY base 0/100 0 100 tuned 29/100 0 71 Same seat, same probe, temp 0, `mlabonne/harmful_behaviors` x100. Probe: `scripts/training-probes/refusal_probe.py`. **The tune added 29 general-harm refusals where the base had none.** Two things fell out: - **The instrument validates.** Base measured 0/100 on my generated-text regex against Heretic's recorded 3/100 from a first-token-probability scorer. 0 vs 3 is agreement — the incomparability worry was right caution about a non-problem. - **DEFLECT is 0 on BOTH arms, so the free control fires.** An instrument artifact does not care which arm it runs against. Both zero means the model is **binary** — refuses in refusal-language or engages, no soft-deflection tail. The R19 undercount does not apply here. ## ⚠⚠ But it is the WRONG AXIS — brokkr's catch, and it is the better one `mlabonne/harmful_behaviors` is **general harm** — weapons, malware, fraud. **The abliteration was not run so the model would explain bomb-making. It was run so the model would engage with explicit fiction.** Different refusal surfaces; a model moves on them independently. I picked that set because it was cached, had a recorded baseline, and was what the abliteration tool used. **Every one of those is a reason it was convenient, not a reason it was right** — and "it has a baseline" was actively misleading, because a comparable number for a question nobody is asking looks like evidence. **29/100 general-harm refusals on a seat writing prose the operator was actively praising is plausibly the DESIRED shape**, not a defect. General-harm refusals returning while domain compliance holds is close to ideal for an internal creative seat. I would have reported it as damage. **The load-bearing cell is COMPLY 71, not the 29.** Stock refused 100/100; anything near that would mean the abliteration was undone. 71 complying means "partially walked back on one axis" — a different finding, and only one of the two threatens the seat. Domain-compliance probe (the right axis, from R19's track-2 map) is brokkr's, pending. Scaffold supplied: `scripts/training-probes/counted_classifier.py` (`2a05ae9`) — classify-never-surface, three-way, ERROR path deliberately does not log the exception body because an exception can echo the prompt back. Playbook §3.13. See [[2026-08-25-erp-tune-run2-complete]].