docs(gemma4-charrp): abliteration measured in isolation — close to free, but it MOVES capability

Second bench window, operator-authorised after an initial decline and reversal.
Stock BF16 against the llmfan46 abliterated BF16: same precision, same pinned
upstream template, same 192 items, CoT off. Abliteration was the only axis that
moved, which is what the previous run could not claim.

Net core cost is 0.6 points — but the headline understates what happened.
Capability MOVED rather than degraded: five items lost on contradiction
detection, four gained on spatial composition, nearly cancelling. A gain was not
predicted by anyone, least of all on that axis.

The decision this was authorised to settle: llmfan46 stands as the trainee base.
No case for re-staging on TrevorJS at KL 0.09 over 0.6 points — the KL gap
between the builds is smaller than the gap this measurement failed to find.

Both limits recorded rather than buried, per brokkr-smithy-dev: the swings are
~5 and ~4 items at n=32, so the -15.6/+12.5 percentages read more precisely than
the measurement supports and only marginals were run; and this says nothing
about quantization, because the stock-NVFP4 T2 figure came from n=16 against
n=32 here — different item counts mean different item sets, so that comparison
is n-confounded and is not being made.

Turnaround was five minutes rather than fifteen because the gemma4-trainee-bench
stack already existed — itself the residue of debugging a 35-restart crash-loop
caused by the production compose hardcoding --quantization compressed-tensors.
The fix outlasted the incident.

gen restored and verified through the gateway; char-rp remains down deliberately;
bench stack env reset to the heretic base for the post-tune gate.
This commit is contained in:
2026-08-24 15:39:43 -07:00
parent 019ccff7e8
commit 5415fd4b30
+31
View File
@@ -52,6 +52,37 @@ propagated through the ecosystem, not that one packager slipped.
For both, point at the upstream file:
`/tank/aimodels/gemma4-26b-a4b-it-bf16/chat_template.jinja`.
### Measured: abliteration is close to free on this base (2026-08-24)
Isolated properly — stock BF16 against llmfan46 BF16, same precision, same
pinned upstream template, same 192 items, CoT off. **Abliteration was the only
axis that moved.**
| task | stock BF16 | heretic BF16 | items |
|---|---|---|---|
| T1 state / T3 constraint / T4 long-context / T5 control | 100% | 100% | 0 |
| T2 contradiction | 75% | 59% | **5** |
| T6 spatial | 75% | 88% | **+4** |
| **core** | **90.0%** | **89.4%** | 0.6 pts |
**Net cost 0.6 points — but it MOVED capability rather than removing it.** Five
items lost on contradiction detection, four gained on spatial composition,
nearly cancelling. Nobody predicted a gain anywhere, least of all that
direction.
Consequence for the base choice: **llmfan46 stands.** There is no case for
re-staging on TrevorJS at KL 0.09 over a 0.6-point net difference — the KL gap
between the two builds is smaller than the gap this measurement failed to find.
⚠ Read those as **~5 items and ~4 items at n=32**, not as 15.6/+12.5 percent.
The percentages read more precisely than the measurement supports, and only
marginals were run — no paired per-item analysis.
**This says nothing about quantization.** Stock BF16 scores T2 75% where stock
NVFP4 scored 94%, but those runs were n=32 and n=16 — different item counts mean
different item sets, and the extra items are not guaranteed equally easy. That
comparison is n-confounded and is not being made.
### Choosing an abliterated base — compare on published damage, not on names
"Low damage" has a measurable proxy and the field spreads widely on it: