docs(gemma4-charrp): abliteration measured in isolation — close to free, but it MOVES capability
Second bench window, operator-authorised after an initial decline and reversal. Stock BF16 against the llmfan46 abliterated BF16: same precision, same pinned upstream template, same 192 items, CoT off. Abliteration was the only axis that moved, which is what the previous run could not claim. Net core cost is 0.6 points — but the headline understates what happened. Capability MOVED rather than degraded: five items lost on contradiction detection, four gained on spatial composition, nearly cancelling. A gain was not predicted by anyone, least of all on that axis. The decision this was authorised to settle: llmfan46 stands as the trainee base. No case for re-staging on TrevorJS at KL 0.09 over 0.6 points — the KL gap between the builds is smaller than the gap this measurement failed to find. Both limits recorded rather than buried, per brokkr-smithy-dev: the swings are ~5 and ~4 items at n=32, so the -15.6/+12.5 percentages read more precisely than the measurement supports and only marginals were run; and this says nothing about quantization, because the stock-NVFP4 T2 figure came from n=16 against n=32 here — different item counts mean different item sets, so that comparison is n-confounded and is not being made. Turnaround was five minutes rather than fifteen because the gemma4-trainee-bench stack already existed — itself the residue of debugging a 35-restart crash-loop caused by the production compose hardcoding --quantization compressed-tensors. The fix outlasted the incident. gen restored and verified through the gateway; char-rp remains down deliberately; bench stack env reset to the heretic base for the post-tune gate.
This commit is contained in:
@@ -52,6 +52,37 @@ propagated through the ecosystem, not that one packager slipped.
|
||||
For both, point at the upstream file:
|
||||
`/tank/aimodels/gemma4-26b-a4b-it-bf16/chat_template.jinja`.
|
||||
|
||||
### Measured: abliteration is close to free on this base (2026-08-24)
|
||||
|
||||
Isolated properly — stock BF16 against llmfan46 BF16, same precision, same
|
||||
pinned upstream template, same 192 items, CoT off. **Abliteration was the only
|
||||
axis that moved.**
|
||||
|
||||
| task | stock BF16 | heretic BF16 | items |
|
||||
|---|---|---|---|
|
||||
| T1 state / T3 constraint / T4 long-context / T5 control | 100% | 100% | 0 |
|
||||
| T2 contradiction | 75% | 59% | **−5** |
|
||||
| T6 spatial | 75% | 88% | **+4** |
|
||||
| **core** | **90.0%** | **89.4%** | −0.6 pts |
|
||||
|
||||
**Net cost 0.6 points — but it MOVED capability rather than removing it.** Five
|
||||
items lost on contradiction detection, four gained on spatial composition,
|
||||
nearly cancelling. Nobody predicted a gain anywhere, least of all that
|
||||
direction.
|
||||
|
||||
Consequence for the base choice: **llmfan46 stands.** There is no case for
|
||||
re-staging on TrevorJS at KL 0.09 over a 0.6-point net difference — the KL gap
|
||||
between the two builds is smaller than the gap this measurement failed to find.
|
||||
|
||||
⚠ Read those as **~5 items and ~4 items at n=32**, not as −15.6/+12.5 percent.
|
||||
The percentages read more precisely than the measurement supports, and only
|
||||
marginals were run — no paired per-item analysis.
|
||||
|
||||
⚠ **This says nothing about quantization.** Stock BF16 scores T2 75% where stock
|
||||
NVFP4 scored 94%, but those runs were n=32 and n=16 — different item counts mean
|
||||
different item sets, and the extra items are not guaranteed equally easy. That
|
||||
comparison is n-confounded and is not being made.
|
||||
|
||||
### Choosing an abliterated base — compare on published damage, not on names
|
||||
|
||||
"Low damage" has a measurable proxy and the field spreads widely on it:
|
||||
|
||||
Reference in New Issue
Block a user