diff --git a/stacks/gemma4-charrp/README.md b/stacks/gemma4-charrp/README.md index 6380a8d..f584ced 100644 --- a/stacks/gemma4-charrp/README.md +++ b/stacks/gemma4-charrp/README.md @@ -52,6 +52,37 @@ propagated through the ecosystem, not that one packager slipped. For both, point at the upstream file: `/tank/aimodels/gemma4-26b-a4b-it-bf16/chat_template.jinja`. +### Measured: abliteration is close to free on this base (2026-08-24) + +Isolated properly — stock BF16 against llmfan46 BF16, same precision, same +pinned upstream template, same 192 items, CoT off. **Abliteration was the only +axis that moved.** + +| task | stock BF16 | heretic BF16 | items | +|---|---|---|---| +| T1 state / T3 constraint / T4 long-context / T5 control | 100% | 100% | 0 | +| T2 contradiction | 75% | 59% | **−5** | +| T6 spatial | 75% | 88% | **+4** | +| **core** | **90.0%** | **89.4%** | −0.6 pts | + +**Net cost 0.6 points — but it MOVED capability rather than removing it.** Five +items lost on contradiction detection, four gained on spatial composition, +nearly cancelling. Nobody predicted a gain anywhere, least of all that +direction. + +Consequence for the base choice: **llmfan46 stands.** There is no case for +re-staging on TrevorJS at KL 0.09 over a 0.6-point net difference — the KL gap +between the two builds is smaller than the gap this measurement failed to find. + +⚠ Read those as **~5 items and ~4 items at n=32**, not as −15.6/+12.5 percent. +The percentages read more precisely than the measurement supports, and only +marginals were run — no paired per-item analysis. + +⚠ **This says nothing about quantization.** Stock BF16 scores T2 75% where stock +NVFP4 scored 94%, but those runs were n=32 and n=16 — different item counts mean +different item sets, and the extra items are not guaranteed equally easy. That +comparison is n-confounded and is not being made. + ### Choosing an abliterated base — compare on published damage, not on names "Low damage" has a measurable proxy and the field spreads widely on it: