feat(coldfusion-abliteration): THESIS PROVEN — in-band-abliterated MTP head accepts 59.1% (beats incumbent ~47%)

Quantized the L35 abliterated model to mixed NVFP4 and measured MTP acceptance
end to end. The experiment's whole premise: Heretic (the incumbent gen seat)
leaves the MTP head a byte-identical base graft its wrapper never loads, whereas
Robinson abliterates the MTP head in-band — the question was whether that in-band
edit survives well enough to spec-decode. It does, better than the graft:

  MTP acceptance  59.1% median (51-65%, 8 cache-busted topics)  vs incumbent ~47%
  decode          118.7 tok/s median (faster; image-confounded, read as not-worse)
  abliteration    survives quant (creative refusals drop, self-harm guardrail
                  intact, coherent)

Output at /tank/aimodels/qwen38-27b-coldfusion-L35-nvfp4-mixed (22.5 GB). Result
JSON in bench/. NOT cut over — the incumbent seat is untouched; making L35 the gen
seat is a separate decision needing the full Stage-3 gate + real multi-turn hold.

Two env foot-guns hardened along the way:
- quant_mixed_nvfp4.py now promotes text_config attention fields
  (num_attention_heads etc.) to the top-level config for the oneshot, then
  restores. transformers 5.10 / llmcompressor 0.12 (this venv moved under us
  since the Aug-15 heresy quant) no longer delegate the top-level lookup, so
  oneshot raised "Cannot determine num_attention_heads". Same "the fight is the
  environment" pattern as the abliteration capture.
- a sub-~23GB quant saves as a single model.safetensors with no index, so the
  post_quant MTP graft needed an index built first — from the safetensors header,
  not safe_open (which mmaps the whole shard and ENOMEMs on ZFS).

post_quant grafted the abliterated MTP (15 tensors, 849 MB) and re-injected
re:^mtp.* into quantization_config.ignore (llm-compressor pruned it again — the
two-rounds-lost 0%-MTP bug, fired and repaired as designed). Probe served on the
pinned nightly (#51113 qwen3_5_mtp fix) to match the live seat's vLLM.
This commit is contained in:
vh
2026-08-20 10:18:48 -07:00
parent c55b1390b7
commit 725c8fdf9e
3 changed files with 100 additions and 0 deletions
@@ -126,6 +126,26 @@ def main():
model = Qwen3_5ForConditionalGeneration.from_pretrained(
a.model, torch_dtype="auto", device_map=None, trust_remote_code=True)
# llmcompressor 0.12 introspects attention structure off the TOP-LEVEL config
# (to size kv-cache/quant params). Qwen3_5 keeps num_attention_heads etc. under
# text_config, and transformers 5.10 no longer delegates the top-level lookup,
# so oneshot raises "Cannot determine num_attention_heads from config". Promote
# them from the authoritative text_config for the duration of quant, then
# restore, so the saved config keeps its canonical text_config-only shape.
# (This env moved under us since the 2026-08-15 heresy quant, where the older
# transformers still delegated — same "the fight is the environment" pattern.)
_promote = ("num_attention_heads", "num_key_value_heads", "hidden_size",
"head_dim", "num_hidden_layers")
_tc = getattr(model.config, "text_config", None)
_orig = {f: getattr(model.config, f, None) for f in _promote}
if _tc is not None:
for f in _promote:
v = getattr(_tc, f, None)
if v is not None:
setattr(model.config, f, v)
print("promoted text_config attention fields to top-level config for oneshot: "
+ ", ".join(f"{f}={getattr(model.config, f)}" for f in _promote), flush=True)
ds = load_calib(a.calib, tok, a.num_samples, a.seqlen)
recipe = build_recipe()
print("oneshot: NVFP4 W4A4 (L0-55 MLP) + FP8 W8A8 (attn/linear_attn/lm_head/L56-63 MLP) "
@@ -133,6 +153,14 @@ def main():
oneshot(model=model, dataset=ds, recipe=recipe,
num_calibration_samples=len(ds), max_seq_length=a.seqlen)
# restore the canonical config shape (undo the promotion above) so the saved
# top-level config matches the known-good heresy output; vLLM reads text_config.
for f, v in _orig.items():
try:
setattr(model.config, f, v)
except Exception:
pass
print(f"saving -> {a.out}", flush=True)
model.save_pretrained(a.out, save_compressed=True)
tok.save_pretrained(a.out)