revert(gen-seat): abandon Cold-Fusion, roll back to heresy — the leak is in the base

Operator directive, given before the result was in: if it's the base, abandon
h300 and the base too. The dose-response said base (18.5% of 22.2%), so it fired.

Live gen seat is /tank/aimodels/qwen38-27b-heresy-nvfp4-mixed again, restored
from .env.bak-coldfusion-L35-20260820. The h300 env is preserved at
.env.bak-h300-abandoned-20260821.

The clincher, same probe pointed at heresy:

  Cold-Fusion stock          P(<think>) 0.1850
  Cold-Fusion L35                       0.2048
  Cold-Fusion h300                      0.2216
  heresy (restored)          not in the top 20, <0.002

A >100x gap between the families, which is why no rollback inside Cold-Fusion
would have helped -- stock and L35 leak at nearly the h300 rate.

Verified after rollback: 0/30 leaks and 0 empty on the same instrument that
scored h300 at 8/30, with the EXISTING enable_thinking:false config; KV pool
403,065 tok / 1.54x, heresy's exact documented baseline; 7/7 aliases; vision
intact. No LiteLLM change was needed, so the chat_template_kwargs fix is left
unapplied -- it worked, but it was a workaround for a base we no longer serve.

Cost, stated plainly: 8/100 refusals becomes 29/100, a 3.6x regression on the
axis the whole Heretic-300 run existed to move. Accepted deliberately.

What carries forward is the methodology, none of which lived in the Cold-Fusion
weights: direction_scope=0 beating per-layer on a merged base, aggression not
being the lever, PR #317 silently dropping the MTP head on save, the MPOA and
sink-screen reasoning, the graft/KL/catatonia/export harnesses, and the finding
that a pristine MTP graft accepts as well as an in-band edit.

New acceptance gate earned here: run think_prior.py on a candidate's STOCK
weights before committing GPU time. It is a ~10s CPU measurement and it would
have disqualified Cold-Fusion before the 300-trial study ever started. Heretic's
objective has no format-compliance term at all -- the same blindness that removed
the self-harm guardrail.

Nothing deleted. Every Cold-Fusion artifact, the 300-trial Optuna journal and
catatonia-T260.json remain on disk. Abandon means stop serving, not rm.
This commit is contained in:
vh
2026-08-21 00:40:23 -07:00
parent 5ee2325820
commit 37e9e1ca7f
2 changed files with 44 additions and 3 deletions
@@ -105,3 +105,35 @@ Run it:
⚠ Abliteration and post-quant outputs are written root-owned `0600` and are
unreadable to `llmuser`; normalize to `llmuser:llmuser 0664` first. The failure
surfaces as a misleading `FileNotFoundError`, not a permission error.
## Resolution — Cold-Fusion abandoned, seat rolled back to `heresy` (2026-08-21)
The operator's call, made before the result was in: *"If it's the base, we abandon
h300 AND the base."* The measurement said base, so it fired. The gen seat is back
on `qwen38-27b-heresy-nvfp4-mixed`.
The clincher is the same instrument, pointed at heresy:
| build | P(`<think>`) at first token |
|---|---|
| Cold-Fusion stock | 0.1850 |
| Cold-Fusion L35 | 0.2048 |
| Cold-Fusion h300 | 0.2216 |
| **`heresy` (restored)** | **not in the top 20 — <0.002** |
A >100x gap. The two families are categorically different here, which is exactly
why no rollback *inside* Cold-Fusion would have helped.
Verified after the rollback, same probes as before:
- `final_validate.py` — **0/30 leaks, 0 empty**, with the *existing*
`enable_thinking: false` config. h300 scored 8/30 on this same instrument.
- KV pool 403,065 tok / 1.54x — heresy's exact documented baseline.
- 7/7 gateway aliases 200; vision intact.
- **No LiteLLM config change was needed.** The `chat_template_kwargs` fix
developed above is left unapplied: it worked, but it was a workaround for a base
we no longer serve.
**The gate this earns: run `think_prior.py` on a candidate's STOCK weights before
committing GPU time to it.** It is a ~10s CPU measurement, and it would have
disqualified Cold-Fusion before the 300-trial Heretic study ever started.