feat(coldfusion-abliteration): dose-response settles the <think> leak — base 83%, our abliteration 17%
Answers "how likely is it that our abliteration caused this?" with a measurement instead of a prior. P(<think>) at the first generated token, template rendered enable_thinking=false so the prompt already carries a CLOSED think pair -- the exact event behind the leak. Raw softmax, bf16, CPU-only, one process per model. Deterministic: stock reproduced to 17 significant figures across two runs. coldfusion-bf16 none (stock) 0.1850 rank 3 coldfusion-abliterated-L35-bf16 Robinson L35, mild 0.2048 rank 2 coldfusion-h300-mtp-bf16 Heretic-300, heavy 0.2216 rank 2 The stock, untouched base already puts 18.5% of first-token mass on opening a think block the template had closed. Abliteration adds a real, monotonic, dose-dependent +3.7 points -- a nudge on a pre-existing base, not the cause. Cold-Fusion is a reasoning-token-compression finetune, i.e. a model trained to think briefly, and the leak's text shape agrees: a compact correct trace with a trained transition marker, which is trained behavior rather than damage. This changes the options. Rolling back to L35 or stock does NOT fix the leak -- at 18.5% under temp 0.7 / top_p 0.8 they leak at nearly the h300 rate. Only leaving the Cold-Fusion family escapes it, at the cost of the 8/100 refusal result. The chat_template_kwargs fix is the correct lever. Durable methodology point: a forward-KL budget cannot catch this. Heretic minimizes forward KL(stock||abliterated), which is near-blind to the model putting new mass on tokens stock barely used -- that is reverse KL's job, and we measured exactly that asymmetry on L35 (reverse 1.43 vs forward 0.70). h300's KL of 0.0136 is not evidence of innocence. For any "did the abliteration break behavior X" question, measure P(token) directly. Ran CPU-only deliberately: 96 EPYC cores and 265 GB of RAM make a 27B forward pass cheap, so this cost no GPU window and no seat downtime, where the obvious route was stopping both GPU0 seats. Also normalizes two more abliteration output dirs from root-owned 0600 to llmuser 0664. The unreadable-model failure surfaces as FileNotFoundError rather than a permission error, which is worth knowing before it wastes a run.
This commit is contained in:
+13
-1
@@ -126,7 +126,19 @@ _As of 2026-08-20 23:30 — **the Heretic-300 session** (see the 🔴 entry abov
|
||||
- **Mechanism.** With `enable_thinking:false` the chat template appends a **pre-closed** `<think>\n\n</think>\n\n` to the prompt (jinja L165-166). The h300 model **opens a fresh `<think>` anyway and never closes it** — verified raw: `has </think>: False`, `finish_reason: stop`, reasoning *and* answer in one `content` blob starting `<think>Ok, let's figure this out:`. vLLM's `qwen3` reasoning parser can't catch it: the prompt already closed the block, so the parser isn't in reasoning state and the tag is just text (`reasoning_content` empty, `reasoning_tokens: 0`). **The client is blameless** — Lobe correctly treats an unterminated `<think>` as still-thinking, so it renders an endless thought bubble and never shows the answer.
|
||||
- **★ It is a SAMPLING event, and the trigger is TEMPERATURE — not presence_penalty.** n=12 per arm on the reproducer: `pp 1.5` → 4 leaks, `pp 0.0` → 4, `pp 0.5` → 3 (all the same), **`temperature 0` → 0**. ⚠️ **This FALSIFIES the standing "presence_penalty 1.5 is the first dial to move" hypothesis** recorded in the litellm config comment and by the operator 2026-08-16 — it is not this bug's cause. Leave that dial alone for this symptom.
|
||||
- **Blast radius = exactly the two temp-0.7 aliases.** `gen` and `summarizer-large` leak (~17-27%); `summarizer`, `classifier`, `image-judge`, `qwen-image-bench` are all **temp=0 and clean at 0/12** — so **nevermore's `summarizer` path is NOT affected**. `gen-reasoning` doesn't leak (its think block is legitimately open) but shows the *other* symptom, empty `content`, at ~1/12.
|
||||
- **New with this build, probably.** The `<think>`/"figure this out" tell appears in `bench/eval_coldfusion_h300.json` (x2) and in **none** of `eval_aeon`, `eval_heresy`, `eval_mixed`, `eval_w4a16`. ⚠ **L35 was never evaled**, so this does NOT separate "a Cold-Fusion base trait" from "a Heretic-300 abliteration artifact" — both are Cold-Fusion. Abliteration optimizes refusal count under a KL cap and has **no notion of format compliance**, so degrading the pre-closed-think-block convention is a plausible cost of the 8/100 result.
|
||||
- **★★ CAUSE SETTLED BY MEASUREMENT 2026-08-21 — it is the BASE, not our abliteration (~83% / ~17%).** Measured P(`<think>`) at the first token with the template rendered `enable_thinking=false` (so the prompt already contains a *closed* think pair), raw softmax, bf16, **CPU-only — no GPU contention, no seat downtime** (96 EPYC cores, 265 GB RAM; ~10s to load a 52 GB model from ARC). Deterministic — stock reproduced to 17 s.f. across two runs.
|
||||
|
||||
| build | abliteration | **P(`<think>`)** | rank |
|
||||
|---|---|---|---|
|
||||
| `coldfusion-bf16` | **none (stock)** | **0.1850** | 3 |
|
||||
| `coldfusion-abliterated-L35-bf16` | Robinson L35, mild | 0.2048 | 2 |
|
||||
| `coldfusion-h300-mtp-bf16` | Heretic-300, heavy | 0.2216 | 2 |
|
||||
|
||||
**The stock, untouched base already puts 18.5% of first-token mass on opening a think block the template had closed.** Abliteration adds a real, monotonic, dose-dependent **+3.7 points** (+2.0 L35, +1.7 more h300) — a nudge on a pre-existing 18.5-point base. Root cause is the Cold-Fusion finetune itself, which our own PROVENANCE describes as **"reasoning-token compression (thinking cut to 1/10..1/2 of stock)"** — a model trained to think, briefly. The leak's text shape confirms it: a compact, *correct* trace with a trained transition marker (`Ok, I'm ready to generate my response:`), which is trained behavior, not damage.
|
||||
- **⚠️ ACTION CONSEQUENCE — ROLLING BACK DOES NOT FIX IT.** At 18.5% mass under temp 0.7 / top_p 0.8, **stock Cold-Fusion and L35 would leak at nearly the h300 rate**. The only rollback that escapes is leaving the Cold-Fusion family entirely (→ `qwen38-27b-heresy-nvfp4-mixed`), which costs the 8/100 refusal result. **The `chat_template_kwargs` fix is the correct lever, not a model swap.**
|
||||
- **★ A FORWARD-KL BUDGET CANNOT CATCH THIS — do not quote KL as evidence of innocence.** Heretic minimizes **forward** KL(stock‖abliterated), which is near-blind to the model putting NEW mass on tokens stock barely used — that is what **reverse** KL penalizes, and we measured exactly that asymmetry on L35 (reverse 1.43 vs forward 0.70, "the mass-where-stock-had-none signature"). So h300's KL 0.0136 proves nothing here; only the direct measurement did. **Generalizes: for any "did the abliteration break behavior X?" question, measure P(token) directly — the KL number is the wrong instrument.**
|
||||
- Consistent with the NVFP4 quant contributing ~nothing: the live quantized seat reads 0.2267 vs its bf16 parent's 0.2216. ⚠ different stacks and a possible vLLM temperature scaling make that suggestive, **not proven**.
|
||||
- Full method + reproducers: `services/gen-seat-mixed-quant/bench/think-leak/` (`think-prior-dose-response.json`, `README.md`).
|
||||
- **FIX, validated n=30 over 4 prompt types + a 3-turn conversation:** `chat_template_kwargs: {enable_thinking: true, reasoning_effort: low}` on `gen` → **0/30 leaks** (current config: **8/30**, worst on prose 4/6). Give the model a legitimately open `<think>` and it closes it properly, the parser does its job, `content` comes out clean. Cost ~+27% completion tokens (257 vs 202 avg) and a residual **1/30 empty-content**. Semantic change: `gen` stops being a non-thinking deployment — **operator's call, not applied.**
|
||||
- **★ PROCESS LESSON: the 7/7 alias smoke test structurally CANNOT catch this.** Trivial prompts ("Reply with exactly: OK-gen") never invite reasoning, so they never sample the leaking token. Same shape as the compose file's own warning that single-turn probes missed the xhigh budget bug. **Probe with a reasoning-inviting prompt at n≥12, and grep the raw `content` for `<think>` — never just check HTTP 200.** The tell was sitting in my own `eval_coldfusion_h300.json` output the night of the cutover and I read past it.
|
||||
|
||||
|
||||
@@ -67,3 +67,41 @@ tokens and leaves a ~3% empty-content residual.
|
||||
**A "7/7 aliases return 200" smoke test cannot catch this.** Trivial prompts
|
||||
never invite reasoning, so they never sample the leaking token. Probe with a
|
||||
reasoning-inviting prompt at n>=12 and grep raw `content` for `<think>`.
|
||||
|
||||
## Did our abliteration cause it? No — the base did (~83% / ~17%)
|
||||
|
||||
`think_prior.py <model-path>` measures P(`<think>`) at the **first generated
|
||||
token**, with the template rendered `enable_thinking=false` so the prompt already
|
||||
contains a *closed* think pair. That is the exact event behind the leak. Raw
|
||||
softmax over the full vocab, bf16, **CPU-only** — no GPU contention and no seat
|
||||
downtime, which is what makes this cheap enough to run on any future candidate.
|
||||
|
||||
| build | abliteration | P(`<think>`) | rank |
|
||||
|---|---|---|---|
|
||||
| `coldfusion-bf16` | **none (stock)** | **0.1850** | 3 |
|
||||
| `coldfusion-abliterated-L35-bf16` | Robinson L35, mild | 0.2048 | 2 |
|
||||
| `coldfusion-h300-mtp-bf16` | Heretic-300, heavy | 0.2216 | 2 |
|
||||
|
||||
The stock base already carries 18.5%. Abliteration adds a real, monotonic,
|
||||
dose-dependent **+3.7 points** — a nudge on a pre-existing base, not the cause.
|
||||
Cold-Fusion is a *reasoning-token-compression* finetune; it is a model trained to
|
||||
think, briefly, and it wants to open that block.
|
||||
|
||||
**Consequence: rolling back does not fix the leak.** At 18.5% mass under temp 0.7
|
||||
/ top_p 0.8 the stock base leaks at nearly the same rate. Only leaving the
|
||||
Cold-Fusion family escapes it. The `chat_template_kwargs` fix is the right lever.
|
||||
|
||||
**A forward-KL budget cannot catch this.** Heretic minimizes forward
|
||||
KL(stock||abliterated), which is near-blind to new mass on tokens stock barely
|
||||
used — that is reverse KL's job, and we measured that asymmetry on L35 (reverse
|
||||
1.43 vs forward 0.70). h300's KL of 0.0136 is not evidence of innocence. For any
|
||||
"did the abliteration break behavior X" question, measure P(token) directly.
|
||||
|
||||
Run it:
|
||||
|
||||
sudo -u llmuser env OMP_NUM_THREADS=48 CUDA_VISIBLE_DEVICES="" \
|
||||
/tank/aimodels/quant-work/.venv/bin/python think_prior.py /tank/aimodels/<build>
|
||||
|
||||
⚠ Abliteration and post-quant outputs are written root-owned `0600` and are
|
||||
unreadable to `llmuser`; normalize to `llmuser:llmuser 0664` first. The failure
|
||||
surfaces as a misleading `FileNotFoundError`, not a permission error.
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"measured": "2026-08-21",
|
||||
"question": "Did our abliteration cause the unterminated-<think> leak on the gen seat?",
|
||||
"method": "P(<think>) at the FIRST generated token, with the chat template rendered at enable_thinking=false so the prompt already contains a CLOSED <think></think> pair. Raw softmax over the full vocab, bf16, CPU-only (no GPU contention, no seat downtime), one process per model. Deterministic: stock reproduced to 17 significant figures across two independent runs.",
|
||||
"prompt": "A farmer has 17 sheep. All but 9 run away. He buys twice as many as he has left, then sells 4. How many now? Explain.",
|
||||
"think_token_id": 248068,
|
||||
"arms": [
|
||||
{"build": "qwen38-27b-coldfusion-bf16", "abliteration": "none (stock DavidAU Cold-Fusion-GAIN-V1.1)", "p_think": 0.18504, "rank": 3},
|
||||
{"build": "qwen38-27b-coldfusion-abliterated-L35-bf16", "abliteration": "Robinson L35, mild", "p_think": 0.20479, "rank": 2},
|
||||
{"build": "qwen38-27b-coldfusion-h300-mtp-bf16", "abliteration": "Heretic-300, heavy (8/100 refusals)", "p_think": 0.22162, "rank": 2}
|
||||
],
|
||||
"verdict": {
|
||||
"base_share": 0.835,
|
||||
"abliteration_share": 0.165,
|
||||
"delta_total": 0.03658,
|
||||
"delta_L35": 0.01975,
|
||||
"delta_h300_over_L35": 0.01683,
|
||||
"reading": "The STOCK base already puts 18.5% of first-token mass on opening a think block the template had closed. Abliteration adds a real, monotonic, dose-dependent +3.7 points on top -- but it is a nudge on a pre-existing 18.5-point base, not the cause. ~83% of the effect is the Cold-Fusion finetune, which our own PROVENANCE describes as reasoning-token compression: a model finetuned to think, briefly.",
|
||||
"action_consequence": "Rolling back to L35 or to stock Cold-Fusion does NOT fix the leak -- at 18.5% mass and temp 0.7 / top_p 0.8 the stock base would leak at nearly the same rate. Only leaving the Cold-Fusion family (back to qwen38-27b-heresy-nvfp4-mixed) would, and that costs the 8/100 refusal result. The chat_template_kwargs fix is the right lever."
|
||||
},
|
||||
"notes": [
|
||||
"The live NVFP4 h300 seat measured p_think 0.2267 through vLLM logprobs vs 0.2216 for its bf16 parent here. Different stacks and a possible temperature scaling on the vLLM side make that not a clean comparison, but it is CONSISTENT WITH the quant contributing ~nothing. Not proven.",
|
||||
"qwen38-27b-coldfusion-heretic300-bf16 (the raw MTP-less Heretic export) could not be measured: its index references model-mtp.safetensors, which that export does not contain. Expected -- it is the known MTP-dropping export -- and irrelevant to this question.",
|
||||
"Both abliterated bf16 dirs were written root-owned 0600 and were unreadable to llmuser; normalized to llmuser:llmuser 0664 to match every other model dir. Same footgun as the h300 quant dir.",
|
||||
"A forward-KL budget CANNOT catch this. Heretic minimizes forward KL(stock||abliterated), which is near-blind to the model putting NEW mass on tokens stock barely used -- that is what reverse KL penalizes, and we measured exactly that asymmetry on L35 (reverse 1.43 vs forward 0.70). The h300 KL of 0.0136 is therefore not evidence of innocence; the measurement above is."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
#!/usr/bin/env python3
|
||||
"""P(<think>) at the first generated token, with enable_thinking=False.
|
||||
|
||||
Measures how much probability mass a checkpoint puts on OPENING a think block
|
||||
when the chat template has ALREADY closed one for it. That is the exact event
|
||||
behind the h300 gen-seat leak. CPU-only: no GPU contention, no seat downtime.
|
||||
"""
|
||||
import json, sys, torch
|
||||
from transformers import AutoTokenizer, AutoModelForCausalLM
|
||||
|
||||
path = sys.argv[1]
|
||||
PROMPT = ("A farmer has 17 sheep. All but 9 run away. He buys twice as many as he "
|
||||
"has left, then sells 4. How many now? Explain.")
|
||||
|
||||
tok = AutoTokenizer.from_pretrained(path)
|
||||
text = tok.apply_chat_template([{"role": "user", "content": PROMPT}],
|
||||
tokenize=False, add_generation_prompt=True,
|
||||
enable_thinking=False)
|
||||
assert text.rstrip().endswith("</think>"), "template did NOT pre-close the think block:\n" + repr(text[-120:])
|
||||
|
||||
ids = tok(text, return_tensors="pt")
|
||||
model = AutoModelForCausalLM.from_pretrained(path, dtype=torch.bfloat16, device_map=None)
|
||||
model.eval()
|
||||
with torch.no_grad():
|
||||
logits = model(**ids).logits[0, -1].float()
|
||||
probs = torch.softmax(logits, dim=-1)
|
||||
|
||||
think_id = tok.convert_tokens_to_ids("<think>")
|
||||
p_think = probs[think_id].item()
|
||||
top = torch.topk(probs, 12)
|
||||
out = {
|
||||
"model": path.rstrip("/").split("/")[-1],
|
||||
"p_think": p_think,
|
||||
"think_token_id": think_id,
|
||||
"think_rank": int((probs > p_think).sum().item()) + 1,
|
||||
"top12": [{"tok": tok.decode([i]), "p": round(probs[i].item(), 5)} for i in top.indices.tolist()],
|
||||
}
|
||||
print("RESULT " + json.dumps(out))
|
||||
Reference in New Issue
Block a user