Files
esh-pfi-infrastructure/services/gen-seat-mixed-quant/bench/think-leak/blast.py
T
vh 91f4cf22e1 fix(gen-seat): diagnose the unterminated-<think> leak — model defect, temp-triggered
Operator reported the new Heretic-300 gen seat "sends CoT but never completes
the turn" through Lobe. Diagnosed; not yet fixed (the fix changes gen's
semantics, so it is the operator's call).

The Qwen3.8 chat template appends a pre-closed <think>\n\n</think>\n\n when
enable_thinking is false. The h300 model opens a fresh <think> anyway and never
closes it. Because the prompt already closed the block, vLLM's qwen3 reasoning
parser is not in reasoning state, so the tag passes through as ordinary text --
reasoning_content empty, reasoning_tokens 0, and the whole reasoning-plus-answer
blob lands in content. Lobe then correctly treats the unterminated tag as
still-thinking and renders no answer. The client and the serving stack are both
behaving correctly; the model is not.

The trigger is TEMPERATURE, not presence_penalty (n=12 per arm):

  temp 0.7, pp 1.5  (current gen)   4/12
  temp 0.7, pp 0.0                  4/12
  temp 0.7, pp 0.5                  3/12
  temp 0,   pp 1.5                  0/12

That falsifies the standing hypothesis, recorded in the litellm config comment
and in the operator's own 2026-08-16 note, that presence_penalty 1.5 is the
first dial to move. It is not this bug's cause.

It also explains the blast radius: only the two temp-0.7 aliases leak, `gen`
and `summarizer-large`. summarizer, classifier, image-judge and qwen-image-bench
all run at temp 0 and are clean, so nevermore's summarizer path is unaffected.

Candidate fix, validated n=30 over 4 prompt types plus a 3-turn conversation:
chat_template_kwargs {enable_thinking: true, reasoning_effort: low} takes 8/30
leaks to 0/30, at ~+27% completion tokens and a ~3% empty-content residual.

The tell appears in eval_coldfusion_h300.json and in none of the aeon, heresy,
mixed or w4a16 evals, so it is new with this build -- but L35 was never evaled,
so this does not separate a Cold-Fusion base trait from a Heretic-300
abliteration artifact.

Reproducers and the full method land in bench/think-leak/. Note in particular
that the 7/7 alias smoke test run at cutover structurally could not catch this:
trivial prompts never invite reasoning, so they never sample the leaking token.
2026-08-21 00:12:30 -07:00

35 lines
2.0 KiB
Python

import json, urllib.request, collections
KEY=open('/home/lkraven/.config/litellm/infra-ops-key').read().strip()
URL="http://10.250.50.70:4000/v1/chat/completions"
PROMPTS=[
("reasoning","A farmer has 17 sheep. All but 9 run away. He buys twice as many as he has left, then sells 4. How many now? Explain."),
("summarize","Summarise these in one sentence each:\n1. Fed holds rates, signals two cuts\n2. Port strike enters third week\n3. Study links microplastics to soil carbon loss"),
("classify","Classify the sentiment of each line as POSITIVE, NEGATIVE or NEUTRAL:\n1. The deploy finally went clean.\n2. Third outage this week.\n3. Meeting moved to Thursday."),
]
def probe(model,n=4,maxtok=2048):
t=collections.Counter(); ex=None
for label,p in PROMPTS:
for i in range(n):
body={"model":model,"messages":[{"role":"user","content":p}],"max_tokens":maxtok}
req=urllib.request.Request(URL,data=json.dumps(body).encode(),
headers={"Authorization":"Bearer "+KEY,"Content-Type":"application/json"})
try:
d=json.loads(urllib.request.urlopen(req,timeout=300).read().decode(),strict=False)
except Exception as e:
t["error"]+=1; continue
c=d["choices"][0]; cont=c["message"].get("content") or ""
t["n"]+=1
if "<think>" in cont:
t["LEAK"]+=1; t[f"leak_{label}"]+=1
if ex is None: ex=(label,cont[:120])
if not cont.strip(): t["empty"]+=1
if c.get("finish_reason")!="stop": t[f"finish_{c.get('finish_reason')}"]+=1
return t,ex
for m in ["gen","summarizer","summarizer-large","classifier","image-judge","qwen-image-bench","gen-reasoning"]:
t,ex=probe(m)
leak=t.get("LEAK",0); n=t.get("n",0)
pct=(100.0*leak/n) if n else 0
print(f" {m:<20} n={n:<3} LEAK={leak:<3} ({pct:4.1f}%) empty={t.get('empty',0)} " + " ".join(f"{k}={v}" for k,v in t.items() if k.startswith(('leak_','finish_'))))
if ex: print(f" e.g. [{ex[0]}] {ex[1]!r}")