7bd38b33b5
Root cause of the long-hunted 'gen goes degenerate in conversation',
isolated 2026-08-16 and operator-confirmed. qwen3_5_mtp speculative
decoding corrupts Qwen3.8-27B output once cumulative multi-turn context
passes ~2,000 tokens: the draft head's bad tokens get accepted and the
reply degenerates into CONTEXT-BLEEDING (a 'describe durian' answer that
contained the Krebs-cycle and winter replies from earlier turns), then
collapses to a few words.
Isolation, each step measured on the varied 7-turn probe:
- not the gateway (identical input -> gateway == direct; echo intact)
- not presence_penalty (1.5/0.5/0.0 all collapse), not temperature
(1.0 collapses harder), not repetition (varied unrelated topics
collapse identically -> it is context length, not template-lock)
- model-INDEPENDENT across all three Qwen3.8-27B quants we serve
(AEON W4A4, unsloth FP8-attn, in-house mixed)
- Qwen3.6 (char-rp-reasoning) and Gemma-4 (char-rp) are CLEAN
- DECISIVE: same Qwen3.8 model + same conversation, MTP OFF -> coherent
through 4k+ tokens, no bleed. MTP is the cause.
Qwen3.6 runs the same qwen3_5_mtp method and is clean, so the 3.6 MTP
head/graft is fine and the 3.8 one is not (suspects: the bf16 MTP graft,
or spec depth 3). COST: ~half decode tok/s without spec decoding.
Accepted as known-good until the 3.8 MTP is fixed; first thing to try on
re-enable is num_speculative_tokens=1. Seat restored to AEON W4A4 (the
production choice); verified clean on the varied series after this change.