diff --git a/docs/pfi/model-quantization-playbook.md b/docs/pfi/model-quantization-playbook.md index 89d0483..53b052a 100644 --- a/docs/pfi/model-quantization-playbook.md +++ b/docs/pfi/model-quantization-playbook.md @@ -131,6 +131,39 @@ There is also a **modelopt-format-specific** version of this: vLLM 0.24 does not modelopt `exclude_modules` to the spec-decode *draft* model, which no checkpoint config can fix (needs a `sitecustomize` runtime patch). Using compressed-tensors avoids it entirely — §3.4. +### 3.7 ⭐ A LOADED MTP head can still corrupt output — Qwen3.8 multi-turn + +§3.3 is about *losing* the head (0% acceptance, silent). This is the opposite and +worse failure: the head loads, acceptance looks healthy, single-turn output is +perfect — and then it **corrupts multi-turn conversations** once cumulative context +passes **~2,000 tokens**. The reply collapses in length *and* bleeds earlier turns +into the current answer (a "describe durian" reply that contained the Krebs-cycle +and winter answers from three turns back). Single-turn probes and the acceptance +gate (§5) **do not catch it** — it only appears as accumulated context grows. + +Isolated 2026-08-16 (operator-confirmed), each step measured on a fixed 7-turn probe: + +- **Not the serving gateway, not sampling, not repetition/template.** Identical + input gateway-vs-direct behaves the same; presence_penalty 1.5/0.5/0.0 all + collapse; higher temperature collapses harder; a conversation of *unrelated* + topics collapses at the same ~2k tokens as a repetitive one → it is context- + length-driven, not template lock-in. +- **Model-independent across every Qwen3.8-27B quant** (AEON W4A4, unsloth + FP8-attn, our in-house mixed) — so not a quant-brand or scheme artifact. +- **DECISIVE: same model + same conversation, MTP OFF → coherent through 4k+ + tokens, zero bleed.** Toggle it back on → collapse returns. MTP is the cause. + +**Qwen3.6-27B running the same `qwen3_5_mtp` method is CLEAN.** So the 3.6 MTP +head/graft is fine and the 3.8 one is not — suspects: the bf16 graft being subtly +wrong for the 3.8 head, or the vLLM `qwen3_5_mtp` impl diverging at `num_speculative_tokens=3`. +Open upstream question (queried dvalin/bil-smithy 2026-08-17). + +**Rule: gate MTP on a MULTI-TURN coherence probe, not just single-shot acceptance.** +Run a 7-turn varied-topic conversation and watch turns past ~2k cumulative tokens +for length-collapse and cross-turn bleed. If it corrupts: disable MTP (costs ~half +decode tok/s) or try `num_speculative_tokens=1` (shallowest draft, least divergence +— untested as of this writing). The gen seat runs MTP **off** as of 2026-08-17. + ### 3.4 Toolchain version deadlocks Both directions have burned us, so the resolution is: **use llm-compressor / compressed-tensors,