1a36e60d3a
Found while answering a question from the operator, relayed via brokkr-smithy-dev, about whether a recorded Qwen3.8 degeneracy at ~1,700 tokens relates to a length sensitivity just measured on the tuned Gemma-4. The record is §3.7 and the number is ~2,000 -- but reading it to answer that question surfaced that the section is stale. §3.7 presented "disable prefix caching, keep MTP" as THE MITIGATION, resolved 2026-08-17, and stated the gen seat runs that config. It does not and has not since that same day: APC-off passed a synthetic 7-turn probe and the operator still saw severe degeneration in real use, so it was reverted. The multi-day hunt resolved to the AEON W4A4 quant being defective, with MTP / prefix-caching / gateway merely amplifying it (§3.8 records the corrected causal story; §3.7 was never updated to match). Verified against the live container rather than against the compose file alone: vllm-gen runs --enable-prefix-caching with qwen3_5_mtp / num_speculative_tokens 3. stacks/gen-seat/compose.yaml carries the full corrected history inline and is the current authority. §3.7's superseded text is kept and fenced rather than deleted -- it is the history of a mitigation that looked right and was not. Added a dated row to §7 per the standing rule that a wrong playbook claim gets a superseded-claims entry, not just a fix. The lesson inside the lesson is worth more than the correction: §3.7's own standing rule is "gate MTP on a multi-turn coherence probe, not just single-shot acceptance." The APC-off mitigation was gated on exactly that probe, passed it, and still failed in real use -- the multi-turn probe was itself too small to gate on. A passing probe is not sufficient evidence at any size that has not been calibrated against real use.