d28a371049
Root cause of the multi-day degeneration hunt, operator-confirmed: the AEON NVFP4 W4A4 quant (sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4, full W4A4 incl. attention) went degenerate ~15-20% of generations in real multi-turn use and forced regenerates. MTP, prefix-caching, and the gateway all merely AMPLIFIED it, which is why MTP-off, APC-off, and the vLLM #51113 fix each 'helped' a synthetic probe without fixing it -- three plausible false root-causes, each passing one clean run then failing in real use. The fix was the WEIGHTS: the in-house JonathanColetti/Heretic mixed NVFP4+FP8 build (qwen38-27b-uncensored-nvfp4-mixed, FP8 attention not W4A4, same base, same MTP) is coherent through long multi-turn with MTP ON. W4A4 *attention* was the defect; FP8 attention is not. This commit: - GEN_MODEL -> the mixed FP8-attn build (primary gen until DavidAU 3.8 lands) - GEN_IMAGE pinned to vllm/vllm-openai:nightly-311b3513... (v0.27.2rc1.dev150, carries #51113; pinned by sha so it does not drift on the next pull) - AEON weights PURGED from /tank (no-good), safety-checked not-in-use first - playbook 3.8: the stochastic-W4A4-degeneration lesson + isolate-weights-early + do-not-declare-a-fix-from-one-probe (it validated three non-fixes) AEON is re-pullable from HF if ever needed, but the operator ruled it no-good.