e6ab51c74a
New stacks/qwen36-27b-aeon: two co-located vLLM serves on ana-ml2 GPU0 — gen (:8015, MTP off) and an RP seat (:8016, native MTP) — dense Qwen3.6-27B (qwen3_5 GDN-hybrid, uncensored/abliterated), ModelOpt-NVFP4, multimodal, 256K context, depends_on-sequenced util split (~0.50/0.45). Each serve carries a base + `-thinking` served-name so the `-reasoning` gateway records target distinct LiteLLM deployments — otherwise a thinking-off request mutates the shared litellm_params and clobbers enable_thinking (the shared-config footgun that silently disabled char-rp-reasoning). Gateway (stacks/litellm/conf/config.yaml): gen / gen-reasoning / summarizer-large -> AEON :8015; char-rp / char-rp-reasoning added -> RP seat :8016 (Qwen-RP sampler recs); gen-reasoning -> `-thinking`, char-rp-reasoning -> `-rp-thinking`. Retired qwen3.5-122-a10b[-reasoning] + qwen-large[-reasoning] (qwopus displaced; those named a 122B that no longer serves gen).