feat(aeon): deploy Qwen3.6-27B AEON as gen + char-rp, displace qwopus

New stacks/qwen36-27b-aeon: two co-located vLLM serves on ana-ml2 GPU0 —
gen (:8015, MTP off) and an RP seat (:8016, native MTP) — dense Qwen3.6-27B
(qwen3_5 GDN-hybrid, uncensored/abliterated), ModelOpt-NVFP4, multimodal,
256K context, depends_on-sequenced util split (~0.50/0.45). Each serve
carries a base + `-thinking` served-name so the `-reasoning` gateway records
target distinct LiteLLM deployments — otherwise a thinking-off request mutates
the shared litellm_params and clobbers enable_thinking (the shared-config
footgun that silently disabled char-rp-reasoning).

Gateway (stacks/litellm/conf/config.yaml): gen / gen-reasoning /
summarizer-large -> AEON :8015; char-rp / char-rp-reasoning added -> RP seat
:8016 (Qwen-RP sampler recs); gen-reasoning -> `-thinking`, char-rp-reasoning
-> `-rp-thinking`. Retired qwen3.5-122-a10b[-reasoning] + qwen-large[-reasoning]
(qwopus displaced; those named a 122B that no longer serves gen).
This commit is contained in:
vh
2026-07-06 00:53:38 -07:00
parent 993decf3eb
commit e6ab51c74a
3 changed files with 333 additions and 81 deletions
+70 -81
View File
@@ -54,13 +54,13 @@ model_list:
model_info:
mode: chat
# alias: summarizer-large -> gen / qwen3.5-122-a10b (operator 2026-06-19). For heavier
# summarization that wants the 122B Qwopus instead of granite-8b. Thinking OFF (matches
# gen). Keep api_base (:8013) + enable_thinking in sync with the gen record below.
# alias: summarizer-large -> gen / qwen3.6-27b-aeon (operator 2026-07-05). For heavier
# summarization that wants the AEON 27B `gen` model instead of granite-8b. Thinking OFF
# (matches gen). Keep api_base (:8015) + enable_thinking in sync with the gen record below.
- model_name: summarizer-large
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
model: hosted_vllm/qwen3.6-27b-aeon
api_base: http://10.250.50.54:8015/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.7
@@ -104,82 +104,24 @@ model_list:
model_info:
mode: chat
# --- Qwopus3.5-122B-A10B (Kimi-distilled, abliterated, NVFP4, VISION-INTACT) — the
# general / `gen` model on ana-ml2 GPU 0. Replaced the bjk110 text-only qwen3.5-122b
# 2026-06-19 (which had replaced mistral-small-4). Served on :8013 via vLLM as plain
# multimodal (no text-only patch), served-name qwen3.5-122-a10b — so these records
# route UNCHANGED. Full 256K (262144) @ fp8 KV + CUDA graphs (92.7 tok/s warm);
# tool-calling via qwen3_coder. Thinking split = chat_template_kwargs.enable_thinking
# + --reasoning-parser qwen3. One upstream fanned out under qwen3.5-122-a10b[-reasoning]
# + aliases qwen-large[-reasoning] + gen[-reasoning]; -reasoning variants enable
# thinking. Keep api_base in sync.
# presence_penalty: 1.0 on ALL these qwen3.5-122-a10b records (+ summarizer-large
# above) — anti-repetition-loop damper for the abliterated/NVFP4 tendency (operator
# 2026-06-27). Gateway-tunable default (callers can override); bake the validated
# value into the vLLM serving def (stacks/qwen3.5-122b, --override-generation-config)
# once confirmed, to also cover direct (non-gateway) callers.
# ⚠️ Worldtree CHARACTER backend (was bound to mistral-small-4) is dark until
# repointed — operator-acknowledged. ---
- model_name: qwen3.5-122-a10b
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.7
top_p: 0.8
extra_body:
top_k: 20
chat_template_kwargs:
enable_thinking: false
model_info:
mode: chat
- model_name: qwen3.5-122-a10b-reasoning
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.6
top_p: 0.95
extra_body:
top_k: 20
chat_template_kwargs:
enable_thinking: true
model_info:
mode: chat
- model_name: qwen-large
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.7
top_p: 0.8
extra_body:
top_k: 20
chat_template_kwargs:
enable_thinking: false
model_info:
mode: chat
- model_name: qwen-large-reasoning
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.6
top_p: 0.95
extra_body:
top_k: 20
chat_template_kwargs:
enable_thinking: true
model_info:
mode: chat
# --- Qwen3.6-27B AEON (uncensored/abliterated, NVFP4 ModelOpt, VISION-INTACT image+video)
# — the general / `gen` model on ana-ml2 GPU 0. REPLACED qwopus3.5-122b 2026-07-05
# (operator: displace qwopus, this assumes gen/gen-reasoning). Dense 27B, qwen3_5
# GDN-hybrid (qwopus's little sibling), native MTP head. Served on :8015 via vLLM,
# served-name qwen3.6-27b-aeon, MTP OFF (spec-decode hurts concurrent aggregate — the
# MTP twin is char-rp below). Thinking split = chat_template_kwargs.enable_thinking +
# --reasoning-parser qwen3; tool-calling qwen3_coder. gen / gen-reasoning + summarizer-
# large route here; -reasoning enables thinking. Keep api_base (:8015) in sync.
# RETIRED with the displacement (→ 404, callers migrate to gen): qwen3.5-122-a10b
# [-reasoning] + qwen-large[-reasoning] — they named a 122B that no longer exists;
# aliasing a 27B under those is the naming footgun the qwen36-vl stack warns against.
# presence_penalty: 1.0 INHERITED from qwopus (same-family abliterated/NVFP4 anti-
# repetition damper, operator 2026-06-27) — RE-VALIDATE for AEON; NOT yet confirmed
# for this model's repetition behavior. ---
- model_name: gen
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
model: hosted_vllm/qwen3.6-27b-aeon
api_base: http://10.250.50.54:8015/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.7
@@ -192,8 +134,10 @@ model_list:
mode: chat
- model_name: gen-reasoning
litellm_params:
model: hosted_vllm/qwen3.5-122-a10b
api_base: http://10.250.50.54:8013/v1
# Distinct served-name so a thinking-off `gen` request can't mutate this deployment's
# enable_thinking (shared-config-mutation footgun). Same backend :8015, different model id.
model: hosted_vllm/qwen3.6-27b-aeon-thinking
api_base: http://10.250.50.54:8015/v1
api_key: os.environ/VLLM_API_KEY
presence_penalty: 1.0
temperature: 0.6
@@ -204,6 +148,51 @@ model_list:
enable_thinking: true
model_info:
mode: chat
# char-rp -> the native-MTP single-seat RP twin (:8016, served qwen3.6-27b-aeon-rp).
# SAME weights as gen, MTP ON (qwen3_5_mtp n=3) for single-stream RP latency. RP-shaped
# sampler defaults (Worldtree character-rp role / callers override); thinking OFF.
# ⚠️ MTP silently drops min_p/logit_bias — don't rely on those through char-rp.
- model_name: char-rp
litellm_params:
model: hosted_vllm/qwen3.6-27b-aeon-rp
api_base: http://10.250.50.54:8016/v1
api_key: os.environ/VLLM_API_KEY
# Qwen3.x non-thinking RP recs (operator 2026-07-05, adapted from Qwen + community
# RP testing). presence_penalty light (0.1) to reduce topic drift; repetition_penalty
# 1.05. min_p skipped (Qwen rec + MTP drops it anyway). DRY off (Qwen3.x artifacts).
temperature: 0.7
top_p: 0.8
presence_penalty: 0.1
extra_body:
top_k: 20
repetition_penalty: 1.05
chat_template_kwargs:
enable_thinking: false
model_info:
mode: chat
# char-rp-reasoning -> same RP seat (:8016, MTP), thinking ON. Reasoning-profile
# sampler (lower temp than char-rp for coherent thought); tunable. ⚠️ The thinking
# TRACE does not yet surface in reasoning_content (chat_template injects <think> in the
# prompt → parser drops the span); pending a template fix, NOT a parser swap.
- model_name: char-rp-reasoning
litellm_params:
# Distinct served-name so a thinking-off char-rp request can't clobber this to
# enable_thinking:false (the bug that broke it). Same backend :8016, different model id.
model: hosted_vllm/qwen3.6-27b-aeon-rp-thinking
api_base: http://10.250.50.54:8016/v1
api_key: os.environ/VLLM_API_KEY
# Same RP profile as char-rp but thinking-mode top_p 0.95 (Qwen thinking rec) + focused
# temp 0.6. (Reasoning-trace surfacing still pending the parser/template fix.)
temperature: 0.6
top_p: 0.95
presence_penalty: 0.1
extra_body:
top_k: 20
repetition_penalty: 1.05
chat_template_kwargs:
enable_thinking: true
model_info:
mode: chat
# --- Selene 1 Mini 8B (AtlaAI judge, FP8) — restored on GPU1 after the
# llama-swap teardown (was the Q6_K GGUF in the swap zoo). vLLM dynamic fp8,