feat(qwen3.5-122b): replace mistral-small-4 as gen (abliterated NVFP4, text-only)

bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4 on ana-ml2 GPU 0 (heretic downed):
- stacks/qwen3.5-122b/ — vLLM serve via the repo's text-only patch (Qwen3.5 MoE is a
  multimodal arch but this checkpoint is text-only weights), --reasoning-parser qwen3,
  GPU 0 pin, :8013; entrypoint+patch mounted from the model dir.
- serve-qwen3.5-122b.yaml — displace heretic + serve + verify.
- litellm: REMOVED dead mistral-small-4 / -reasoning; added qwen3.5-122-a10b[-reasoning]
  + aliases qwen-large[-reasoning] + repointed gen[-reasoning] -> qwen (thinking split via
  chat_template_kwargs.enable_thinking + --reasoning-parser qwen3).

Verified live: qwen healthy on :8013; gen / qwen-large / qwen3.5-122-a10b route, and
gen-reasoning returns reasoning_content; mistral-small-4 removed.
NOTE: Worldtree character backend (was bound to mistral-small-4) is dark until repointed
(operator-acknowledged).
This commit is contained in:
vh
2026-06-19 00:49:04 -07:00
parent 67102b5b94
commit 89c83c4271
4 changed files with 229 additions and 31 deletions
+23
View File
@@ -0,0 +1,23 @@
# qwen3.5-122b (bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4, text-only) — ana-ml2
# GPU 0 tunables. Real .env lives at /opt/docker/compose/qwen3.5-122b/.env.
# vllm/vllm-openai:latest per the repo (the text-only patch targets latest).
# MUTABLE tag — pin a digest once a known-good version is established.
QWEN35_IMAGE=vllm/vllm-openai:latest
QWEN35_CONTAINER_NAME=vllm-qwen35-122b
# Own port (8010=mistral heretic [downed], 8007=qwen36, 8011=selene — 8013 free).
QWEN35_PORT=8013
QWEN35_GPU_ID=0
# vLLM served-model-name; litellm fans out qwen3.5-122-a10b[-reasoning] + aliases
# (qwen-large, gen) onto this, differentiated by chat_template_kwargs.enable_thinking.
QWEN35_SERVED_NAME=qwen3.5-122-a10b
QWEN35_MAX_MODEL_LEN=131072
QWEN35_MAX_NUM_SEQS=4
QWEN35_GPU_MEM_UTIL=0.90
QWEN35_MAX_NUM_BATCHED_TOKENS=32768
# Optional upstream vLLM API key (empty = no auth; internal net only).
API_KEY=