feat(ana-ml2): replace Qwen3.5-9B vision with Qwen3.6-35B-A3B FP8 on GPU 1

Retire qwen35-vl (Qwen3.5-9B); add qwen36-vl serving the official FP8
Qwen3.6-35B-A3B vision MoE on :8007 under its TRUE name only — no alias.
qwen3.5-9b-fp8 is killed at vLLM AND the litellm gateway (404/400); a model is
never served under a prior model's name. Consumer (comfy-dev/arbo) notified +
migrated; arbo vkeys flipped to all-proxy-models; shared all-agents-local key
repointed to qwen3.6-35b-a3b.

GPU-1 rebalance for the heavier FP8 weights (~34 GB): granite 0.35->0.24 /
131K->64K, embed/rerank 0.05->0.03 (reclaimed util-reservation waste). Verified:
vision correct, 20-concurrent/endpoint load test = no OOM (~7.5 GB headroom).

Drop the llama-swap qwen3.5-9b GPU-0 pin (GPU 0 freed for the creative-writing
hot-swap card). NVFP4 was the lighter fit (~21 GB) but its vLLM ModelOpt-MoE
loader is broken (KeyError w2_input_scale / lm_head.input_scale, vllm #44081);
revisit when fixed.
This commit is contained in:
vh
2026-06-14 14:41:49 -07:00
parent b45d0cd86d
commit a0fed13801
8 changed files with 168 additions and 221 deletions
+11 -40
View File
@@ -117,24 +117,10 @@ models:
--presence-penalty 1.5
--chat-template-kwargs '{"enable_thinking":true}'
"qwen3.5-9b":
name: "Qwen 3.5 9B UD-Q4_K_XL"
description: "Dense 9B model. Lightweight general-purpose chat and reasoning."
ttl: 0 # pinned — member of the `pinned` group, never unloads
cmd: |
/app/llama-server
--context-shift
--model /models/unsloth_Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf
--port ${PORT}
--n-gpu-layers 999
--ctx-size 32768
--flash-attn on
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.00
--presence-penalty 1.5
--chat-template-kwargs '{"enable_thinking":true}'
# "qwen3.5-9b" REMOVED 2026-06-14 — the GPU-0 GGUF pin is dropped. General
# chat + vision now served by Qwen3.6-35B-A3B (official FP8) on GPU 1 via vLLM
# (stacks/qwen36-vl). GPU 0 is freed for the creative-writing hot-swap card.
# The GGUF stays on disk (/models/unsloth_Qwen3.5-9B-GGUF/) if ever wanted.
# --------------------------------------------------------------------------
# Qwen 3.6 — uses -hf syntax, reads from HF_HOME=/hfcache (host pre-download)
@@ -630,25 +616,10 @@ groups:
- "qwen3-embedding-0.6B"
- "qwen3-reranker-0.6B"
# Pinned general-purpose / utility models. Coexist in VRAM, never
# unload. Members also have ttl: 0 individually so idle-timeout can't
# drop them.
#
# Current pins:
# qwen3.5-9b — ~6 GB at Q4 + KV. General-purpose chat baseline.
# VRAM budget: ~6 GB persistent in the pin slot. llama-swap is pinned to
# GPU 0 (a single RTX PRO 6000 Blackwell, 96 GB), so this leaves ~90 GB
# for whichever non-pinned model the user invokes alongside.
#
# granite-4-small WAS pinned here; removed 2026-06-04 — superseded by
# phi4-mini (vLLM FP8, stacks/vllm → vllm-phi4). Freed ~24 GB (120K KV).
#
# qwen3.6-35-a3b WAS in this group; removed 2026-04-27 because its
# ~29 GB at Q6_K_XL pushed concurrent loads OOM. Now lives outside
# with ttl: 0 — never idle-unloads but evictable under memory pressure.
"pinned":
swap: false
exclusive: false
persistent: true
members:
- "qwen3.5-9b"
# Pinned group RETIRED 2026-06-14 — its only pin (qwen3.5-9b) was dropped, so
# there is no active `pinned` group. General chat/vision moved to Qwen3.6-35B-
# A3B FP8 on GPU 1 (vLLM), not a GGUF pin on GPU 0; GPU 0 is now fully free for
# the creative-writing hot-swap card. Re-add a `pinned` group here if a
# persistent GPU-0 model is ever wanted again.
# (History: granite-4-small removed 2026-06-04 → phi4-mini vLLM FP8;
# qwen3.6-35-a3b removed 2026-04-27 — ~29 GB Q6_K_XL OOM'd concurrent loads.)