feat(ana-ml2): replace Qwen3.5-9B vision with Qwen3.6-35B-A3B FP8 on GPU 1
Retire qwen35-vl (Qwen3.5-9B); add qwen36-vl serving the official FP8 Qwen3.6-35B-A3B vision MoE on :8007 under its TRUE name only — no alias. qwen3.5-9b-fp8 is killed at vLLM AND the litellm gateway (404/400); a model is never served under a prior model's name. Consumer (comfy-dev/arbo) notified + migrated; arbo vkeys flipped to all-proxy-models; shared all-agents-local key repointed to qwen3.6-35b-a3b. GPU-1 rebalance for the heavier FP8 weights (~34 GB): granite 0.35->0.24 / 131K->64K, embed/rerank 0.05->0.03 (reclaimed util-reservation waste). Verified: vision correct, 20-concurrent/endpoint load test = no OOM (~7.5 GB headroom). Drop the llama-swap qwen3.5-9b GPU-0 pin (GPU 0 freed for the creative-writing hot-swap card). NVFP4 was the lighter fit (~21 GB) but its vLLM ModelOpt-MoE loader is broken (KeyError w2_input_scale / lm_head.input_scale, vllm #44081); revisit when fixed.
This commit is contained in:
@@ -117,24 +117,10 @@ models:
|
||||
--presence-penalty 1.5
|
||||
--chat-template-kwargs '{"enable_thinking":true}'
|
||||
|
||||
"qwen3.5-9b":
|
||||
name: "Qwen 3.5 9B UD-Q4_K_XL"
|
||||
description: "Dense 9B model. Lightweight general-purpose chat and reasoning."
|
||||
ttl: 0 # pinned — member of the `pinned` group, never unloads
|
||||
cmd: |
|
||||
/app/llama-server
|
||||
--context-shift
|
||||
--model /models/unsloth_Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf
|
||||
--port ${PORT}
|
||||
--n-gpu-layers 999
|
||||
--ctx-size 32768
|
||||
--flash-attn on
|
||||
--temp 1.0
|
||||
--top-p 0.95
|
||||
--top-k 20
|
||||
--min-p 0.00
|
||||
--presence-penalty 1.5
|
||||
--chat-template-kwargs '{"enable_thinking":true}'
|
||||
# "qwen3.5-9b" REMOVED 2026-06-14 — the GPU-0 GGUF pin is dropped. General
|
||||
# chat + vision now served by Qwen3.6-35B-A3B (official FP8) on GPU 1 via vLLM
|
||||
# (stacks/qwen36-vl). GPU 0 is freed for the creative-writing hot-swap card.
|
||||
# The GGUF stays on disk (/models/unsloth_Qwen3.5-9B-GGUF/) if ever wanted.
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Qwen 3.6 — uses -hf syntax, reads from HF_HOME=/hfcache (host pre-download)
|
||||
@@ -630,25 +616,10 @@ groups:
|
||||
- "qwen3-embedding-0.6B"
|
||||
- "qwen3-reranker-0.6B"
|
||||
|
||||
# Pinned general-purpose / utility models. Coexist in VRAM, never
|
||||
# unload. Members also have ttl: 0 individually so idle-timeout can't
|
||||
# drop them.
|
||||
#
|
||||
# Current pins:
|
||||
# qwen3.5-9b — ~6 GB at Q4 + KV. General-purpose chat baseline.
|
||||
# VRAM budget: ~6 GB persistent in the pin slot. llama-swap is pinned to
|
||||
# GPU 0 (a single RTX PRO 6000 Blackwell, 96 GB), so this leaves ~90 GB
|
||||
# for whichever non-pinned model the user invokes alongside.
|
||||
#
|
||||
# granite-4-small WAS pinned here; removed 2026-06-04 — superseded by
|
||||
# phi4-mini (vLLM FP8, stacks/vllm → vllm-phi4). Freed ~24 GB (120K KV).
|
||||
#
|
||||
# qwen3.6-35-a3b WAS in this group; removed 2026-04-27 because its
|
||||
# ~29 GB at Q6_K_XL pushed concurrent loads OOM. Now lives outside
|
||||
# with ttl: 0 — never idle-unloads but evictable under memory pressure.
|
||||
"pinned":
|
||||
swap: false
|
||||
exclusive: false
|
||||
persistent: true
|
||||
members:
|
||||
- "qwen3.5-9b"
|
||||
# Pinned group RETIRED 2026-06-14 — its only pin (qwen3.5-9b) was dropped, so
|
||||
# there is no active `pinned` group. General chat/vision moved to Qwen3.6-35B-
|
||||
# A3B FP8 on GPU 1 (vLLM), not a GGUF pin on GPU 0; GPU 0 is now fully free for
|
||||
# the creative-writing hot-swap card. Re-add a `pinned` group here if a
|
||||
# persistent GPU-0 model is ever wanted again.
|
||||
# (History: granite-4-small removed 2026-06-04 → phi4-mini vLLM FP8;
|
||||
# qwen3.6-35-a3b removed 2026-04-27 — ~29 GB Q6_K_XL OOM'd concurrent loads.)
|
||||
|
||||
Reference in New Issue
Block a user