feat(ana-ml2): replace Qwen3.5-9B vision with Qwen3.6-35B-A3B FP8 on GPU 1
Retire qwen35-vl (Qwen3.5-9B); add qwen36-vl serving the official FP8 Qwen3.6-35B-A3B vision MoE on :8007 under its TRUE name only — no alias. qwen3.5-9b-fp8 is killed at vLLM AND the litellm gateway (404/400); a model is never served under a prior model's name. Consumer (comfy-dev/arbo) notified + migrated; arbo vkeys flipped to all-proxy-models; shared all-agents-local key repointed to qwen3.6-35b-a3b. GPU-1 rebalance for the heavier FP8 weights (~34 GB): granite 0.35->0.24 / 131K->64K, embed/rerank 0.05->0.03 (reclaimed util-reservation waste). Verified: vision correct, 20-concurrent/endpoint load test = no OOM (~7.5 GB headroom). Drop the llama-swap qwen3.5-9b GPU-0 pin (GPU 0 freed for the creative-writing hot-swap card). NVFP4 was the lighter fit (~21 GB) but its vLLM ModelOpt-MoE loader is broken (KeyError w2_input_scale / lm_head.input_scale, vllm #44081); revisit when fixed.
This commit is contained in:
@@ -0,0 +1,102 @@
|
||||
# qwen36-vl — Qwen3.6-35B-A3B vision-language MoE (official FP8) on ana-ml2.
|
||||
#
|
||||
# Replaces the qwen35-vl stack (Qwen3.5-9B) 2026-06-14. Co-located on GPU 1 with
|
||||
# the granite summarizer + embed/rerank/reward trio (GPU 0 stays free for the
|
||||
# llama-swap creative-writing hot-swap card). Serves on :8007.
|
||||
#
|
||||
# WHY official FP8 (not NVFP4): NVFP4 (nvidia/Qwen3.6-35B-A3B-NVFP4, ~21 GB) is the
|
||||
# lighter fit but its vLLM ModelOpt-MoE loader is BROKEN as of 0.19.1/0.22.0
|
||||
# (KeyError w2_input_scale / lm_head.input_scale — vLLM #44081). The official
|
||||
# Qwen pre-quantized FP8 (~34 GB weights) loads clean on :latest and — unlike
|
||||
# the old qwen35-vl — needs NO pinned nightly digest: that hack existed because
|
||||
# vLLM DYNAMIC `--quantization fp8` quantized the vision tower to noise. This
|
||||
# checkpoint is PRE-quantized, so we OMIT --quantization (vLLM auto-detects the
|
||||
# checkpoint's own fp8) and the vision tower is preserved. Validated 2026-06-14
|
||||
# on GPU 0: loads in 34.2 GiB, image test returns correct ("Blue"). Revisit
|
||||
# NVFP4 (frees ~13 GB) once vLLM's loader is fixed.
|
||||
#
|
||||
# NAMING: served ONLY as its TRUE name `qwen3.6-35b-a3b`. A model is never aliased
|
||||
# under a prior model's name — a caller asking for `qwen3.5-9b-fp8` (a 9B dense)
|
||||
# must NOT be silently handed this 35B-A3B MoE; that's a downstream-confusion
|
||||
# footgun. The legacy `qwen3.5-9b-fp8` name is RETIRED. Consumers (Arbo's vision
|
||||
# hero-judge, stacks/arbo v0.11.3+) migrate to `qwen3.6-35b-a3b` — they 404 on the
|
||||
# old name until they repoint, which is the correct loud signal (notified 2026-06-14).
|
||||
#
|
||||
# WHY util 0.42 / max-len 131072: hybrid attn (10 of 40 layers full-attn, ~10 KB/
|
||||
# tok KV) → KV is cheap, so big context is nearly free; the 34 GB weights are the
|
||||
# cost. 0.42 (~40 GB) = weights + graph + generous KV. Granite drops to 0.25/64K
|
||||
# to make room (the FP8-vs-maxed-granite tradeoff, operator-approved 2026-06-14).
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
name: qwen36-vl
|
||||
|
||||
services:
|
||||
vllm-qwen36:
|
||||
image: ${QWEN_IMAGE}
|
||||
container_name: ${QWEN_CONTAINER_NAME}
|
||||
restart: unless-stopped
|
||||
ipc: host
|
||||
ports:
|
||||
- "${QWEN_PORT}:8000"
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/hfcache
|
||||
environment:
|
||||
- HF_HOME=/hfcache
|
||||
- HF_HUB_CACHE=/hfcache/hub
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_API_KEY=${API_KEY:-}
|
||||
command:
|
||||
- ${QWEN_MODEL}
|
||||
# Pre-quantized FP8 checkpoint → NO --quantization (vLLM auto-detects; a
|
||||
# forced flag would re-quantize the vision tower to noise, see header).
|
||||
- --served-model-name
|
||||
- qwen3.6-35b-a3b
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8000"
|
||||
- --gpu-memory-utilization
|
||||
- ${QWEN_GPU_MEM_UTIL}
|
||||
- --max-model-len
|
||||
- ${QWEN_MAX_MODEL_LEN}
|
||||
# Cap concurrency: vLLM warms the sampler with max_num_seqs dummy requests,
|
||||
# and this model's 248K vocab makes that warmup tensor huge — the default
|
||||
# 1024 OOMs on a shared GPU even though weights+KV fit. 32 is ample for a
|
||||
# vision endpoint (the summarizer carries the concurrency, not this).
|
||||
- --max-num-seqs
|
||||
- ${QWEN_MAX_NUM_SEQS}
|
||||
- --kv-cache-dtype
|
||||
- fp8
|
||||
- --trust-remote-code
|
||||
- --dtype
|
||||
- auto
|
||||
- --enable-prefix-caching
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids:
|
||||
- "${QWEN_GPU_ID}"
|
||||
capabilities:
|
||||
- gpu
|
||||
healthcheck:
|
||||
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 300s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=Qwen3.6-35B-A3B VL (FP8)
|
||||
- homepage.icon=mdi-image-search
|
||||
- homepage.description=Qwen3.6-35B-A3B vision-language MoE (FP8) via vLLM (ana-ml2)
|
||||
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
Reference in New Issue
Block a user