feat(qwen36-vl): swap FP8→NVFP4 + GPU1 rebalance (granite restored)
The nvidia ModelOpt NVFP4 MoE that failed on vLLM 0.19.1/0.22.0 (#44081) loads clean on 0.23.0. Cut prod qwen36 FP8→NVFP4: ~20.4 GiB weights vs ~34 (~40% lighter, ~13 GB reclaimed on GPU 1), faster single-stream on Blackwell FP4 cores, vision tower preserved (comfy-dev real anatomy-judge A/B on 16 prod images: PASS; brokkr text/speed: parity bar a minor multi-step-chained-reasoning slip that doesn't bite the judge role). - compose: pin image by 0.23.0 digest, drop --kv-cache-dtype fp8 (fp16 KV — the freed room buys full-precision KV), util 0.46→0.32. - GPU1 rebalance (pinned): granite restored 0.24→0.34 / 65536→131072 (undoes the FP8-era sacrifice); trio unchanged; total ~0.82, ~24 GB free. - gateway model name qwen3.6-35b-a3b unchanged (now NVFP4 behind it); thinking-split (enable_thinking=false default) intact — the judge needs it.
This commit is contained in:
@@ -1,34 +1,35 @@
|
||||
# Qwen3.6-35B-A3B VL (official FP8) on ana-ml2 — copy to .env on the host and fill.
|
||||
# Qwen3.6-35B-A3B VL (official NVFP4) on ana-ml2 — copy to .env on the host and fill.
|
||||
# Real .env lives on ana-ml2 at /opt/docker/compose/qwen36-vl/.env (gitignored).
|
||||
#
|
||||
# Replaces qwen35-vl (Qwen3.5-9B) 2026-06-14. See compose.yaml header for the
|
||||
# FP8-over-NVFP4 rationale (vLLM NVFP4 MoE loader broken, #44081) and why this
|
||||
# uses plain :latest with NO --quantization (pre-quantized checkpoint; a forced
|
||||
# flag would noise-quantize the vision tower like the old qwen35-vl).
|
||||
# Swapped FP8→NVFP4 2026-06-15 (the nvidia ModelOpt NVFP4 MoE now loads on vLLM
|
||||
# 0.23.0 — #44081 fixed). See compose.yaml header for the full rationale + the
|
||||
# comfy-dev vision A/B that cleared it + the GPU-1 rebalance.
|
||||
|
||||
# :latest is fine — the official FP8 checkpoint loads + serves vision correctly
|
||||
# on 0.19.1 (validated 2026-06-14). NO pinned nightly digest needed.
|
||||
QWEN_IMAGE=vllm/vllm-openai:latest
|
||||
# PINNED by digest — NVFP4 needs vLLM >= 0.23.0 (the related ModelOpt NVFP4 MoE
|
||||
# path broke on 0.19.1/0.22.0). Pin guards against a :latest regression. This
|
||||
# digest = 0.23.0, validated to load + serve this checkpoint (vision incl.).
|
||||
QWEN_IMAGE=vllm/vllm-openai@sha256:6d8429e38e3747723ca07ee1b17972e09bb9c51c4032b266f24fb1cc3b22ed8f
|
||||
|
||||
QWEN_CONTAINER_NAME=vllm-qwen36
|
||||
QWEN_MODEL=Qwen/Qwen3.6-35B-A3B-FP8
|
||||
QWEN_MODEL=nvidia/Qwen3.6-35B-A3B-NVFP4
|
||||
QWEN_PORT=8007
|
||||
|
||||
# GPU 1 = shared with the granite summarizer + embed/rerank/reward trio.
|
||||
# GPU 0 is kept free for the llama-swap creative-writing hot-swap card.
|
||||
# (GPU 0 now hosts Mistral Small 4, not the old llama-swap card.)
|
||||
QWEN_GPU_ID=1
|
||||
|
||||
# util 0.42 (~40 GB) — official FP8 weights load in ~34.2 GiB; 0.42 covers
|
||||
# weights + CUDA-graph + a generous KV pool. Hybrid attn (10 of 40 layers full-
|
||||
# attn, ~10 KB/tok KV) makes long context nearly free, so max-len is generous.
|
||||
# Budget (2026-06-14): qwen36 0.42 + granite 0.28 + trio 0.20 = 0.90 total,
|
||||
# ~10 GB graph headroom. Bring qwen36 up LAST so capture sees the free room.
|
||||
QWEN_GPU_MEM_UTIL=0.42
|
||||
# util 0.32 (~31 GB) — NVFP4 weights load in ~20.4 GiB; 0.32 covers weights +
|
||||
# fp16 KV + CUDA-graph. fp16 KV (compose drops --kv-cache-dtype fp8): the NVFP4
|
||||
# swap freed enough room to run full-precision KV. Hybrid attn (10/40 full-attn)
|
||||
# keeps even fp16 KV cheap. GPU-1 budget (2026-06-15 rebalance, pinned): qwen36
|
||||
# 0.32 + granite 0.34/131072 (RESTORED from the FP8-era 0.24/64K) + trio 0.16 =
|
||||
# ~0.82, ~24 GB free headroom. Recreate ONE service at a time (profiling race).
|
||||
QWEN_GPU_MEM_UTIL=0.32
|
||||
QWEN_MAX_MODEL_LEN=131072
|
||||
# Sampler-warmup OOM guard on the shared GPU (248K vocab × default 1024 seqs is
|
||||
# a huge transient). 32 is plenty for a vision endpoint.
|
||||
QWEN_MAX_NUM_SEQS=32
|
||||
|
||||
# Optional
|
||||
# Optional — checkpoint is ungated.
|
||||
HF_TOKEN=
|
||||
API_KEY=
|
||||
|
||||
Reference in New Issue
Block a user