# Qwen3.6-35B-A3B VL (official FP8) on ana-ml2 — copy to .env on the host and fill. # Real .env lives on ana-ml2 at /opt/docker/compose/qwen36-vl/.env (gitignored). # # Replaces qwen35-vl (Qwen3.5-9B) 2026-06-14. See compose.yaml header for the # FP8-over-NVFP4 rationale (vLLM NVFP4 MoE loader broken, #44081) and why this # uses plain :latest with NO --quantization (pre-quantized checkpoint; a forced # flag would noise-quantize the vision tower like the old qwen35-vl). # :latest is fine — the official FP8 checkpoint loads + serves vision correctly # on 0.19.1 (validated 2026-06-14). NO pinned nightly digest needed. QWEN_IMAGE=vllm/vllm-openai:latest QWEN_CONTAINER_NAME=vllm-qwen36 QWEN_MODEL=Qwen/Qwen3.6-35B-A3B-FP8 QWEN_PORT=8007 # GPU 1 = shared with the granite summarizer + embed/rerank/reward trio. # GPU 0 is kept free for the llama-swap creative-writing hot-swap card. QWEN_GPU_ID=1 # util 0.42 (~40 GB) — official FP8 weights load in ~34.2 GiB; 0.42 covers # weights + CUDA-graph + a generous KV pool. Hybrid attn (10 of 40 layers full- # attn, ~10 KB/tok KV) makes long context nearly free, so max-len is generous. # Budget (2026-06-14): qwen36 0.42 + granite 0.28 + trio 0.20 = 0.90 total, # ~10 GB graph headroom. Bring qwen36 up LAST so capture sees the free room. QWEN_GPU_MEM_UTIL=0.42 QWEN_MAX_MODEL_LEN=131072 # Sampler-warmup OOM guard on the shared GPU (248K vocab × default 1024 seqs is # a huge transient). 32 is plenty for a vision endpoint. QWEN_MAX_NUM_SEQS=32 # Optional HF_TOKEN= API_KEY=