Files
esh-pfi-infrastructure/stacks/qwen36-vl/.env.example
T
vh c6d76051a4 feat(qwen36-vl): swap FP8→NVFP4 + GPU1 rebalance (granite restored)
The nvidia ModelOpt NVFP4 MoE that failed on vLLM 0.19.1/0.22.0 (#44081)
loads clean on 0.23.0. Cut prod qwen36 FP8→NVFP4: ~20.4 GiB weights vs
~34 (~40% lighter, ~13 GB reclaimed on GPU 1), faster single-stream on
Blackwell FP4 cores, vision tower preserved (comfy-dev real anatomy-judge
A/B on 16 prod images: PASS; brokkr text/speed: parity bar a minor
multi-step-chained-reasoning slip that doesn't bite the judge role).

- compose: pin image by 0.23.0 digest, drop --kv-cache-dtype fp8 (fp16 KV
  — the freed room buys full-precision KV), util 0.46→0.32.
- GPU1 rebalance (pinned): granite restored 0.24→0.34 / 65536→131072
  (undoes the FP8-era sacrifice); trio unchanged; total ~0.82, ~24 GB free.
- gateway model name qwen3.6-35b-a3b unchanged (now NVFP4 behind it);
  thinking-split (enable_thinking=false default) intact — the judge needs it.
2026-06-15 17:28:15 -07:00

36 lines
1.7 KiB
Bash
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Qwen3.6-35B-A3B VL (official NVFP4) on ana-ml2 — copy to .env on the host and fill.
# Real .env lives on ana-ml2 at /opt/docker/compose/qwen36-vl/.env (gitignored).
#
# Swapped FP8→NVFP4 2026-06-15 (the nvidia ModelOpt NVFP4 MoE now loads on vLLM
# 0.23.0 — #44081 fixed). See compose.yaml header for the full rationale + the
# comfy-dev vision A/B that cleared it + the GPU-1 rebalance.
# PINNED by digest — NVFP4 needs vLLM >= 0.23.0 (the related ModelOpt NVFP4 MoE
# path broke on 0.19.1/0.22.0). Pin guards against a :latest regression. This
# digest = 0.23.0, validated to load + serve this checkpoint (vision incl.).
QWEN_IMAGE=vllm/vllm-openai@sha256:6d8429e38e3747723ca07ee1b17972e09bb9c51c4032b266f24fb1cc3b22ed8f
QWEN_CONTAINER_NAME=vllm-qwen36
QWEN_MODEL=nvidia/Qwen3.6-35B-A3B-NVFP4
QWEN_PORT=8007
# GPU 1 = shared with the granite summarizer + embed/rerank/reward trio.
# (GPU 0 now hosts Mistral Small 4, not the old llama-swap card.)
QWEN_GPU_ID=1
# util 0.32 (~31 GB) — NVFP4 weights load in ~20.4 GiB; 0.32 covers weights +
# fp16 KV + CUDA-graph. fp16 KV (compose drops --kv-cache-dtype fp8): the NVFP4
# swap freed enough room to run full-precision KV. Hybrid attn (10/40 full-attn)
# keeps even fp16 KV cheap. GPU-1 budget (2026-06-15 rebalance, pinned): qwen36
# 0.32 + granite 0.34/131072 (RESTORED from the FP8-era 0.24/64K) + trio 0.16 =
# ~0.82, ~24 GB free headroom. Recreate ONE service at a time (profiling race).
QWEN_GPU_MEM_UTIL=0.32
QWEN_MAX_MODEL_LEN=131072
# Sampler-warmup OOM guard on the shared GPU (248K vocab × default 1024 seqs is
# a huge transient). 32 is plenty for a vision endpoint.
QWEN_MAX_NUM_SEQS=32
# Optional — checkpoint is ungated.
HF_TOKEN=
API_KEY=