Files
esh-pfi-infrastructure/stacks/qwen36-vl
vh c6d76051a4 feat(qwen36-vl): swap FP8→NVFP4 + GPU1 rebalance (granite restored)
The nvidia ModelOpt NVFP4 MoE that failed on vLLM 0.19.1/0.22.0 (#44081)
loads clean on 0.23.0. Cut prod qwen36 FP8→NVFP4: ~20.4 GiB weights vs
~34 (~40% lighter, ~13 GB reclaimed on GPU 1), faster single-stream on
Blackwell FP4 cores, vision tower preserved (comfy-dev real anatomy-judge
A/B on 16 prod images: PASS; brokkr text/speed: parity bar a minor
multi-step-chained-reasoning slip that doesn't bite the judge role).

- compose: pin image by 0.23.0 digest, drop --kv-cache-dtype fp8 (fp16 KV
  — the freed room buys full-precision KV), util 0.46→0.32.
- GPU1 rebalance (pinned): granite restored 0.24→0.34 / 65536→131072
  (undoes the FP8-era sacrifice); trio unchanged; total ~0.82, ~24 GB free.
- gateway model name qwen3.6-35b-a3b unchanged (now NVFP4 behind it);
  thinking-split (enable_thinking=false default) intact — the judge needs it.
2026-06-15 17:28:15 -07:00
..