2e3dcc2d3d
Qwen3.5-9B vision-language served FP8 on ana-ml2 GPU1 (co-located with granite + the embed/rerank/reward trio; GPU0 kept free for hot-loading large models), :8007, fronted by LiteLLM as qwen3.5-9b-fp8. Pinned to vllm/vllm-openai nightly@sha256:49211ab2 — :latest (v0.19.1) quantizes the VL vision tower under fp8 and garbles vision; the nightly correctly excludes it (LM stays FP8, vision tower BF16). util 0.40 (~38GB) on the shared card (vLLM needs free>=util*total here). Vision verified end-to-end through the gateway.