feat(mistral-small-4): pin v0.22.0 for working VISION baseline + reasoning entry

Operator needs a verified-working vision tower as the abliteration/tuning
baseline. vLLM 0.23.0 crashes Mistral multimodal at startup (#44911
fetch_images regression, ~0.22.1+). Pinned the Mistral container to
v0.22.0 — the last pre-regression release — which loads the NVFP4
(compressed-tensors) AND serves vision: verified a half-blue/half-red
image read correctly ('left blue, right red'). Dropped --limit-mm
(vision re-enabled). qwen36 stays on 0.23.0 (separate container; needs it
for its ModelOpt NVFP4).

- gateway: add mistral-small-4-reasoning. Operator asked for effort=medium
  but Mistral's reasoning_effort is BINARY (none/high only — medium 400s);
  set to 'high' (sole reasoning-ON level). NOTE: reasoning fires but
  reasoning_content-splitting is unreliable on v0.22.0 (lands in content);
  clean split would need 0.23.0, which breaks vision — vision prioritized.
- mistral-small-4 (instant) + mistral-small-4-reasoning both gateway-live.
This commit is contained in:
vh
2026-06-15 17:56:37 -07:00
parent c77a9aa4d8
commit 9a49963d07
3 changed files with 32 additions and 14 deletions
+9 -5
View File
@@ -3,11 +3,15 @@
#
# See compose.yaml header for the NVFP4/TP=1/MLA rationale and the vLLM>=0.20 floor.
# PINNED by digest, not :latest — this model is version-sensitive (needs vLLM
# >= 0.20 for day-0 support; the related nvidia-ModelOpt NVFP4 MoE path broke on
# 0.19.1/0.22.0). Pin protects against a :latest regression. This digest = vLLM
# 0.23.0, the version validated to load this checkpoint. Bump deliberately.
MISTRAL_IMAGE=vllm/vllm-openai@sha256:6d8429e38e3747723ca07ee1b17972e09bb9c51c4032b266f24fb1cc3b22ed8f
# PINNED to v0.22.0 — the last release BEFORE the Mistral multimodal regression
# (#44911, MistralCommonImageProcessor.fetch_images, landed ~0.22.1; 0.23.0 is
# affected). v0.22.0 loads the NVFP4 (compressed-tensors) AND serves VISION —
# verified: half-blue/half-red image read correctly ("left blue, right red").
# This gives a working vision tower as the abliteration/tuning baseline. Do NOT
# bump to 0.23.0 (breaks vision). reasoning_effort works (none/high only) but
# reasoning_content-splitting is unreliable on this version — vision is the
# priority. Revisit when vLLM patches the Mistral mm path on a newer release.
MISTRAL_IMAGE=vllm/vllm-openai:v0.22.0
MISTRAL_CONTAINER_NAME=vllm-mistral4
MISTRAL_MODEL=mistralai/Mistral-Small-4-119B-2603-NVFP4