feat(mistral-small-4): pin v0.22.0 for working VISION baseline + reasoning entry

Operator needs a verified-working vision tower as the abliteration/tuning
baseline. vLLM 0.23.0 crashes Mistral multimodal at startup (#44911
fetch_images regression, ~0.22.1+). Pinned the Mistral container to
v0.22.0 — the last pre-regression release — which loads the NVFP4
(compressed-tensors) AND serves vision: verified a half-blue/half-red
image read correctly ('left blue, right red'). Dropped --limit-mm
(vision re-enabled). qwen36 stays on 0.23.0 (separate container; needs it
for its ModelOpt NVFP4).

- gateway: add mistral-small-4-reasoning. Operator asked for effort=medium
  but Mistral's reasoning_effort is BINARY (none/high only — medium 400s);
  set to 'high' (sole reasoning-ON level). NOTE: reasoning fires but
  reasoning_content-splitting is unreliable on v0.22.0 (lands in content);
  clean split would need 0.23.0, which breaks vision — vision prioritized.
- mistral-small-4 (instant) + mistral-small-4-reasoning both gateway-live.
This commit is contained in:
vh
2026-06-15 17:56:37 -07:00
parent c77a9aa4d8
commit 9a49963d07
3 changed files with 32 additions and 14 deletions
+8 -9
View File
@@ -74,15 +74,14 @@ services:
- mistral
- --max-num-seqs
- ${MISTRAL_MAX_NUM_SEQS}
# TEXT-ONLY (2026-06-15): vLLM 0.23.0's Mistral multimodal processor crashes
# at startup dummy-image profiling — `MistralCommonImageProcessor has no
# attribute fetch_images` (vLLM↔mistral_common incompat; same class hit
# Mistral-3.1/Devstral/Magistral). Setting image/video limit to 0 skips the
# vision profiling so the model loads text-only — which is all the creative-
# writing use needs. REMOVE this flag to restore vision once vLLM patches the
# Mistral mm path (track: the model is natively multimodal).
- --limit-mm-per-prompt
- '{"image":0,"video":0}'
# VISION ENABLED. vLLM is pinned to v0.22.0 in .env — the last release BEFORE
# the Mistral multimodal regression (#44911, `MistralCommonImageProcessor has
# no attribute fetch_images`, landed ~0.22.1+; 0.23.0 is affected). v0.22.0
# still has Mistral-Small-4 arch + compressed-tensors NVFP4 support (the
# #44081 ModelOpt-NVFP4 bug on 0.22.0 is a DIFFERENT quant path, doesn't touch
# this compressed-tensors checkpoint). Gives a verified working vision tower
# as the abliteration/tuning baseline. (qwen36 stays on 0.23.0 — separate
# container; it NEEDS 0.23.0 for its ModelOpt NVFP4.)
- --dtype
- auto
- --enable-prefix-caching