feat(mistral-small-4): pin v0.22.0 for working VISION baseline + reasoning entry

Operator needs a verified-working vision tower as the abliteration/tuning
baseline. vLLM 0.23.0 crashes Mistral multimodal at startup (#44911
fetch_images regression, ~0.22.1+). Pinned the Mistral container to
v0.22.0 — the last pre-regression release — which loads the NVFP4
(compressed-tensors) AND serves vision: verified a half-blue/half-red
image read correctly ('left blue, right red'). Dropped --limit-mm
(vision re-enabled). qwen36 stays on 0.23.0 (separate container; needs it
for its ModelOpt NVFP4).

- gateway: add mistral-small-4-reasoning. Operator asked for effort=medium
  but Mistral's reasoning_effort is BINARY (none/high only — medium 400s);
  set to 'high' (sole reasoning-ON level). NOTE: reasoning fires but
  reasoning_content-splitting is unreliable on v0.22.0 (lands in content);
  clean split would need 0.23.0, which breaks vision — vision prioritized.
- mistral-small-4 (instant) + mistral-small-4-reasoning both gateway-live.
This commit is contained in:
vh
2026-06-15 17:56:37 -07:00
parent c77a9aa4d8
commit 9a49963d07
3 changed files with 32 additions and 14 deletions
+15
View File
@@ -80,6 +80,21 @@ model_list:
model_info:
mode: chat
# mistral-small-4-reasoning: same upstream checkpoint, reasoning ON (2026-06-15,
# operator wanted "medium" — but Mistral's reasoning_effort is BINARY, only
# 'none' or 'high' (a medium/low request 400s). 'high' is the sole reasoning-ON
# level, so it carries the -reasoning intent. LiteLLM forwards extra_body to
# vLLM; the mistral reasoning-parser splits [THINK]…[/THINK] into reasoning_content.
- model_name: mistral-small-4-reasoning
litellm_params:
model: hosted_vllm/mistral-small-4
api_base: http://10.250.50.54:8010/v1
api_key: os.environ/VLLM_API_KEY
extra_body:
reasoning_effort: high
model_info:
mode: chat
# --- Qwen3 embeddings ---
- model_name: qwen3-embedding
litellm_params: