# mistral-small-4 — Mistral-Small-4-119B-2603 (official NVFP4) on ana-ml2 GPU 0. # # Mistral Small 4 is a 119B-total / 6.5B-active MoE (128 experts, 4 active), # 256K context, multimodal, Apache-2.0 (released 2026-03). This serves the # OFFICIAL NVFP4 checkpoint (mistralai/Mistral-Small-4-119B-2603-NVFP4) — 74.4 GB # of compressed-tensors (llm-compressor, a vLLM + Red Hat collaboration, day-0 # vLLM support). It is the GPU-0 tenant (the slot formerly reserved for a # creative-writing pick — operator reassigned 2026-06-15; tune-for-creative- # writing comes after base-characteristic probing). # # WHY NVFP4 (not FP8/bf16): on a SINGLE 96 GB card, NVFP4 (74.4 GB weights) is # the only variant that fits at TP=1 — FP8 (~119 GB) and bf16 (~238 GB) need both # GPUs. The card is Blackwell (sm_120) with FP4 tensor cores, so NVFP4 gets a real # speedup, not just a VRAM save. NOTE: this is the COMPRESSED-TENSORS NVFP4 path # (vendor-shipped, vLLM-tested) — distinct from the nvidia-ModelOpt NVFP4 MoE # loader that broke on Qwen3.6 (#44081); different code path, day-0 supported. # # WHY TP=1 here: Mistral's official card uses --tensor-parallel-size 2 (their # reference 80 GB cards can't fit 74.4 GB + context on one). The 96 GB Blackwell # flips that to single-card: 74.4 GB weights + ~5 GB overhead leaves ~17 GB for # KV. Mistral Small 4 uses MLA attention (TRITON_MLA) so KV is compressed/cheap — # big context stays affordable even on a constrained KV pool. We serve the FULL # native 256K (max-model-len 262144) — the KV pool is util-bound (~862K tokens) # so 256K costs no extra VRAM, it just lets one request use up to 256K (max # concurrency 3.29x at full length). For novel-length consistency-checking. # # vLLM PIN: v0.22.0 (in .env) — the last release with WORKING Mistral vision # (#44911 fetch_images regression hit 0.22.1+/0.23.0). See the .env header. # # Serve flags mirror Mistral's official command (cited in README), adapted for # single-card: TP 2->1, util 0.8->0.93, max-num-seqs 128->32 (32 keeps the # 256K warmup transient flat). All tunables live in .env — edit that, not this file. name: mistral-small-4 services: vllm-mistral4: image: ${MISTRAL_IMAGE} container_name: ${MISTRAL_CONTAINER_NAME} restart: unless-stopped ipc: host ports: - "${MISTRAL_PORT}:8000" volumes: - /tank/aimodels/huggingface:/hfcache environment: - HF_HOME=/hfcache - HF_HUB_CACHE=/hfcache/hub - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-} - VLLM_API_KEY=${API_KEY:-} command: - ${MISTRAL_MODEL} # Pre-quantized NVFP4 (compressed-tensors) — vLLM auto-detects the quant; # no --quantization flag. - --served-model-name - mistral-small-4 - --host - 0.0.0.0 - --port - "8000" - --tensor-parallel-size - "1" - --gpu-memory-utilization - ${MISTRAL_GPU_MEM_UTIL} - --max-model-len - ${MISTRAL_MAX_MODEL_LEN} # MLA attention backend (DeepSeek-style latent KV → compressed, cheap KV). - --attention-backend - TRITON_MLA # Mistral tool-calling + configurable reasoning (per the official card). - --tool-call-parser - mistral - --enable-auto-tool-choice - --reasoning-parser - mistral - --max-num-seqs - ${MISTRAL_MAX_NUM_SEQS} # VISION ENABLED. vLLM is pinned to v0.22.0 in .env — the last release BEFORE # the Mistral multimodal regression (#44911, `MistralCommonImageProcessor has # no attribute fetch_images`, landed ~0.22.1+; 0.23.0 is affected). v0.22.0 # still has Mistral-Small-4 arch + compressed-tensors NVFP4 support (the # #44081 ModelOpt-NVFP4 bug on 0.22.0 is a DIFFERENT quant path, doesn't touch # this compressed-tensors checkpoint). Gives a verified working vision tower # as the abliteration/tuning baseline. (qwen36 stays on 0.23.0 — separate # container; it NEEDS 0.23.0 for its ModelOpt NVFP4.) - --dtype - auto - --enable-prefix-caching deploy: resources: reservations: devices: - driver: nvidia device_ids: - "${MISTRAL_GPU_ID}" capabilities: - gpu healthcheck: test: ["CMD", "curl", "-f", "http://localhost:8000/health"] interval: 30s timeout: 10s retries: 3 start_period: 600s networks: - tnet labels: - homepage.group=AI Systems - homepage.name=Mistral Small 4 (NVFP4) - homepage.icon=mdi-creation - homepage.description=Mistral-Small-4-119B-2603 MoE (NVFP4) via vLLM (ana-ml2 GPU 0) - homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs networks: tnet: name: traefik-net external: true