Files
esh-pfi-infrastructure/stacks/mistral-small-4/compose.yaml
T
vh c985ede07b feat(selene+mistral): restore Selene judge (FP8, GPU1) + push Mistral to 256K
selene: AtlaAI Selene-1-Mini-Llama-3.1-8B judge restored on vLLM after the
llama-swap teardown took its Q6_K GGUF offline. FP8 (dynamic --quantization
fp8; FP8 >= the validated Q6_K fidelity, and text-only Llama so no vision-
tower-noise risk; NVFP4's W4A4 too aggressive for a precision judge). GPU1
util 0.13 (8.51 GiB weights + 2.53 GiB KV, 32K ctx, 1.27x concurrency),
~11 GB GPU1 buffer left. Gateway selene-1-mini-8b → :8011 (shadows the *
wildcard that used to reach it via llama-swap). Judge smoke: scored an
unfaithful claim 1/5 correctly.

mistral-small-4: max-model-len 131072 → 262144 (full native 256K) for
novel-length consistency-checking. KV pool is util-bound (~862K tokens), so
256K costs no extra VRAM — max concurrency just drops to 3.29x at full length.
max-num-seqs 64 → 32 keeps the warmup transient flat (scales with seqs × len),
so it fits the tight GPU0 (free unchanged at 5.2 GB). Verified loaded + healthy.
2026-06-15 18:52:26 -07:00

117 lines
4.7 KiB
YAML

# mistral-small-4 — Mistral-Small-4-119B-2603 (official NVFP4) on ana-ml2 GPU 0.
#
# Mistral Small 4 is a 119B-total / 6.5B-active MoE (128 experts, 4 active),
# 256K context, multimodal, Apache-2.0 (released 2026-03). This serves the
# OFFICIAL NVFP4 checkpoint (mistralai/Mistral-Small-4-119B-2603-NVFP4) — 74.4 GB
# of compressed-tensors (llm-compressor, a vLLM + Red Hat collaboration, day-0
# vLLM support). It is the GPU-0 tenant (the slot formerly reserved for a
# creative-writing pick — operator reassigned 2026-06-15; tune-for-creative-
# writing comes after base-characteristic probing).
#
# WHY NVFP4 (not FP8/bf16): on a SINGLE 96 GB card, NVFP4 (74.4 GB weights) is
# the only variant that fits at TP=1 — FP8 (~119 GB) and bf16 (~238 GB) need both
# GPUs. The card is Blackwell (sm_120) with FP4 tensor cores, so NVFP4 gets a real
# speedup, not just a VRAM save. NOTE: this is the COMPRESSED-TENSORS NVFP4 path
# (vendor-shipped, vLLM-tested) — distinct from the nvidia-ModelOpt NVFP4 MoE
# loader that broke on Qwen3.6 (#44081); different code path, day-0 supported.
#
# WHY TP=1 here: Mistral's official card uses --tensor-parallel-size 2 (their
# reference 80 GB cards can't fit 74.4 GB + context on one). The 96 GB Blackwell
# flips that to single-card: 74.4 GB weights + ~5 GB overhead leaves ~17 GB for
# KV. Mistral Small 4 uses MLA attention (TRITON_MLA) so KV is compressed/cheap —
# big context stays affordable even on a constrained KV pool. We serve the FULL
# native 256K (max-model-len 262144) — the KV pool is util-bound (~862K tokens)
# so 256K costs no extra VRAM, it just lets one request use up to 256K (max
# concurrency 3.29x at full length). For novel-length consistency-checking.
#
# vLLM PIN: v0.22.0 (in .env) — the last release with WORKING Mistral vision
# (#44911 fetch_images regression hit 0.22.1+/0.23.0). See the .env header.
#
# Serve flags mirror Mistral's official command (cited in README), adapted for
# single-card: TP 2->1, util 0.8->0.93, max-num-seqs 128->32 (32 keeps the
# 256K warmup transient flat). All tunables live in .env — edit that, not this file.
name: mistral-small-4
services:
vllm-mistral4:
image: ${MISTRAL_IMAGE}
container_name: ${MISTRAL_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${MISTRAL_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${MISTRAL_MODEL}
# Pre-quantized NVFP4 (compressed-tensors) — vLLM auto-detects the quant;
# no --quantization flag.
- --served-model-name
- mistral-small-4
- --host
- 0.0.0.0
- --port
- "8000"
- --tensor-parallel-size
- "1"
- --gpu-memory-utilization
- ${MISTRAL_GPU_MEM_UTIL}
- --max-model-len
- ${MISTRAL_MAX_MODEL_LEN}
# MLA attention backend (DeepSeek-style latent KV → compressed, cheap KV).
- --attention-backend
- TRITON_MLA
# Mistral tool-calling + configurable reasoning (per the official card).
- --tool-call-parser
- mistral
- --enable-auto-tool-choice
- --reasoning-parser
- mistral
- --max-num-seqs
- ${MISTRAL_MAX_NUM_SEQS}
# VISION ENABLED. vLLM is pinned to v0.22.0 in .env — the last release BEFORE
# the Mistral multimodal regression (#44911, `MistralCommonImageProcessor has
# no attribute fetch_images`, landed ~0.22.1+; 0.23.0 is affected). v0.22.0
# still has Mistral-Small-4 arch + compressed-tensors NVFP4 support (the
# #44081 ModelOpt-NVFP4 bug on 0.22.0 is a DIFFERENT quant path, doesn't touch
# this compressed-tensors checkpoint). Gives a verified working vision tower
# as the abliteration/tuning baseline. (qwen36 stays on 0.23.0 — separate
# container; it NEEDS 0.23.0 for its ModelOpt NVFP4.)
- --dtype
- auto
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${MISTRAL_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 600s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Mistral Small 4 (NVFP4)
- homepage.icon=mdi-creation
- homepage.description=Mistral-Small-4-119B-2603 MoE (NVFP4) via vLLM (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs
networks:
tnet:
name: traefik-net
external: true