feat(selene+mistral): restore Selene judge (FP8, GPU1) + push Mistral to 256K
selene: AtlaAI Selene-1-Mini-Llama-3.1-8B judge restored on vLLM after the llama-swap teardown took its Q6_K GGUF offline. FP8 (dynamic --quantization fp8; FP8 >= the validated Q6_K fidelity, and text-only Llama so no vision- tower-noise risk; NVFP4's W4A4 too aggressive for a precision judge). GPU1 util 0.13 (8.51 GiB weights + 2.53 GiB KV, 32K ctx, 1.27x concurrency), ~11 GB GPU1 buffer left. Gateway selene-1-mini-8b → :8011 (shadows the * wildcard that used to reach it via llama-swap). Judge smoke: scored an unfaithful claim 1/5 correctly. mistral-small-4: max-model-len 131072 → 262144 (full native 256K) for novel-length consistency-checking. KV pool is util-bound (~862K tokens), so 256K costs no extra VRAM — max concurrency just drops to 3.29x at full length. max-num-seqs 64 → 32 keeps the warmup transient flat (scales with seqs × len), so it fits the tight GPU0 (free unchanged at 5.2 GB). Verified loaded + healthy.
This commit is contained in:
@@ -19,16 +19,17 @@
|
||||
# reference 80 GB cards can't fit 74.4 GB + context on one). The 96 GB Blackwell
|
||||
# flips that to single-card: 74.4 GB weights + ~5 GB overhead leaves ~17 GB for
|
||||
# KV. Mistral Small 4 uses MLA attention (TRITON_MLA) so KV is compressed/cheap —
|
||||
# big context stays affordable even on a constrained KV pool. max-model-len is
|
||||
# capped to 131072 on first bring-up (raise toward the native 256K once real KV
|
||||
# headroom is measured).
|
||||
# big context stays affordable even on a constrained KV pool. We serve the FULL
|
||||
# native 256K (max-model-len 262144) — the KV pool is util-bound (~862K tokens)
|
||||
# so 256K costs no extra VRAM, it just lets one request use up to 256K (max
|
||||
# concurrency 3.29x at full length). For novel-length consistency-checking.
|
||||
#
|
||||
# vLLM FLOOR: needs >= 0.20 (Mistral Small 4 day-0 support); validated on 0.23.0.
|
||||
# Do NOT reuse the qwen36-vl 0.19.1 image — it predates this model.
|
||||
# vLLM PIN: v0.22.0 (in .env) — the last release with WORKING Mistral vision
|
||||
# (#44911 fetch_images regression hit 0.22.1+/0.23.0). See the .env header.
|
||||
#
|
||||
# Serve flags mirror Mistral's official command (cited in README), adapted for
|
||||
# single-card: TP 2->1, util 0.8->0.93, max-len 262144->131072, max-num-seqs
|
||||
# 128->64. All tunables live in .env — edit that, not this file.
|
||||
# single-card: TP 2->1, util 0.8->0.93, max-num-seqs 128->32 (32 keeps the
|
||||
# 256K warmup transient flat). All tunables live in .env — edit that, not this file.
|
||||
|
||||
name: mistral-small-4
|
||||
|
||||
|
||||
Reference in New Issue
Block a user