c985ede07b
selene: AtlaAI Selene-1-Mini-Llama-3.1-8B judge restored on vLLM after the llama-swap teardown took its Q6_K GGUF offline. FP8 (dynamic --quantization fp8; FP8 >= the validated Q6_K fidelity, and text-only Llama so no vision- tower-noise risk; NVFP4's W4A4 too aggressive for a precision judge). GPU1 util 0.13 (8.51 GiB weights + 2.53 GiB KV, 32K ctx, 1.27x concurrency), ~11 GB GPU1 buffer left. Gateway selene-1-mini-8b → :8011 (shadows the * wildcard that used to reach it via llama-swap). Judge smoke: scored an unfaithful claim 1/5 correctly. mistral-small-4: max-model-len 131072 → 262144 (full native 256K) for novel-length consistency-checking. KV pool is util-bound (~862K tokens), so 256K costs no extra VRAM — max concurrency just drops to 3.29x at full length. max-num-seqs 64 → 32 keeps the warmup transient flat (scales with seqs × len), so it fits the tight GPU0 (free unchanged at 5.2 GB). Verified loaded + healthy.
27 lines
1.0 KiB
Bash
27 lines
1.0 KiB
Bash
# Selene 1 Mini 8B judge (FP8) on ana-ml2 GPU 1 — copy to .env on the host.
|
|
# Real .env lives on ana-ml2 at /opt/docker/compose/selene/.env (gitignored).
|
|
#
|
|
# See compose.yaml header for the FP8-over-NVFP4 (judge-fidelity) rationale.
|
|
|
|
# 0.23.0 — Llama 3.1 + dynamic fp8 is rock-solid here (same digest as qwen36).
|
|
SELENE_IMAGE=vllm/vllm-openai@sha256:6d8429e38e3747723ca07ee1b17972e09bb9c51c4032b266f24fb1cc3b22ed8f
|
|
|
|
SELENE_CONTAINER_NAME=vllm-selene
|
|
SELENE_MODEL=AtlaAI/Selene-1-Mini-Llama-3.1-8B
|
|
SELENE_PORT=8011
|
|
|
|
# GPU 1 = shared with qwen36 (NVFP4) + granite + embed/rerank/reward.
|
|
SELENE_GPU_ID=1
|
|
|
|
# util 0.13 (~12.5 GB) — measured: 8.51 GiB FP8 weights + ~1.5 GiB graph +
|
|
# 2.53 GiB fp8 KV (41,456-token pool, 1.27x concurrency at full 32K). util 0.12
|
|
# was too thin (1.86 GiB KV < the 2.0 GiB a single 32K request needs → crash).
|
|
# Keeps GPU 1 total ~0.95 → ~5 GB buffer; ctx 32768 mirrors the old judge config.
|
|
SELENE_GPU_MEM_UTIL=0.13
|
|
SELENE_MAX_MODEL_LEN=32768
|
|
SELENE_MAX_NUM_SEQS=16
|
|
|
|
# Optional — checkpoint is ungated (Apache-2.0).
|
|
HF_TOKEN=
|
|
API_KEY=
|