# selene — RETIRED 2026-08-23 AtlaAI Selene 1 Mini (Llama 3.1 8B, dynamic FP8), served on ana-ml2 GPU 1 at `:8011` as `selene-1-mini-8b`. An LLM-as-judge: it scored and critiqued other models' output rather than generating for users (hallucination / RAG-faithfulness checks, Worldtree's Domari and selene-judgment roles). **The seat is down and the compose file is kept for reference only.** It is not deployed. `stacks-mirror/` will no longer show it on the host. ## Why it was retired Benchmarked head-to-head against `gen` (qwen3.8-27b-uncensored, ana-ml2 GPU 0 `:8015`) **on selene's own job** — 24 designed judge items with checkable ground truth, pairwise and absolute-scoring modes, 3 repeats each, run on two prompt templates. 288 calls total, all free local. | template | selene | gen | |---|---|---| | neutral JSON prompt | 20/24 (83%) | **23/24 (96%)** | | Selene's native Atla template | 21/24 (88%) | **22/24 (92%)** | **gen won on both templates, and selene's best score sat below gen's worst.** Selene was given its own fine-tuned prompt format as a fairness check — it gained one point, not the three it needed. **The decisive defect: selene cannot emit "tie."** On both tie items, on both templates, it forced a winner (0/2 each time). gen returned `tie` correctly on the JSON template. For eval work, close pairs are precisely the case that matters; a judge that manufactures a preference on every equivalent pair is producing noise exactly where it is most trusted. Other findings: - **Calibration.** gen used the full 1–5 range decisively (1s for bad answers, 5s for good). Selene clustered at 2s and 4s, compressing the scale. - **Stability at temperature 0.** Selene had 2 unstable items on the JSON template; gen had 0. They swapped on the native template (gen 4, selene 0) — longer free-text reasoning costs determinism. - **Latency was selene's only win** — roughly 3× faster (0.23s vs 0.40s median on JSON). Unexercised: it served ~60 calls/day with zero queueing across five weeks of uptime. - **Shared blind spot.** Both preferred a response containing an arithmetic error on the JSON template. gen caught it on the reasoning-first template. Neither is trustworthy for numerically-checkable judgments without a prompt that forces reasoning before the verdict. Best measured configuration overall: **`gen` + the neutral JSON prompt** — 23/24, zero instability, ties handled correctly. ## What happened to the names **`chat-judge` → repointed to `gen`.** It is a *role* alias, and ADR-0012 is explicit that consumers bind the capability, not a concrete model. Sampler profile copied from `image-judge` (temperature 0, top_p 1.0, top_k 1, thinking off) so the served config matches the benchmarked condition. **`selene-1-mini-8b` → removed outright. It 404s by design.** It was NOT aliased to `gen`. A served-name is a contract about what the model *is*; answering it with a different model hides a material change behind a stable string, and the caller has no way to know. Failing loud forces a conscious migration. Operator ruling, 2026-08-23: > never repoint a named model at a different model's endpoint — that is > intentionally misleading Direct precedent on this gateway: `qwen-image-bench` / `image-judge` had their dedicated backend retired 2026-07-15 to reclaim GPU 1 VRAM and were repointed to `gen` as role aliases. Counter-example worth remembering: when `qwen3-reranker` was retired, `nevermore` was pinned to it *by name* and the cutover moved the `reranker` alias but never moved nevermore — which is why name-pinned consumers get notified explicitly rather than assumed covered. ## Reclaimed ``` GPU 1 free: 1,818 MiB → 19,450 MiB ``` 17.2 GiB, on a card that had under 2 GiB of headroom. `gen` was already running on GPU 0, so the judge role moved onto an existing seat rather than allocating anything new. ## To bring it back `compose.yaml` and the host `.env` are intact. GPU 1 must have ~17 GiB free (`SELENE_GPU_MEM_UTIL=0.17`); re-add the `selene-1-mini-8b` entry to the LiteLLM config **under its own true name**, never as an alias for something else.