Files
esh-pfi-infrastructure/stacks/selene/README.md
T
vh ca3c984f93 feat(selene): retire the seat; chat-judge -> gen, selene-1-mini-8b 404s by design
Benchmarked selene against gen on selene's own job: 24 designed judge items
with checkable ground truth, pairwise + absolute modes, 3 repeats, on BOTH a
neutral JSON prompt and Selene's native Atla template. 288 calls, all free
local.

  neutral JSON     selene 20/24 (83%)   gen 23/24 (96%)
  native Atla      selene 21/24 (88%)   gen 22/24 (92%)

gen won on both templates and selene's BEST sat below gen's WORST. Selene was
given its own fine-tuned template as a fairness check; it gained one point,
not the three it needed.

Decisive defect: selene cannot emit "tie" -- 0/2 on both templates, forcing a
winner on every equivalent pair. For eval work that is the case that matters
most. gen returned tie correctly on the JSON template. Selene also compressed
the 1-5 scale (clustered at 2s and 4s) where gen used it fully. Selene's only
win was ~3x latency, unexercised at ~60 calls/day with zero queueing.

TWO NAMES, TWO DIFFERENT TREATMENTS, deliberately:

- chat-judge -> repointed to gen. It is a ROLE alias and ADR-0012 says
  consumers bind the capability, not a concrete model. Sampler profile copied
  from image-judge (temp 0, top_p 1.0, top_k 1, thinking off) so the served
  config matches the benchmarked condition.

- selene-1-mini-8b -> REMOVED. It 404s. It was NOT aliased to gen. A
  served-name is a contract about what the model IS; answering it with a
  different model hides a material change behind a stable string. Operator
  ruling: "never repoint a named model at a different model's endpoint --
  that is intentionally misleading." Verified: the gateway now returns
  HTTP 400 "Invalid model name" for it.

Reclaimed 17.2 GiB on ana-ml2 GPU 1 (free 1,818 -> 19,450 MiB) on a card that
had under 2 GiB of headroom. gen already runs on GPU 0, so the judge role
moved onto an existing seat rather than allocating anything new.

Canonical litellm config synced from the host; ana-ml2 README and
recommended-model-settings updated. compose.yaml kept for reference, not
deployed.
2026-08-23 05:07:01 -07:00

90 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# selene — RETIRED 2026-08-23
AtlaAI Selene 1 Mini (Llama 3.1 8B, dynamic FP8), served on ana-ml2 GPU 1 at
`:8011` as `selene-1-mini-8b`. An LLM-as-judge: it scored and critiqued other
models' output rather than generating for users (hallucination /
RAG-faithfulness checks, Worldtree's Domari and selene-judgment roles).
**The seat is down and the compose file is kept for reference only.** It is not
deployed. `stacks-mirror/` will no longer show it on the host.
## Why it was retired
Benchmarked head-to-head against `gen` (qwen3.8-27b-uncensored, ana-ml2 GPU 0
`:8015`) **on selene's own job** — 24 designed judge items with checkable
ground truth, pairwise and absolute-scoring modes, 3 repeats each, run on two
prompt templates. 288 calls total, all free local.
| template | selene | gen |
|---|---|---|
| neutral JSON prompt | 20/24 (83%) | **23/24 (96%)** |
| Selene's native Atla template | 21/24 (88%) | **22/24 (92%)** |
**gen won on both templates, and selene's best score sat below gen's worst.**
Selene was given its own fine-tuned prompt format as a fairness check — it
gained one point, not the three it needed.
**The decisive defect: selene cannot emit "tie."** On both tie items, on both
templates, it forced a winner (0/2 each time). gen returned `tie` correctly on
the JSON template. For eval work, close pairs are precisely the case that
matters; a judge that manufactures a preference on every equivalent pair is
producing noise exactly where it is most trusted.
Other findings:
- **Calibration.** gen used the full 15 range decisively (1s for bad answers,
5s for good). Selene clustered at 2s and 4s, compressing the scale.
- **Stability at temperature 0.** Selene had 2 unstable items on the JSON
template; gen had 0. They swapped on the native template (gen 4, selene 0) —
longer free-text reasoning costs determinism.
- **Latency was selene's only win** — roughly 3× faster (0.23s vs 0.40s median
on JSON). Unexercised: it served ~60 calls/day with zero queueing across five
weeks of uptime.
- **Shared blind spot.** Both preferred a response containing an arithmetic
error on the JSON template. gen caught it on the reasoning-first template.
Neither is trustworthy for numerically-checkable judgments without a prompt
that forces reasoning before the verdict.
Best measured configuration overall: **`gen` + the neutral JSON prompt** —
23/24, zero instability, ties handled correctly.
## What happened to the names
**`chat-judge` → repointed to `gen`.** It is a *role* alias, and ADR-0012 is
explicit that consumers bind the capability, not a concrete model. Sampler
profile copied from `image-judge` (temperature 0, top_p 1.0, top_k 1,
thinking off) so the served config matches the benchmarked condition.
**`selene-1-mini-8b` → removed outright. It 404s by design.** It was NOT
aliased to `gen`. A served-name is a contract about what the model *is*;
answering it with a different model hides a material change behind a stable
string, and the caller has no way to know. Failing loud forces a conscious
migration. Operator ruling, 2026-08-23:
> never repoint a named model at a different model's endpoint — that is
> intentionally misleading
Direct precedent on this gateway: `qwen-image-bench` / `image-judge` had their
dedicated backend retired 2026-07-15 to reclaim GPU 1 VRAM and were repointed
to `gen` as role aliases. Counter-example worth remembering: when
`qwen3-reranker` was retired, `nevermore` was pinned to it *by name* and the
cutover moved the `reranker` alias but never moved nevermore — which is why
name-pinned consumers get notified explicitly rather than assumed covered.
## Reclaimed
```
GPU 1 free: 1,818 MiB → 19,450 MiB
```
17.2 GiB, on a card that had under 2 GiB of headroom. `gen` was already
running on GPU 0, so the judge role moved onto an existing seat rather than
allocating anything new.
## To bring it back
`compose.yaml` and the host `.env` are intact. GPU 1 must have ~17 GiB free
(`SELENE_GPU_MEM_UTIL=0.17`); re-add the `selene-1-mini-8b` entry to the
LiteLLM config **under its own true name**, never as an alias for something
else.