chore(vllm): retire LFM2.5-2.6B permanently; audit finds nevermore on the harmful reranker

Operator directive: lfm2.5-2.6b goes down permanently.

  - stacks/vllm/compose.yaml   vllm-lfm25 service removed (replaced by a
                               tombstone comment), pushed live to ana-ml2
  - ana-ml2                    container docker rm -f'd, 8,721 MiB freed on GPU1
                               (95,388 -> 86,667 of 97,887)
  - litellm config             lfm2.5-2.6b alias deleted, live + canonical,
                               28 -> 27 models

It was an EVAL-ONLY bake-off seat against granite-4.1-8b that never received
the operator ruling it was pending; the comparator was retired from the roster
on 2026-08-15; it was deliberately never wired into any default or fallback
routing chain; and spend logs show 0 calls in the 4-day window to 2026-08-21.
Weights stay in the shared HF cache -- nothing deleted from disk.

The gateway restart that makes the alias deletion take effect is HELD so it can
batch with a pending reranker change. Until then the name is still routable
in-memory and will error against a dead backend.

Auditing the three reranker seats while answering "why do we have three" turned
up a real problem. The design is one production, one rollback, one fallback --
but the traffic is backwards:

  :8013 A3 bge-v2-m3      PRODUCTION, backs `reranker`     0 calls / 4 days
  :8002 Qwen3-Reranker    RETIRED incumbent, rollback only 7 calls, 12-hourly
  :8014 A4 gte-modernbert "fallback"                       no alias at all

nevermore is hard-wired to the incumbent by name (NEVERMORE_RERANK_MODEL=
qwen3-reranker), so the R43 cutover never moved it -- the cutover repointed the
`reranker` alias and correctly left `qwen3-reranker` naming the Qwen model.
Brokkr R43 measured that model harming 80/90 fleet queries, so nevermore's
twice-daily rerank pass is likely degrading its own briefing.

Fix is one line in nevermore's .env plus a nevermore restart, and it must land
before :8002 is retired. Recorded in persistent-memory with the A4 alias also
noted as absent (global CLAUDE.md names reranker-a4-gte-modernbert; it does not
exist).
This commit is contained in:
vh
2026-08-20 23:23:22 -07:00
parent e3ce713f7f
commit b990951d80
3 changed files with 26 additions and 90 deletions
+6 -16
View File
@@ -513,22 +513,12 @@ model_list:
# model names. Removed so unknown models now fail loudly (404). Re-add an
# explicit per-model entry if a swappable zoo ever returns. ---
# --- lfm2.5-2.6b -> LiquidAI LFM2.5-2.6B (ana-ml2 GPU1 :8021, vLLM). NON-PROD bake-off
# vs granite-4.1-8b (brokkr R-target 2026-08-10). LFM Open License v1.0 (<USD 10M-rev
# commercial) - EVAL-ONLY pending operator ruling; NOT in any default/fallback chain.
# Reasoning model served raw (no vLLM reasoning-parser) so content is non-empty.
# Vendor sampling (temp 0.1 / top_k 50 / rep_pen 1.1) baked as the alias default. ---
- model_name: lfm2.5-2.6b
litellm_params:
model: hosted_vllm/lfm2.5-2.6b
api_base: http://10.250.50.54:8021/v1
api_key: os.environ/VLLM_API_KEY
temperature: 0.1
extra_body:
top_k: 50
repetition_penalty: 1.1
model_info:
mode: chat
# --- lfm2.5-2.6b -> RETIRED PERMANENTLY 2026-08-20 (operator directive). The
# LiquidAI LFM2.5-2.6B seat (ana-ml2 GPU1 :8021) was an EVAL-ONLY bake-off
# against granite-4.1-8b that never got its operator ruling; its comparator
# was retired 2026-08-15 and spend logs showed 0 calls in the 4 days to
# 2026-08-21. Container removed, service deleted from stacks/vllm. The alias
# is deleted rather than repointed so the name 404s cleanly. ---
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY