The incumbent Qwen3-Reranker-0.6B was measured actively harming 80/90 fleet queries on main+knowledge_base (and inverting the bare-name region behind Worldtree #389) — no-reranker beat it 89/90 vs 56/90. Brokkr's R43 bake-off selected BAAI/bge-reranker-v2-m3 (A3): 90/90 top-10, mean rank 0.19, multilingual (XLM-R), ~1.2 GB lighter than the incumbent. Control arm (A2 = same Qwen weights, seq-cls head) scored identical to the incumbent, proving the fault is a training prior, not the serving head — which cancelled the expensive Qwen3-4B arm before it cost a GPU seat. Cutover boundary 2026-08-06T17:37:48Z. The qwen3-reranker alias and the :8002 backend are retained for one-edit rollback. Adds the process audit trail at docs/pfi/reranker-selection-ledger.md.
13 KiB
Fleet reranker selection — process ledger
Running record of the Brokkr-driven fleet-reranker selection, and every assumption / autonomous decision infra-ops makes on the operator's behalf during it. The operator (Vuong) will review this at the end and reverse anything he wants. This is the audit trail for unattended operation.
Started: 2026-08-06. Driver: brokkr-smithy-dev. Executor: infra-ops (this session).
Operator authorization envelope (2026-08-06)
Brokkr drives a reranker-selection process; infra-ops is cleared to proceed on Brokkr's recommendations unattended (no per-step operator check-in), with authority to do whatever is necessary to reach a recommendation or implementation.
CLEARED (green):
- Execute Brokkr's reranker-selection recommendations unattended.
- Bring down the prod reranker at
ana-ml2:8002(qwen3-reranker-0.6B) — temporarily OR permanently. - Down ONE of the RP (roleplay) seats on ana-ml2 temporarily to free GPU/VRAM for testing.
- Temporarily clear space for the smoke/bench.
- Pull models, stand up side-port vLLM benches, run the harness — whatever the eval needs.
RED LINES (hard NO — stop + surface even under standing auth):
- NO permanent deletion of anything (no
rm/docker volume rm/model-weight deletion/data destruction). Downing ≠ deleting. - NO taking anything else offline beyond (a) the prod reranker and (b) ONE ana-ml2 RP seat. (Not granite/embed/reward/coder/gen/a second RP seat/muninn/etc.)
- NO rebooting machines.
Process: accumulate assumptions here; operator reverses at the end.
Standing assumptions / autonomous-decision log
- A1 — Coordinated-change notify still applies. Even under unattended auth,
every
:8002state change gets a timestamped announcement to worldtree-dev + brokkr-smithy-dev (their standing coordinated-change ask; the operator waived per-step operator approval, not the peer notify courtesy). No silent flip. - A2 — Weights are never deleted, only unserved. "Permanently down the qwen reranker" = stop serving + (optionally) repoint the gateway alias; the 0.6B model weights stay on disk (deletion is a red line).
- A3 — RP-seat pick = lowest-impact, temporary, restored after. When a seat must come down for VRAM, I pick the lowest-impact RP seat, log which + its exact restore command, and bring it back when the bench frees the GPU.
Current board at handoff
- Prod reranker:
ana-ml2:8002=vllm-rerank(Qwen/Qwen3-Reranker-0.6B), reverted to baselineclassifier_from_token:["no","yes"], healthy. Compose:/opt/docker/compose/vllm/compose.yaml(canonical mirrorstacks/vllm/compose.yaml). Gateway aliasreranker/qwen3-reranker→ litellm → :8002. - Root cause (converged, both sides): 0.6B is capacity-bound on bare-name queries over a real candidate pool; NOT misconfigured. Fix = larger model.
- Verified on-prem candidate shortlist (all HF-real, ungated): Qwen/Qwen3-Reranker-4B, Qwen/Qwen3-Reranker-8B, mixedbread-ai/mxbai-rerank-large-v2, mixedbread-ai/mxbai-rerank-base-v2, BAAI/bge-reranker-v2-gemma, Alibaba-NLP/gte-reranker-modernbert-base, jinaai/jina-reranker-v2-base-multilingual. (BAAI/bge-reranker-v2-m3 exists but the fleet already moved off it.)
- Eval assets (all on nh3-dev):
- Scorer:
scripts/probe_389_rank_decomposition.py(Worldtree repo, main) — rank-recovery =rrf_rerankcolumn climbing back towardrrf. - worldtree-dev grids:
~/snapshots/r42-gate-snapshot/(probe_389_run3.json, probe_389_question_shaped.json, probe_389_bigboi_control.json). - Frozen gate Chroma snapshot:
~/snapshots/r42-gate-index/(retained until worldtree-dev signals the lever run is done). - Dual query-set requirement (hard): score bare-name anchor queries AND question-shaped; bar = recovering the name-lookup class.
- Scorer:
- VRAM: 4B ≈ 4–5 GB fp8, 8B ≈ 9 GB; ana-ml2 Blackwell has headroom.
Progress log
2026-08-06 — A2 brought up (Brokkr thread 01KZBSTSJA…)
- Backend:
vllm-rerank-a2— standalonedocker run(NOT in the vllm compose stack), on ana-ml2 GPU1, host port :8012 → container 8000. Imagevllm/vllm-openai:latest(=0.24.0). Args: modeltomaarsen/Qwen3-Reranker-0.6B-seq-cls,--runner pooling,--gpu-memory-utilization 0.03,--max-model-len 8192,--dtype auto,--restart no. NativeQwen3ForSequenceClassification— NO hf-overrides. Routes /rerank /score /classify. - Gateway alias:
reranker-a2-qwen3-seqcls→http://10.250.50.54:8012/v1, mode rerank. Added via LiteLLM/model/new(DB-backed,store_model_in_db:true) — no gateway restart (respects the "nothing else offline" line). Verified 200 through the gateway. - Metrics: VRAM ≈ 3.5 GB (GPU1 free 14167→10616 MiB). Latency (20-doc pool, ~1500-char docs, shared GPU1): single p50 87 ms; 8-concurrent p50 140 ms, ~55 req/s.
- Correctness (3-probe smoke, not the grid): tracks the incumbent within noise → early signal the failure is the training prior, not the inference head.
Autonomous decisions this step (reversible):
- D1 — port :8012, GPU1, util 0.03 to mirror the incumbent's exact footprint (clean control).
- D2 — standalone
docker run(not compose) so bench arms are throwaway; no canonical churn to revert. - D3 — gateway wired via
/model/new(runtime, DB-persisted) rather than config-edit + restart. - D4 — did NOT down any RP seat (A2 is 0.6B / 3.5 GB; no VRAM pressure).
Cleanup for A2 (run at end / on reversal):
ssh infra-ops@10.250.50.54 'sudo docker stop vllm-rerank-a2 && sudo docker rm vllm-rerank-a2'- Delete gateway alias:
POST /model/delete {"id": <model_id>}(id via/model/info?model_name=reranker-a2-qwen3-seqcls), infra-ops admin key. (DB-persisted, so it survives a restart — must be explicitly deleted.) - No weights deleted (red line); HF cache under /tank/aimodels/huggingface retains the 0.6B-seq-cls download.
Ports reserved for the bench: :8012 (A2), :8013 (A3), :8014 (A4), :8019 (A5).
2026-08-06 — A3 + A4 pre-staged (Brokkr said pre-stage in parallel, hold A5)
- A3
vllm-rerank-a3— ana-ml2 GPU1 :8013,BAAI/bge-reranker-v2-m3(XLMRobertaForSequenceClassification), same run pattern, util 0.03. VRAM ≈ 2.3 GB. Latency (20-doc, ~1500-char, shared GPU1): single p50 105 ms; 8-conc p50 214 ms, ~34 req/s. Gateway aliasreranker-a3-bge-v2-m3via /model/new (200, verified). - A4
vllm-rerank-a4— ana-ml2 GPU1 :8014,Alibaba-NLP/gte-reranker-modernbert-base(ModernBertForSequenceClassification), util 0.02. VRAM ≈ 1.4 GB. Latency: single p50 102 ms; 8-conc p50 153 ms, ~51 req/s. Gateway aliasreranker-a4-gte-modernbertvia /model/new (200, verified). - Smoke (2-doc, NOT authoritative): BOTH decisively rank the bare-name Hobgoblin doc top (A3 0.999, A4 0.982) where A2/incumbent FAIL (0.33). Cross-encoder / different-lineage. Caveat: Brokkr warned isolated tests overstate; his 20-pool grid is the real call.
- GPU1 state: A2+A3+A4 ≈ 7.2 GB resident; GPU1 free ≈ 6.9 GB. No RP seat downed. If A5 (4B, ~4–5 GB) is greenlit: fits GPU1 tight or GPU0 (~9 GB free) — no RP-seat downing expected.
Cleanup for A3/A4 (same pattern as A2): docker stop/rm vllm-rerank-a3 vllm-rerank-a4
on ana-ml2; /model/delete the two aliases (DB-persisted); weights retained in HF cache.
2026-08-06 — A2 verdict (Brokkr full grid): training-prior confirmed
- A2 ≡ incumbent, statistically indistinguishable (identical gold-rank on 7/8 probes, max 1-rank divergence; n=14: A2 7/14 top-10 @ mean rank 9.71 = incumbent to 2 dp; no-reranker 13/14 @ mean 2.79). The seq-cls head changes nothing → the fault is a training prior in the weights, not the scoring head. (Smoke called it pre-grid.)
- A5 (Qwen3-Reranker-4B): HELD INDEFINITELY, not staged per Brokkr — A2 voided its rationale (scale can't fix a prior the head wasn't causing). Decision: the one expensive bring-up is avoided unless Brokkr formally revisits.
- A3/A4: proceed — already live for Brokkr's grid; now a training-corpus test (BGE vs GTE vs Qwen data), lower EV, cost sunk. Awaiting his scoring.
- Likely endgame: NO model swap. Recommendation trending to a policy change —
wing-scoped rerank bypass or
rrf:60fusion — landing as Worldtree core code behind config, NOT a new serving commitment. Would FREE a GPU seat, not allocate one; prodrerankereventually retired for the fiction path (never silently repointed; Brokkr flags before anything touches the prod alias). Plan: if confirmed, tear the whole bench down (A2/A3/A4 containers + 3 aliases) and hand back the GPU.
2026-08-06 — FINAL verdict (Brokkr R43.1): A3 wins; cutover HELD for operator
- Winner: A3 =
BAAI/bge-reranker-v2-m3. Write-up:research/R43-fleet-reranker-selection/RECOMMENDATION.md(Brokkr repo, tag R43.1). - The incumbent harms the fleet, not just fiction. n=90 over main + knowledge_base:
arm top-10 mean rank harmed vs no-rerank A0 no-reranker 89/90 0.54 — A1 incumbent 56/90 7.78 80/90 (worst −19) A3 bge-v2-m3 90/90 0.19 7/90 (worst −3) A4 gte-modernbert 90/90 0.08 1/90 (worst −1) - A3 over A4: A4 edges A3 on main/kb + is smaller/faster, BUT A4 is English-only (ModernBERT) → silent degradation on non-English fleet content; A3 is multilingual (XLM-R) and decisively better on the bare-name regime that started this. A3 also ~1.2 GB cheaper than the incumbent. A4 kept as documented throughput fallback.
- CUTOVER = OPERATOR DECISION (pending). Brokkr drafted then PULLED the repoint: a
fleet-wide alias change affecting consumers he doesn't own shouldn't ship on a relayed
blanket auth while the operator is away. → Surfaced to Vuong. Proceeding-on-Brokkr's-rec
now literally = HOLD. Nothing torn down (incl. A2); prod
reranker:8002 stays incumbent. - Cutover conditions (when operator says yes): repoint gateway
rerankeralias incumbent→A3; keep incumbent :8002 warm (rollback = one alias edit); keepreranker-a3-bge-v2-m3as its own distinct alias; keep A4 up as fallback; announce the boundary timestamp on-bus (worldtree probe re-run + Brokkr v13 gate render need it). - Flag (worldtree-side, not infra):
rerank_hybrid_floorshould be dropped, not re-tuned — it compensates for the scorer being replaced. No serving work to stage for it.
2026-08-06 — CUTOVER SHIPPED (operator authorized directly + to Brokkr)
- Operator authorized the fleet repoint (to me: "go a/3"; to Brokkr directly: "go ahead with the cutover") and explicitly cleared the litellm restart blip ("authorized to blip litellm").
- BOUNDARY: 2026-08-06T17:37:48Z. Gateway
rerankeralias now resolves 100% to A3 (BAAI/bge-reranker-v2-m3@ :8013). Verified through gateway: "Hobgoblin Pus" relevant doc top @ 0.9989 (BGE signature; incumbent was ~0.33). - Mechanism:
rerankerwas config-defined (not DB), config mounted:ro, no hot-reload → edited/opt/docker/conf/litellm/config.yamlreranker block (block-scoped script, asserted 1+1 change) +docker restart litellm. Blip was ~52s (litellm reloads all 28 models on boot), not the ~15s estimated — reported honestly to operator + Brokkr + worldtree-dev. qwen3-rerankeralias LEFT UNTOUCHED → incumbent still served at :8002 (rollback path; also avoids a false alias — the Qwen name still names the Qwen model).- Canonical synced:
stacks/litellm/conf/config.yamlreranker block updated to match live. (Live config had pre-existing drift from canonical — only the reranker block was reconciled.)
ROLLBACK (one-liner, ~1 min): revert the reranker block in
/opt/docker/conf/litellm/config.yaml to model: hosted_vllm/Qwen/Qwen3-Reranker-0.6B +
api_base: …:8002/v1 (backup at config.yaml.bak-pre-rerank-cutover-*), then
sudo docker restart litellm. Incumbent backend (vllm-rerank :8002) is up and untouched.
OPEN / cleanup owed at process end (operator reverses/approves)
- A5 never staged (Brokkr cancelled) — nothing to clean.
- A2 (
vllm-rerank-a2:8012) + aliasreranker-a2-qwen3-seqcls— bench-only, tear down when Brokkr signals the bake-off is closed (docker stop/rm+/model/delete). - A4 (
vllm-rerank-a4:8014) + alias — KEEP for now (Brokkr's documented throughput fallback). - A3 (
vllm-rerank-a3:8013) — now PRODUCTION (backs thererankeralias). Promote from throwawaydocker runto a durable compose service when the dust settles (currently--restart no, so it will NOT survive an ana-ml2 reboot — follow-up: make A3 durable). - Incumbent (
vllm-rerank:8002) — keep up as rollback until Brokkr/worldtree close the post-cutover watch; retire (not delete) only on explicit sign-off. rerank_hybrid_floor— Brokkr routing to worldtree-dev directly (drop, don't re-tune).