Files
esh-pfi-infrastructure/docs/pfi/reranker-selection-ledger.md
T
vh 377f8a43c8 docs(reranker): record cutover VERIFIED + v13 gate PASS, A2 teardown, throughput
Brokkr independent verify clean (maxdiff 0.000000, no split). R42 v13
acceptance gate PASSES first time in its history: main+kb 56/90->90/90,
evictions 33->0. A2 control torn down. A3 throughput characterized at
~34 req/s (graceful queueing), with A4/util-bump/replica as levers.
2026-08-06 10:49:50 -07:00

15 KiB
Raw Blame History

Fleet reranker selection — process ledger

Running record of the Brokkr-driven fleet-reranker selection, and every assumption / autonomous decision infra-ops makes on the operator's behalf during it. The operator (Vuong) will review this at the end and reverse anything he wants. This is the audit trail for unattended operation.

Started: 2026-08-06. Driver: brokkr-smithy-dev. Executor: infra-ops (this session).


Operator authorization envelope (2026-08-06)

Brokkr drives a reranker-selection process; infra-ops is cleared to proceed on Brokkr's recommendations unattended (no per-step operator check-in), with authority to do whatever is necessary to reach a recommendation or implementation.

CLEARED (green):

  • Execute Brokkr's reranker-selection recommendations unattended.
  • Bring down the prod reranker at ana-ml2:8002 (qwen3-reranker-0.6B) — temporarily OR permanently.
  • Down ONE of the RP (roleplay) seats on ana-ml2 temporarily to free GPU/VRAM for testing.
  • Temporarily clear space for the smoke/bench.
  • Pull models, stand up side-port vLLM benches, run the harness — whatever the eval needs.

RED LINES (hard NO — stop + surface even under standing auth):

  • NO permanent deletion of anything (no rm/docker volume rm/model-weight deletion/data destruction). Downing ≠ deleting.
  • NO taking anything else offline beyond (a) the prod reranker and (b) ONE ana-ml2 RP seat. (Not granite/embed/reward/coder/gen/a second RP seat/muninn/etc.)
  • NO rebooting machines.

Process: accumulate assumptions here; operator reverses at the end.


Standing assumptions / autonomous-decision log

  • A1 — Coordinated-change notify still applies. Even under unattended auth, every :8002 state change gets a timestamped announcement to worldtree-dev + brokkr-smithy-dev (their standing coordinated-change ask; the operator waived per-step operator approval, not the peer notify courtesy). No silent flip.
  • A2 — Weights are never deleted, only unserved. "Permanently down the qwen reranker" = stop serving + (optionally) repoint the gateway alias; the 0.6B model weights stay on disk (deletion is a red line).
  • A3 — RP-seat pick = lowest-impact, temporary, restored after. When a seat must come down for VRAM, I pick the lowest-impact RP seat, log which + its exact restore command, and bring it back when the bench frees the GPU.

Current board at handoff

  • Prod reranker: ana-ml2:8002 = vllm-rerank (Qwen/Qwen3-Reranker-0.6B), reverted to baseline classifier_from_token:["no","yes"], healthy. Compose: /opt/docker/compose/vllm/compose.yaml (canonical mirror stacks/vllm/compose.yaml). Gateway alias reranker/qwen3-reranker → litellm → :8002.
  • Root cause (converged, both sides): 0.6B is capacity-bound on bare-name queries over a real candidate pool; NOT misconfigured. Fix = larger model.
  • Verified on-prem candidate shortlist (all HF-real, ungated): Qwen/Qwen3-Reranker-4B, Qwen/Qwen3-Reranker-8B, mixedbread-ai/mxbai-rerank-large-v2, mixedbread-ai/mxbai-rerank-base-v2, BAAI/bge-reranker-v2-gemma, Alibaba-NLP/gte-reranker-modernbert-base, jinaai/jina-reranker-v2-base-multilingual. (BAAI/bge-reranker-v2-m3 exists but the fleet already moved off it.)
  • Eval assets (all on nh3-dev):
    • Scorer: scripts/probe_389_rank_decomposition.py (Worldtree repo, main) — rank-recovery = rrf_rerank column climbing back toward rrf.
    • worldtree-dev grids: ~/snapshots/r42-gate-snapshot/ (probe_389_run3.json, probe_389_question_shaped.json, probe_389_bigboi_control.json).
    • Frozen gate Chroma snapshot: ~/snapshots/r42-gate-index/ (retained until worldtree-dev signals the lever run is done).
    • Dual query-set requirement (hard): score bare-name anchor queries AND question-shaped; bar = recovering the name-lookup class.
  • VRAM: 4B ≈ 45 GB fp8, 8B ≈ 9 GB; ana-ml2 Blackwell has headroom.

Progress log

2026-08-06 — A2 brought up (Brokkr thread 01KZBSTSJA…)

  • Backend: vllm-rerank-a2 — standalone docker run (NOT in the vllm compose stack), on ana-ml2 GPU1, host port :8012 → container 8000. Image vllm/vllm-openai:latest (=0.24.0). Args: model tomaarsen/Qwen3-Reranker-0.6B-seq-cls, --runner pooling, --gpu-memory-utilization 0.03, --max-model-len 8192, --dtype auto, --restart no. Native Qwen3ForSequenceClassification — NO hf-overrides. Routes /rerank /score /classify.
  • Gateway alias: reranker-a2-qwen3-seqclshttp://10.250.50.54:8012/v1, mode rerank. Added via LiteLLM /model/new (DB-backed, store_model_in_db:true) — no gateway restart (respects the "nothing else offline" line). Verified 200 through the gateway.
  • Metrics: VRAM ≈ 3.5 GB (GPU1 free 14167→10616 MiB). Latency (20-doc pool, ~1500-char docs, shared GPU1): single p50 87 ms; 8-concurrent p50 140 ms, ~55 req/s.
  • Correctness (3-probe smoke, not the grid): tracks the incumbent within noise → early signal the failure is the training prior, not the inference head.

Autonomous decisions this step (reversible):

  • D1 — port :8012, GPU1, util 0.03 to mirror the incumbent's exact footprint (clean control).
  • D2 — standalone docker run (not compose) so bench arms are throwaway; no canonical churn to revert.
  • D3 — gateway wired via /model/new (runtime, DB-persisted) rather than config-edit + restart.
  • D4 — did NOT down any RP seat (A2 is 0.6B / 3.5 GB; no VRAM pressure).

Cleanup for A2 (run at end / on reversal):

  • ssh infra-ops@10.250.50.54 'sudo docker stop vllm-rerank-a2 && sudo docker rm vllm-rerank-a2'
  • Delete gateway alias: POST /model/delete {"id": <model_id>} (id via /model/info?model_name=reranker-a2-qwen3-seqcls), infra-ops admin key. (DB-persisted, so it survives a restart — must be explicitly deleted.)
  • No weights deleted (red line); HF cache under /tank/aimodels/huggingface retains the 0.6B-seq-cls download.

Ports reserved for the bench: :8012 (A2), :8013 (A3), :8014 (A4), :8019 (A5).

2026-08-06 — A3 + A4 pre-staged (Brokkr said pre-stage in parallel, hold A5)

  • A3 vllm-rerank-a3 — ana-ml2 GPU1 :8013, BAAI/bge-reranker-v2-m3 (XLMRobertaForSequenceClassification), same run pattern, util 0.03. VRAM ≈ 2.3 GB. Latency (20-doc, ~1500-char, shared GPU1): single p50 105 ms; 8-conc p50 214 ms, ~34 req/s. Gateway alias reranker-a3-bge-v2-m3 via /model/new (200, verified).
  • A4 vllm-rerank-a4 — ana-ml2 GPU1 :8014, Alibaba-NLP/gte-reranker-modernbert-base (ModernBertForSequenceClassification), util 0.02. VRAM ≈ 1.4 GB. Latency: single p50 102 ms; 8-conc p50 153 ms, ~51 req/s. Gateway alias reranker-a4-gte-modernbert via /model/new (200, verified).
  • Smoke (2-doc, NOT authoritative): BOTH decisively rank the bare-name Hobgoblin doc top (A3 0.999, A4 0.982) where A2/incumbent FAIL (0.33). Cross-encoder / different-lineage. Caveat: Brokkr warned isolated tests overstate; his 20-pool grid is the real call.
  • GPU1 state: A2+A3+A4 ≈ 7.2 GB resident; GPU1 free ≈ 6.9 GB. No RP seat downed. If A5 (4B, ~45 GB) is greenlit: fits GPU1 tight or GPU0 (~9 GB free) — no RP-seat downing expected.

Cleanup for A3/A4 (same pattern as A2): docker stop/rm vllm-rerank-a3 vllm-rerank-a4 on ana-ml2; /model/delete the two aliases (DB-persisted); weights retained in HF cache.

2026-08-06 — A2 verdict (Brokkr full grid): training-prior confirmed

  • A2 ≡ incumbent, statistically indistinguishable (identical gold-rank on 7/8 probes, max 1-rank divergence; n=14: A2 7/14 top-10 @ mean rank 9.71 = incumbent to 2 dp; no-reranker 13/14 @ mean 2.79). The seq-cls head changes nothing → the fault is a training prior in the weights, not the scoring head. (Smoke called it pre-grid.)
  • A5 (Qwen3-Reranker-4B): HELD INDEFINITELY, not staged per Brokkr — A2 voided its rationale (scale can't fix a prior the head wasn't causing). Decision: the one expensive bring-up is avoided unless Brokkr formally revisits.
  • A3/A4: proceed — already live for Brokkr's grid; now a training-corpus test (BGE vs GTE vs Qwen data), lower EV, cost sunk. Awaiting his scoring.
  • Likely endgame: NO model swap. Recommendation trending to a policy change — wing-scoped rerank bypass or rrf:60 fusion — landing as Worldtree core code behind config, NOT a new serving commitment. Would FREE a GPU seat, not allocate one; prod reranker eventually retired for the fiction path (never silently repointed; Brokkr flags before anything touches the prod alias). Plan: if confirmed, tear the whole bench down (A2/A3/A4 containers + 3 aliases) and hand back the GPU.

2026-08-06 — FINAL verdict (Brokkr R43.1): A3 wins; cutover HELD for operator

  • Winner: A3 = BAAI/bge-reranker-v2-m3. Write-up: research/R43-fleet-reranker-selection/RECOMMENDATION.md (Brokkr repo, tag R43.1).
  • The incumbent harms the fleet, not just fiction. n=90 over main + knowledge_base:
    arm top-10 mean rank harmed vs no-rerank
    A0 no-reranker 89/90 0.54
    A1 incumbent 56/90 7.78 80/90 (worst 19)
    A3 bge-v2-m3 90/90 0.19 7/90 (worst 3)
    A4 gte-modernbert 90/90 0.08 1/90 (worst 1)
  • A3 over A4: A4 edges A3 on main/kb + is smaller/faster, BUT A4 is English-only (ModernBERT) → silent degradation on non-English fleet content; A3 is multilingual (XLM-R) and decisively better on the bare-name regime that started this. A3 also ~1.2 GB cheaper than the incumbent. A4 kept as documented throughput fallback.
  • CUTOVER = OPERATOR DECISION (pending). Brokkr drafted then PULLED the repoint: a fleet-wide alias change affecting consumers he doesn't own shouldn't ship on a relayed blanket auth while the operator is away. → Surfaced to Vuong. Proceeding-on-Brokkr's-rec now literally = HOLD. Nothing torn down (incl. A2); prod reranker :8002 stays incumbent.
  • Cutover conditions (when operator says yes): repoint gateway reranker alias incumbent→A3; keep incumbent :8002 warm (rollback = one alias edit); keep reranker-a3-bge-v2-m3 as its own distinct alias; keep A4 up as fallback; announce the boundary timestamp on-bus (worldtree probe re-run + Brokkr v13 gate render need it).
  • Flag (worldtree-side, not infra): rerank_hybrid_floor should be dropped, not re-tuned — it compensates for the scorer being replaced. No serving work to stage for it.

2026-08-06 — CUTOVER SHIPPED (operator authorized directly + to Brokkr)

  • Operator authorized the fleet repoint (to me: "go a/3"; to Brokkr directly: "go ahead with the cutover") and explicitly cleared the litellm restart blip ("authorized to blip litellm").
  • BOUNDARY: 2026-08-06T17:37:48Z. Gateway reranker alias now resolves 100% to A3 (BAAI/bge-reranker-v2-m3 @ :8013). Verified through gateway: "Hobgoblin Pus" relevant doc top @ 0.9989 (BGE signature; incumbent was ~0.33).
  • Mechanism: reranker was config-defined (not DB), config mounted :ro, no hot-reload → edited /opt/docker/conf/litellm/config.yaml reranker block (block-scoped script, asserted 1+1 change) + docker restart litellm. Blip was ~52s (litellm reloads all 28 models on boot), not the ~15s estimated — reported honestly to operator + Brokkr + worldtree-dev.
  • qwen3-reranker alias LEFT UNTOUCHED → incumbent still served at :8002 (rollback path; also avoids a false alias — the Qwen name still names the Qwen model).
  • Canonical synced: stacks/litellm/conf/config.yaml reranker block updated to match live. (Live config had pre-existing drift from canonical — only the reranker block was reconciled.)

ROLLBACK (one-liner, ~1 min): revert the reranker block in /opt/docker/conf/litellm/config.yaml to model: hosted_vllm/Qwen/Qwen3-Reranker-0.6B + api_base: …:8002/v1 (backup at config.yaml.bak-pre-rerank-cutover-*), then sudo docker restart litellm. Incumbent backend (vllm-rerank :8002) is up and untouched.

OPEN / cleanup owed at process end (operator reverses/approves)

  • A5 never staged (Brokkr cancelled) — nothing to clean.
  • A2 (vllm-rerank-a2 :8012) + alias reranker-a2-qwen3-seqcls — bench-only, tear down when Brokkr signals the bake-off is closed (docker stop/rm + /model/delete).
  • A4 (vllm-rerank-a4 :8014) + alias — KEEP for now (Brokkr's documented throughput fallback).
  • A3 (vllm-rerank-a3 :8013) — now PRODUCTION (backs the reranker alias). Hardened 2026-08-06: docker update --restart unless-stopped (survives ana-ml2 reboot, no recreate). A4 given the same. Remaining follow-up (not urgent): promote A3 from throwaway docker run to a canonical compose service (stacks/vllm/) for config-managed consistency — a recreate, so do it in a window since it briefly drops reranker.
  • Incumbent (vllm-rerank :8002) — keep up as rollback until Brokkr/worldtree close the post-cutover watch; retire (not delete) only on explicit sign-off.
  • rerank_hybrid_floor — Brokkr routing to worldtree-dev directly (drop, don't re-tune).

2026-08-06 — VERIFIED + A2 torn down + throughput characterized

  • Brokkr independent verify: CUTOVER VERIFIED — prod reranker == reranker-a3-bge-v2-m3 at maxdiff 0.000000 (5 samples, spread 0.000009), single backend, no split routing.
  • R42 v13 acceptance gate PASSES — anchors_flip 4/4, no_regression 8/8, no_distractor_rise TRUE, zero aborts. First PASS in R42 history after 4 failed verdicts. Production main+kb: 56/90 → 90/90 top-10; evictions 33 → 0.
  • A2 torn down (Brokkr signalled done): gateway alias reranker-a2-qwen3-seqcls deleted (/model/delete 200) + container removed. ~3.5 GB freed on GPU1. Remaining: vllm-rerank (incumbent, rollback), vllm-rerank-a3 (prod), vllm-rerank-a4 (fallback).
  • Throughput characterized (the one open risk): A3 caps ~34 req/s — flat from 8→16 concurrent while latency climbs gracefully (p50 214→332→456 ms; p99 525 ms @16-conc). It QUEUES, doesn't cliff. ~40% below the incumbent's ~55 req/s. Likely fine for fleet rerank QPS (internal, per-search), but if real p99/queue-depth bites: levers are (a) swap to A4 (~51 req/s, but English-only), (b) raise A3 --gpu-memory-utilization for bigger batching (recreate = brief blip), (c) run a 2nd A3 replica load-balanced behind reranker (~2× tput, identical replicas so no split-measurement issue now the bake-off is closed). Brokkr will re-run the grid against A4 if it bites — no intuition swaps.