377f8a43c8
Brokkr independent verify clean (maxdiff 0.000000, no split). R42 v13 acceptance gate PASSES first time in its history: main+kb 56/90->90/90, evictions 33->0. A2 control torn down. A3 throughput characterized at ~34 req/s (graceful queueing), with A4/util-bump/replica as levers.
231 lines
15 KiB
Markdown
231 lines
15 KiB
Markdown
# Fleet reranker selection — process ledger
|
||
|
||
Running record of the Brokkr-driven fleet-reranker selection, and every
|
||
assumption / autonomous decision infra-ops makes on the operator's behalf
|
||
during it. The operator (Vuong) will review this at the end and reverse
|
||
anything he wants. **This is the audit trail for unattended operation.**
|
||
|
||
Started: 2026-08-06. Driver: **brokkr-smithy-dev**. Executor: **infra-ops** (this session).
|
||
|
||
---
|
||
|
||
## Operator authorization envelope (2026-08-06)
|
||
|
||
Brokkr drives a reranker-selection process; infra-ops is cleared to proceed on
|
||
Brokkr's recommendations **unattended** (no per-step operator check-in), with
|
||
authority to do whatever is necessary to reach a recommendation **or**
|
||
implementation.
|
||
|
||
**CLEARED (green):**
|
||
- Execute Brokkr's reranker-selection recommendations unattended.
|
||
- Bring **down the prod reranker** at `ana-ml2:8002` (qwen3-reranker-0.6B) —
|
||
**temporarily OR permanently**.
|
||
- Down **ONE** of the RP (roleplay) seats on ana-ml2 **temporarily** to free
|
||
GPU/VRAM for testing.
|
||
- Temporarily clear space for the smoke/bench.
|
||
- Pull models, stand up side-port vLLM benches, run the harness — whatever the
|
||
eval needs.
|
||
|
||
**RED LINES (hard NO — stop + surface even under standing auth):**
|
||
- **NO permanent deletion of anything** (no `rm`/`docker volume rm`/model-weight
|
||
deletion/data destruction). Downing ≠ deleting.
|
||
- **NO taking anything else offline** beyond (a) the prod reranker and (b) ONE
|
||
ana-ml2 RP seat. (Not granite/embed/reward/coder/gen/a second RP seat/muninn/etc.)
|
||
- **NO rebooting machines.**
|
||
|
||
**Process:** accumulate assumptions here; operator reverses at the end.
|
||
|
||
---
|
||
|
||
## Standing assumptions / autonomous-decision log
|
||
|
||
- **A1 — Coordinated-change notify still applies.** Even under unattended auth,
|
||
every `:8002` state change gets a timestamped announcement to worldtree-dev +
|
||
brokkr-smithy-dev (their standing coordinated-change ask; the operator waived
|
||
per-step *operator* approval, not the peer *notify* courtesy). No silent flip.
|
||
- **A2 — Weights are never deleted, only unserved.** "Permanently down the qwen
|
||
reranker" = stop serving + (optionally) repoint the gateway alias; the 0.6B
|
||
model weights stay on disk (deletion is a red line).
|
||
- **A3 — RP-seat pick = lowest-impact, temporary, restored after.** When a seat
|
||
must come down for VRAM, I pick the lowest-impact RP seat, log which + its
|
||
exact restore command, and bring it back when the bench frees the GPU.
|
||
|
||
---
|
||
|
||
## Current board at handoff
|
||
|
||
- **Prod reranker:** `ana-ml2:8002` = `vllm-rerank` (Qwen/Qwen3-Reranker-0.6B),
|
||
reverted to baseline `classifier_from_token:["no","yes"]`, healthy. Compose:
|
||
`/opt/docker/compose/vllm/compose.yaml` (canonical mirror
|
||
`stacks/vllm/compose.yaml`). Gateway alias `reranker`/`qwen3-reranker` →
|
||
litellm → :8002.
|
||
- **Root cause (converged, both sides):** 0.6B is capacity-bound on bare-name
|
||
queries over a real candidate pool; NOT misconfigured. Fix = larger model.
|
||
- **Verified on-prem candidate shortlist (all HF-real, ungated):**
|
||
Qwen/Qwen3-Reranker-4B, Qwen/Qwen3-Reranker-8B, mixedbread-ai/mxbai-rerank-large-v2,
|
||
mixedbread-ai/mxbai-rerank-base-v2, BAAI/bge-reranker-v2-gemma,
|
||
Alibaba-NLP/gte-reranker-modernbert-base, jinaai/jina-reranker-v2-base-multilingual.
|
||
(BAAI/bge-reranker-v2-m3 exists but the fleet already moved off it.)
|
||
- **Eval assets (all on nh3-dev):**
|
||
- Scorer: `scripts/probe_389_rank_decomposition.py` (Worldtree repo, main) —
|
||
rank-recovery = `rrf_rerank` column climbing back toward `rrf`.
|
||
- worldtree-dev grids: `~/snapshots/r42-gate-snapshot/` (probe_389_run3.json,
|
||
probe_389_question_shaped.json, probe_389_bigboi_control.json).
|
||
- Frozen gate Chroma snapshot: `~/snapshots/r42-gate-index/` (retained until
|
||
worldtree-dev signals the lever run is done).
|
||
- **Dual query-set requirement (hard):** score bare-name anchor queries AND
|
||
question-shaped; bar = recovering the name-lookup class.
|
||
- **VRAM:** 4B ≈ 4–5 GB fp8, 8B ≈ 9 GB; ana-ml2 Blackwell has headroom.
|
||
|
||
---
|
||
|
||
## Progress log
|
||
|
||
### 2026-08-06 — A2 brought up (Brokkr thread 01KZBSTSJA…)
|
||
|
||
- **Backend:** `vllm-rerank-a2` — standalone `docker run` (NOT in the vllm compose
|
||
stack), on ana-ml2 **GPU1**, host port **:8012** → container 8000. Image
|
||
`vllm/vllm-openai:latest` (=0.24.0). Args: model
|
||
`tomaarsen/Qwen3-Reranker-0.6B-seq-cls`, `--runner pooling`, `--gpu-memory-utilization
|
||
0.03`, `--max-model-len 8192`, `--dtype auto`, `--restart no`. Native
|
||
`Qwen3ForSequenceClassification` — NO hf-overrides. Routes /rerank /score /classify.
|
||
- **Gateway alias:** `reranker-a2-qwen3-seqcls` → `http://10.250.50.54:8012/v1`,
|
||
mode rerank. Added via LiteLLM **`/model/new`** (DB-backed, `store_model_in_db:true`)
|
||
— **no gateway restart** (respects the "nothing else offline" line). Verified 200
|
||
through the gateway.
|
||
- **Metrics:** VRAM ≈ **3.5 GB** (GPU1 free 14167→10616 MiB). Latency (20-doc pool,
|
||
~1500-char docs, shared GPU1): single p50 **87 ms**; 8-concurrent p50 **140 ms**,
|
||
~**55 req/s**.
|
||
- **Correctness (3-probe smoke, not the grid):** tracks the incumbent within noise →
|
||
early signal the failure is the **training prior, not the inference head**.
|
||
|
||
**Autonomous decisions this step (reversible):**
|
||
- D1 — port :8012, GPU1, util 0.03 to mirror the incumbent's exact footprint (clean control).
|
||
- D2 — standalone `docker run` (not compose) so bench arms are throwaway; no canonical churn to revert.
|
||
- D3 — gateway wired via `/model/new` (runtime, DB-persisted) rather than config-edit + restart.
|
||
- D4 — did NOT down any RP seat (A2 is 0.6B / 3.5 GB; no VRAM pressure).
|
||
|
||
**Cleanup for A2 (run at end / on reversal):**
|
||
- `ssh infra-ops@10.250.50.54 'sudo docker stop vllm-rerank-a2 && sudo docker rm vllm-rerank-a2'`
|
||
- Delete gateway alias: `POST /model/delete {"id": <model_id>}` (id via `/model/info?model_name=reranker-a2-qwen3-seqcls`), infra-ops admin key. (DB-persisted, so it survives a restart — must be explicitly deleted.)
|
||
- No weights deleted (red line); HF cache under /tank/aimodels/huggingface retains the 0.6B-seq-cls download.
|
||
|
||
**Ports reserved for the bench:** :8012 (A2), :8013 (A3), :8014 (A4), :8019 (A5).
|
||
|
||
### 2026-08-06 — A3 + A4 pre-staged (Brokkr said pre-stage in parallel, hold A5)
|
||
|
||
- **A3** `vllm-rerank-a3` — ana-ml2 GPU1 :8013, `BAAI/bge-reranker-v2-m3`
|
||
(XLMRobertaForSequenceClassification), same run pattern, util 0.03. VRAM ≈ **2.3 GB**.
|
||
Latency (20-doc, ~1500-char, shared GPU1): single p50 **105 ms**; 8-conc p50 214 ms, ~34 req/s.
|
||
Gateway alias `reranker-a3-bge-v2-m3` via /model/new (200, verified).
|
||
- **A4** `vllm-rerank-a4` — ana-ml2 GPU1 :8014, `Alibaba-NLP/gte-reranker-modernbert-base`
|
||
(ModernBertForSequenceClassification), util 0.02. VRAM ≈ **1.4 GB**. Latency: single
|
||
p50 **102 ms**; 8-conc p50 153 ms, ~51 req/s. Gateway alias `reranker-a4-gte-modernbert`
|
||
via /model/new (200, verified).
|
||
- **Smoke (2-doc, NOT authoritative):** BOTH decisively rank the bare-name Hobgoblin doc top
|
||
(A3 0.999, A4 0.982) where A2/incumbent FAIL (0.33). Cross-encoder / different-lineage.
|
||
Caveat: Brokkr warned isolated tests overstate; his 20-pool grid is the real call.
|
||
- **GPU1 state:** A2+A3+A4 ≈ 7.2 GB resident; GPU1 free ≈ **6.9 GB**. No RP seat downed.
|
||
If A5 (4B, ~4–5 GB) is greenlit: fits GPU1 tight or GPU0 (~9 GB free) — no RP-seat downing expected.
|
||
|
||
**Cleanup for A3/A4 (same pattern as A2):** `docker stop/rm vllm-rerank-a3 vllm-rerank-a4`
|
||
on ana-ml2; `/model/delete` the two aliases (DB-persisted); weights retained in HF cache.
|
||
|
||
### 2026-08-06 — A2 verdict (Brokkr full grid): training-prior confirmed
|
||
|
||
- **A2 ≡ incumbent, statistically indistinguishable** (identical gold-rank on 7/8 probes,
|
||
max 1-rank divergence; n=14: A2 7/14 top-10 @ mean rank 9.71 = incumbent to 2 dp;
|
||
no-reranker 13/14 @ mean 2.79). The seq-cls head changes nothing → the fault is a
|
||
**training prior in the weights**, not the scoring head. (Smoke called it pre-grid.)
|
||
- **A5 (Qwen3-Reranker-4B): HELD INDEFINITELY, not staged** per Brokkr — A2 voided its
|
||
rationale (scale can't fix a prior the head wasn't causing). *Decision: the one expensive
|
||
bring-up is avoided unless Brokkr formally revisits.*
|
||
- **A3/A4:** proceed — already live for Brokkr's grid; now a training-corpus test (BGE vs
|
||
GTE vs Qwen data), lower EV, cost sunk. Awaiting his scoring.
|
||
- **Likely endgame:** NO model swap. Recommendation trending to a **policy change** —
|
||
wing-scoped rerank bypass or `rrf:60` fusion — landing as Worldtree core code behind
|
||
config, NOT a new serving commitment. Would FREE a GPU seat, not allocate one; prod
|
||
`reranker` eventually retired for the fiction path (never silently repointed; Brokkr
|
||
flags before anything touches the prod alias). *Plan: if confirmed, tear the whole bench
|
||
down (A2/A3/A4 containers + 3 aliases) and hand back the GPU.*
|
||
|
||
### 2026-08-06 — FINAL verdict (Brokkr R43.1): A3 wins; cutover HELD for operator
|
||
|
||
- **Winner: A3 = `BAAI/bge-reranker-v2-m3`.** Write-up:
|
||
`research/R43-fleet-reranker-selection/RECOMMENDATION.md` (Brokkr repo, tag R43.1).
|
||
- **The incumbent harms the fleet, not just fiction.** n=90 over main + knowledge_base:
|
||
| arm | top-10 | mean rank | harmed vs no-rerank |
|
||
|---|:--:|:--:|:--:|
|
||
| A0 no-reranker | 89/90 | 0.54 | — |
|
||
| A1 incumbent | 56/90 | 7.78 | **80/90 (worst −19)** |
|
||
| **A3 bge-v2-m3** | 90/90 | 0.19 | 7/90 (worst −3) |
|
||
| A4 gte-modernbert | 90/90 | 0.08 | 1/90 (worst −1) |
|
||
- **A3 over A4:** A4 edges A3 on main/kb + is smaller/faster, BUT A4 is **English-only
|
||
(ModernBERT)** → silent degradation on non-English fleet content; A3 is **multilingual
|
||
(XLM-R)** and decisively better on the bare-name regime that started this. A3 also ~1.2 GB
|
||
*cheaper* than the incumbent. A4 kept as documented throughput fallback.
|
||
- **CUTOVER = OPERATOR DECISION (pending).** Brokkr drafted then PULLED the repoint: a
|
||
fleet-wide alias change affecting consumers he doesn't own shouldn't ship on a relayed
|
||
blanket auth while the operator is away. → Surfaced to Vuong. Proceeding-on-Brokkr's-rec
|
||
now literally = HOLD. **Nothing torn down (incl. A2); prod `reranker` :8002 stays incumbent.**
|
||
- **Cutover conditions (when operator says yes):** repoint gateway `reranker` alias
|
||
incumbent→A3; keep incumbent :8002 warm (rollback = one alias edit); keep
|
||
`reranker-a3-bge-v2-m3` as its own distinct alias; keep A4 up as fallback; **announce the
|
||
boundary timestamp on-bus** (worldtree probe re-run + Brokkr v13 gate render need it).
|
||
- **Flag (worldtree-side, not infra):** `rerank_hybrid_floor` should be **dropped, not
|
||
re-tuned** — it compensates for the scorer being replaced. No serving work to stage for it.
|
||
|
||
### 2026-08-06 — CUTOVER SHIPPED (operator authorized directly + to Brokkr)
|
||
|
||
- **Operator authorized** the fleet repoint (to me: "go a/3"; to Brokkr directly: "go ahead
|
||
with the cutover") and explicitly cleared the litellm restart blip ("authorized to blip litellm").
|
||
- **BOUNDARY: 2026-08-06T17:37:48Z.** Gateway `reranker` alias now resolves 100% to A3
|
||
(`BAAI/bge-reranker-v2-m3` @ :8013). Verified through gateway: "Hobgoblin Pus" relevant
|
||
doc top @ 0.9989 (BGE signature; incumbent was ~0.33).
|
||
- **Mechanism:** `reranker` was config-defined (not DB), config mounted `:ro`, no hot-reload →
|
||
edited `/opt/docker/conf/litellm/config.yaml` reranker block (block-scoped script, asserted
|
||
1+1 change) + `docker restart litellm`. **Blip was ~52s** (litellm reloads all 28 models on
|
||
boot), not the ~15s estimated — reported honestly to operator + Brokkr + worldtree-dev.
|
||
- **`qwen3-reranker` alias LEFT UNTOUCHED** → incumbent still served at :8002 (rollback path;
|
||
also avoids a false alias — the Qwen name still names the Qwen model).
|
||
- **Canonical synced:** `stacks/litellm/conf/config.yaml` reranker block updated to match live.
|
||
(Live config had pre-existing drift from canonical — only the reranker block was reconciled.)
|
||
|
||
**ROLLBACK (one-liner, ~1 min):** revert the `reranker` block in
|
||
`/opt/docker/conf/litellm/config.yaml` to `model: hosted_vllm/Qwen/Qwen3-Reranker-0.6B` +
|
||
`api_base: …:8002/v1` (backup at `config.yaml.bak-pre-rerank-cutover-*`), then
|
||
`sudo docker restart litellm`. Incumbent backend (`vllm-rerank` :8002) is up and untouched.
|
||
|
||
### OPEN / cleanup owed at process end (operator reverses/approves)
|
||
- **A5** never staged (Brokkr cancelled) — nothing to clean.
|
||
- **A2 (`vllm-rerank-a2` :8012)** + alias `reranker-a2-qwen3-seqcls` — bench-only, tear down when
|
||
Brokkr signals the bake-off is closed (`docker stop/rm` + `/model/delete`).
|
||
- **A4 (`vllm-rerank-a4` :8014)** + alias — KEEP for now (Brokkr's documented throughput fallback).
|
||
- **A3 (`vllm-rerank-a3` :8013)** — now PRODUCTION (backs the `reranker` alias). Hardened
|
||
2026-08-06: `docker update --restart unless-stopped` (survives ana-ml2 reboot, no recreate).
|
||
A4 given the same. **Remaining follow-up (not urgent): promote A3 from throwaway `docker run`
|
||
to a canonical compose service** (`stacks/vllm/`) for config-managed consistency — a recreate,
|
||
so do it in a window since it briefly drops `reranker`.
|
||
- **Incumbent (`vllm-rerank` :8002)** — keep up as rollback until Brokkr/worldtree close the
|
||
post-cutover watch; retire (not delete) only on explicit sign-off.
|
||
- **`rerank_hybrid_floor`** — Brokkr routing to worldtree-dev directly (drop, don't re-tune).
|
||
|
||
### 2026-08-06 — VERIFIED + A2 torn down + throughput characterized
|
||
|
||
- **Brokkr independent verify: CUTOVER VERIFIED** — prod `reranker` == `reranker-a3-bge-v2-m3`
|
||
at maxdiff 0.000000 (5 samples, spread 0.000009), single backend, no split routing.
|
||
- **R42 v13 acceptance gate PASSES** — anchors_flip 4/4, no_regression 8/8, no_distractor_rise
|
||
TRUE, zero aborts. **First PASS in R42 history after 4 failed verdicts.** Production main+kb:
|
||
56/90 → 90/90 top-10; evictions 33 → 0.
|
||
- **A2 torn down** (Brokkr signalled done): gateway alias `reranker-a2-qwen3-seqcls` deleted
|
||
(/model/delete 200) + container removed. ~3.5 GB freed on GPU1. Remaining: `vllm-rerank`
|
||
(incumbent, rollback), `vllm-rerank-a3` (prod), `vllm-rerank-a4` (fallback).
|
||
- **Throughput characterized (the one open risk):** A3 caps ~34 req/s — flat from 8→16
|
||
concurrent while latency climbs gracefully (p50 214→332→456 ms; p99 525 ms @16-conc). It
|
||
QUEUES, doesn't cliff. ~40% below the incumbent's ~55 req/s. Likely fine for fleet rerank
|
||
QPS (internal, per-search), but if real p99/queue-depth bites: levers are (a) swap to A4
|
||
(~51 req/s, but English-only), (b) raise A3 `--gpu-memory-utilization` for bigger batching
|
||
(recreate = brief blip), (c) run a 2nd A3 replica load-balanced behind `reranker` (~2× tput,
|
||
identical replicas so no split-measurement issue now the bake-off is closed). Brokkr will
|
||
re-run the grid against A4 if it bites — no intuition swaps.
|