The box is physically at Fountain Valley, renamed, renumbered onto 10.251/16, and serving inference again. This lands the repo half of that. Host: hostname ana-ml2 -> fv-ml1, pinned to 10.251.50.54 by a dnsmasq reservation so the address the runbook, DNS and LiteLLM all assume is the address it actually has. Its headscale node is renamed too. The sweep ran from scripts/fv-ml1-rename-sweep.sh, whose allowlist is the reason this diff touches current-state files and not the record. Dated persistent-memory entries, archival-memory and incident notes still say ana-ml2 in 31 and 62 places respectively, because that is what the box was when those things happened. Rewriting them would make the history lie. LiteLLM was the load-bearing piece and needed more than the api_base sed the runbook describes. Twenty api_base entries repointed, but a grep-and-verify pass also caught a LIVE pass_through_endpoints target for the scalar-judge reward route still on the old address -- an api_base-only substitution would have left it dead. Four prose references describing current state were repointed as well; one historical note recording where a hand-test was run is deliberately left pointing at 10.250.50.54. Two facts in the server tables were wrong and are corrected here. The site is Fountain Valley, not Anaheim. And the box has FOUR RTX PRO 6000 Blackwell Max-Q, not two -- verified by nvidia-smi -L and independently by PCI enumeration of four GB202GL devices. That is 391 GB of VRAM rather than 196, which changes what fits on it. DNS: fv-ml1, fv-ml1-bmc and fv-gw added under the fv site via the piggyback approach, scriberr re-homed, and the ana-ml2 records removed. Applied to all three resolvers. The BMC record carries a warning that its 802.1q VLAN tag must stay disabled -- it shipped tagging VLAN 250 into an untagged port, which made it invisible to every network-side diagnostic and is the reason it appeared dead through several cable changes. Verified end to end: summarizer and sec both answer through the Anaheim gateway across the mesh to FV seats on different ports.
231 lines
15 KiB
Markdown
231 lines
15 KiB
Markdown
# Fleet reranker selection — process ledger
|
||
|
||
Running record of the Brokkr-driven fleet-reranker selection, and every
|
||
assumption / autonomous decision infra-ops makes on the operator's behalf
|
||
during it. The operator (Vuong) will review this at the end and reverse
|
||
anything he wants. **This is the audit trail for unattended operation.**
|
||
|
||
Started: 2026-08-06. Driver: **brokkr-smithy-dev**. Executor: **infra-ops** (this session).
|
||
|
||
---
|
||
|
||
## Operator authorization envelope (2026-08-06)
|
||
|
||
Brokkr drives a reranker-selection process; infra-ops is cleared to proceed on
|
||
Brokkr's recommendations **unattended** (no per-step operator check-in), with
|
||
authority to do whatever is necessary to reach a recommendation **or**
|
||
implementation.
|
||
|
||
**CLEARED (green):**
|
||
- Execute Brokkr's reranker-selection recommendations unattended.
|
||
- Bring **down the prod reranker** at `fv-ml1:8002` (qwen3-reranker-0.6B) —
|
||
**temporarily OR permanently**.
|
||
- Down **ONE** of the RP (roleplay) seats on fv-ml1 **temporarily** to free
|
||
GPU/VRAM for testing.
|
||
- Temporarily clear space for the smoke/bench.
|
||
- Pull models, stand up side-port vLLM benches, run the harness — whatever the
|
||
eval needs.
|
||
|
||
**RED LINES (hard NO — stop + surface even under standing auth):**
|
||
- **NO permanent deletion of anything** (no `rm`/`docker volume rm`/model-weight
|
||
deletion/data destruction). Downing ≠ deleting.
|
||
- **NO taking anything else offline** beyond (a) the prod reranker and (b) ONE
|
||
fv-ml1 RP seat. (Not granite/embed/reward/coder/gen/a second RP seat/muninn/etc.)
|
||
- **NO rebooting machines.**
|
||
|
||
**Process:** accumulate assumptions here; operator reverses at the end.
|
||
|
||
---
|
||
|
||
## Standing assumptions / autonomous-decision log
|
||
|
||
- **A1 — Coordinated-change notify still applies.** Even under unattended auth,
|
||
every `:8002` state change gets a timestamped announcement to worldtree-dev +
|
||
brokkr-smithy-dev (their standing coordinated-change ask; the operator waived
|
||
per-step *operator* approval, not the peer *notify* courtesy). No silent flip.
|
||
- **A2 — Weights are never deleted, only unserved.** "Permanently down the qwen
|
||
reranker" = stop serving + (optionally) repoint the gateway alias; the 0.6B
|
||
model weights stay on disk (deletion is a red line).
|
||
- **A3 — RP-seat pick = lowest-impact, temporary, restored after.** When a seat
|
||
must come down for VRAM, I pick the lowest-impact RP seat, log which + its
|
||
exact restore command, and bring it back when the bench frees the GPU.
|
||
|
||
---
|
||
|
||
## Current board at handoff
|
||
|
||
- **Prod reranker:** `fv-ml1:8002` = `vllm-rerank` (Qwen/Qwen3-Reranker-0.6B),
|
||
reverted to baseline `classifier_from_token:["no","yes"]`, healthy. Compose:
|
||
`/opt/docker/compose/vllm/compose.yaml` (canonical mirror
|
||
`stacks/vllm/compose.yaml`). Gateway alias `reranker`/`qwen3-reranker` →
|
||
litellm → :8002.
|
||
- **Root cause (converged, both sides):** 0.6B is capacity-bound on bare-name
|
||
queries over a real candidate pool; NOT misconfigured. Fix = larger model.
|
||
- **Verified on-prem candidate shortlist (all HF-real, ungated):**
|
||
Qwen/Qwen3-Reranker-4B, Qwen/Qwen3-Reranker-8B, mixedbread-ai/mxbai-rerank-large-v2,
|
||
mixedbread-ai/mxbai-rerank-base-v2, BAAI/bge-reranker-v2-gemma,
|
||
Alibaba-NLP/gte-reranker-modernbert-base, jinaai/jina-reranker-v2-base-multilingual.
|
||
(BAAI/bge-reranker-v2-m3 exists but the fleet already moved off it.)
|
||
- **Eval assets (all on nh3-dev):**
|
||
- Scorer: `scripts/probe_389_rank_decomposition.py` (Worldtree repo, main) —
|
||
rank-recovery = `rrf_rerank` column climbing back toward `rrf`.
|
||
- worldtree-dev grids: `~/snapshots/r42-gate-snapshot/` (probe_389_run3.json,
|
||
probe_389_question_shaped.json, probe_389_bigboi_control.json).
|
||
- Frozen gate Chroma snapshot: `~/snapshots/r42-gate-index/` (retained until
|
||
worldtree-dev signals the lever run is done).
|
||
- **Dual query-set requirement (hard):** score bare-name anchor queries AND
|
||
question-shaped; bar = recovering the name-lookup class.
|
||
- **VRAM:** 4B ≈ 4–5 GB fp8, 8B ≈ 9 GB; fv-ml1 Blackwell has headroom.
|
||
|
||
---
|
||
|
||
## Progress log
|
||
|
||
### 2026-08-06 — A2 brought up (Brokkr thread 01KZBSTSJA…)
|
||
|
||
- **Backend:** `vllm-rerank-a2` — standalone `docker run` (NOT in the vllm compose
|
||
stack), on fv-ml1 **GPU1**, host port **:8012** → container 8000. Image
|
||
`vllm/vllm-openai:latest` (=0.24.0). Args: model
|
||
`tomaarsen/Qwen3-Reranker-0.6B-seq-cls`, `--runner pooling`, `--gpu-memory-utilization
|
||
0.03`, `--max-model-len 8192`, `--dtype auto`, `--restart no`. Native
|
||
`Qwen3ForSequenceClassification` — NO hf-overrides. Routes /rerank /score /classify.
|
||
- **Gateway alias:** `reranker-a2-qwen3-seqcls` → `http://10.251.50.54:8012/v1`,
|
||
mode rerank. Added via LiteLLM **`/model/new`** (DB-backed, `store_model_in_db:true`)
|
||
— **no gateway restart** (respects the "nothing else offline" line). Verified 200
|
||
through the gateway.
|
||
- **Metrics:** VRAM ≈ **3.5 GB** (GPU1 free 14167→10616 MiB). Latency (20-doc pool,
|
||
~1500-char docs, shared GPU1): single p50 **87 ms**; 8-concurrent p50 **140 ms**,
|
||
~**55 req/s**.
|
||
- **Correctness (3-probe smoke, not the grid):** tracks the incumbent within noise →
|
||
early signal the failure is the **training prior, not the inference head**.
|
||
|
||
**Autonomous decisions this step (reversible):**
|
||
- D1 — port :8012, GPU1, util 0.03 to mirror the incumbent's exact footprint (clean control).
|
||
- D2 — standalone `docker run` (not compose) so bench arms are throwaway; no canonical churn to revert.
|
||
- D3 — gateway wired via `/model/new` (runtime, DB-persisted) rather than config-edit + restart.
|
||
- D4 — did NOT down any RP seat (A2 is 0.6B / 3.5 GB; no VRAM pressure).
|
||
|
||
**Cleanup for A2 (run at end / on reversal):**
|
||
- `ssh infra-ops@10.251.50.54 'sudo docker stop vllm-rerank-a2 && sudo docker rm vllm-rerank-a2'`
|
||
- Delete gateway alias: `POST /model/delete {"id": <model_id>}` (id via `/model/info?model_name=reranker-a2-qwen3-seqcls`), infra-ops admin key. (DB-persisted, so it survives a restart — must be explicitly deleted.)
|
||
- No weights deleted (red line); HF cache under /tank/aimodels/huggingface retains the 0.6B-seq-cls download.
|
||
|
||
**Ports reserved for the bench:** :8012 (A2), :8013 (A3), :8014 (A4), :8019 (A5).
|
||
|
||
### 2026-08-06 — A3 + A4 pre-staged (Brokkr said pre-stage in parallel, hold A5)
|
||
|
||
- **A3** `vllm-rerank-a3` — fv-ml1 GPU1 :8013, `BAAI/bge-reranker-v2-m3`
|
||
(XLMRobertaForSequenceClassification), same run pattern, util 0.03. VRAM ≈ **2.3 GB**.
|
||
Latency (20-doc, ~1500-char, shared GPU1): single p50 **105 ms**; 8-conc p50 214 ms, ~34 req/s.
|
||
Gateway alias `reranker-a3-bge-v2-m3` via /model/new (200, verified).
|
||
- **A4** `vllm-rerank-a4` — fv-ml1 GPU1 :8014, `Alibaba-NLP/gte-reranker-modernbert-base`
|
||
(ModernBertForSequenceClassification), util 0.02. VRAM ≈ **1.4 GB**. Latency: single
|
||
p50 **102 ms**; 8-conc p50 153 ms, ~51 req/s. Gateway alias `reranker-a4-gte-modernbert`
|
||
via /model/new (200, verified).
|
||
- **Smoke (2-doc, NOT authoritative):** BOTH decisively rank the bare-name Hobgoblin doc top
|
||
(A3 0.999, A4 0.982) where A2/incumbent FAIL (0.33). Cross-encoder / different-lineage.
|
||
Caveat: Brokkr warned isolated tests overstate; his 20-pool grid is the real call.
|
||
- **GPU1 state:** A2+A3+A4 ≈ 7.2 GB resident; GPU1 free ≈ **6.9 GB**. No RP seat downed.
|
||
If A5 (4B, ~4–5 GB) is greenlit: fits GPU1 tight or GPU0 (~9 GB free) — no RP-seat downing expected.
|
||
|
||
**Cleanup for A3/A4 (same pattern as A2):** `docker stop/rm vllm-rerank-a3 vllm-rerank-a4`
|
||
on fv-ml1; `/model/delete` the two aliases (DB-persisted); weights retained in HF cache.
|
||
|
||
### 2026-08-06 — A2 verdict (Brokkr full grid): training-prior confirmed
|
||
|
||
- **A2 ≡ incumbent, statistically indistinguishable** (identical gold-rank on 7/8 probes,
|
||
max 1-rank divergence; n=14: A2 7/14 top-10 @ mean rank 9.71 = incumbent to 2 dp;
|
||
no-reranker 13/14 @ mean 2.79). The seq-cls head changes nothing → the fault is a
|
||
**training prior in the weights**, not the scoring head. (Smoke called it pre-grid.)
|
||
- **A5 (Qwen3-Reranker-4B): HELD INDEFINITELY, not staged** per Brokkr — A2 voided its
|
||
rationale (scale can't fix a prior the head wasn't causing). *Decision: the one expensive
|
||
bring-up is avoided unless Brokkr formally revisits.*
|
||
- **A3/A4:** proceed — already live for Brokkr's grid; now a training-corpus test (BGE vs
|
||
GTE vs Qwen data), lower EV, cost sunk. Awaiting his scoring.
|
||
- **Likely endgame:** NO model swap. Recommendation trending to a **policy change** —
|
||
wing-scoped rerank bypass or `rrf:60` fusion — landing as Worldtree core code behind
|
||
config, NOT a new serving commitment. Would FREE a GPU seat, not allocate one; prod
|
||
`reranker` eventually retired for the fiction path (never silently repointed; Brokkr
|
||
flags before anything touches the prod alias). *Plan: if confirmed, tear the whole bench
|
||
down (A2/A3/A4 containers + 3 aliases) and hand back the GPU.*
|
||
|
||
### 2026-08-06 — FINAL verdict (Brokkr R43.1): A3 wins; cutover HELD for operator
|
||
|
||
- **Winner: A3 = `BAAI/bge-reranker-v2-m3`.** Write-up:
|
||
`research/R43-fleet-reranker-selection/RECOMMENDATION.md` (Brokkr repo, tag R43.1).
|
||
- **The incumbent harms the fleet, not just fiction.** n=90 over main + knowledge_base:
|
||
| arm | top-10 | mean rank | harmed vs no-rerank |
|
||
|---|:--:|:--:|:--:|
|
||
| A0 no-reranker | 89/90 | 0.54 | — |
|
||
| A1 incumbent | 56/90 | 7.78 | **80/90 (worst −19)** |
|
||
| **A3 bge-v2-m3** | 90/90 | 0.19 | 7/90 (worst −3) |
|
||
| A4 gte-modernbert | 90/90 | 0.08 | 1/90 (worst −1) |
|
||
- **A3 over A4:** A4 edges A3 on main/kb + is smaller/faster, BUT A4 is **English-only
|
||
(ModernBERT)** → silent degradation on non-English fleet content; A3 is **multilingual
|
||
(XLM-R)** and decisively better on the bare-name regime that started this. A3 also ~1.2 GB
|
||
*cheaper* than the incumbent. A4 kept as documented throughput fallback.
|
||
- **CUTOVER = OPERATOR DECISION (pending).** Brokkr drafted then PULLED the repoint: a
|
||
fleet-wide alias change affecting consumers he doesn't own shouldn't ship on a relayed
|
||
blanket auth while the operator is away. → Surfaced to Vuong. Proceeding-on-Brokkr's-rec
|
||
now literally = HOLD. **Nothing torn down (incl. A2); prod `reranker` :8002 stays incumbent.**
|
||
- **Cutover conditions (when operator says yes):** repoint gateway `reranker` alias
|
||
incumbent→A3; keep incumbent :8002 warm (rollback = one alias edit); keep
|
||
`reranker-a3-bge-v2-m3` as its own distinct alias; keep A4 up as fallback; **announce the
|
||
boundary timestamp on-bus** (worldtree probe re-run + Brokkr v13 gate render need it).
|
||
- **Flag (worldtree-side, not infra):** `rerank_hybrid_floor` should be **dropped, not
|
||
re-tuned** — it compensates for the scorer being replaced. No serving work to stage for it.
|
||
|
||
### 2026-08-06 — CUTOVER SHIPPED (operator authorized directly + to Brokkr)
|
||
|
||
- **Operator authorized** the fleet repoint (to me: "go a/3"; to Brokkr directly: "go ahead
|
||
with the cutover") and explicitly cleared the litellm restart blip ("authorized to blip litellm").
|
||
- **BOUNDARY: 2026-08-06T17:37:48Z.** Gateway `reranker` alias now resolves 100% to A3
|
||
(`BAAI/bge-reranker-v2-m3` @ :8013). Verified through gateway: "Hobgoblin Pus" relevant
|
||
doc top @ 0.9989 (BGE signature; incumbent was ~0.33).
|
||
- **Mechanism:** `reranker` was config-defined (not DB), config mounted `:ro`, no hot-reload →
|
||
edited `/opt/docker/conf/litellm/config.yaml` reranker block (block-scoped script, asserted
|
||
1+1 change) + `docker restart litellm`. **Blip was ~52s** (litellm reloads all 28 models on
|
||
boot), not the ~15s estimated — reported honestly to operator + Brokkr + worldtree-dev.
|
||
- **`qwen3-reranker` alias LEFT UNTOUCHED** → incumbent still served at :8002 (rollback path;
|
||
also avoids a false alias — the Qwen name still names the Qwen model).
|
||
- **Canonical synced:** `stacks/litellm/conf/config.yaml` reranker block updated to match live.
|
||
(Live config had pre-existing drift from canonical — only the reranker block was reconciled.)
|
||
|
||
**ROLLBACK (one-liner, ~1 min):** revert the `reranker` block in
|
||
`/opt/docker/conf/litellm/config.yaml` to `model: hosted_vllm/Qwen/Qwen3-Reranker-0.6B` +
|
||
`api_base: …:8002/v1` (backup at `config.yaml.bak-pre-rerank-cutover-*`), then
|
||
`sudo docker restart litellm`. Incumbent backend (`vllm-rerank` :8002) is up and untouched.
|
||
|
||
### OPEN / cleanup owed at process end (operator reverses/approves)
|
||
- **A5** never staged (Brokkr cancelled) — nothing to clean.
|
||
- **A2 (`vllm-rerank-a2` :8012)** + alias `reranker-a2-qwen3-seqcls` — bench-only, tear down when
|
||
Brokkr signals the bake-off is closed (`docker stop/rm` + `/model/delete`).
|
||
- **A4 (`vllm-rerank-a4` :8014)** + alias — KEEP for now (Brokkr's documented throughput fallback).
|
||
- **A3 (`vllm-rerank-a3` :8013)** — now PRODUCTION (backs the `reranker` alias). Hardened
|
||
2026-08-06: `docker update --restart unless-stopped` (survives fv-ml1 reboot, no recreate).
|
||
A4 given the same. **Remaining follow-up (not urgent): promote A3 from throwaway `docker run`
|
||
to a canonical compose service** (`stacks/vllm/`) for config-managed consistency — a recreate,
|
||
so do it in a window since it briefly drops `reranker`.
|
||
- **Incumbent (`vllm-rerank` :8002)** — keep up as rollback until Brokkr/worldtree close the
|
||
post-cutover watch; retire (not delete) only on explicit sign-off.
|
||
- **`rerank_hybrid_floor`** — Brokkr routing to worldtree-dev directly (drop, don't re-tune).
|
||
|
||
### 2026-08-06 — VERIFIED + A2 torn down + throughput characterized
|
||
|
||
- **Brokkr independent verify: CUTOVER VERIFIED** — prod `reranker` == `reranker-a3-bge-v2-m3`
|
||
at maxdiff 0.000000 (5 samples, spread 0.000009), single backend, no split routing.
|
||
- **R42 v13 acceptance gate PASSES** — anchors_flip 4/4, no_regression 8/8, no_distractor_rise
|
||
TRUE, zero aborts. **First PASS in R42 history after 4 failed verdicts.** Production main+kb:
|
||
56/90 → 90/90 top-10; evictions 33 → 0.
|
||
- **A2 torn down** (Brokkr signalled done): gateway alias `reranker-a2-qwen3-seqcls` deleted
|
||
(/model/delete 200) + container removed. ~3.5 GB freed on GPU1. Remaining: `vllm-rerank`
|
||
(incumbent, rollback), `vllm-rerank-a3` (prod), `vllm-rerank-a4` (fallback).
|
||
- **Throughput characterized (the one open risk):** A3 caps ~34 req/s — flat from 8→16
|
||
concurrent while latency climbs gracefully (p50 214→332→456 ms; p99 525 ms @16-conc). It
|
||
QUEUES, doesn't cliff. ~40% below the incumbent's ~55 req/s. Likely fine for fleet rerank
|
||
QPS (internal, per-search), but if real p99/queue-depth bites: levers are (a) swap to A4
|
||
(~51 req/s, but English-only), (b) raise A3 `--gpu-memory-utilization` for bigger batching
|
||
(recreate = brief blip), (c) run a 2nd A3 replica load-balanced behind `reranker` (~2× tput,
|
||
identical replicas so no split-measurement issue now the bake-off is closed). Brokkr will
|
||
re-run the grid against A4 if it bites — no intuition swaps.
|