feat(embed-rerank): TEI is the fleet embed/rerank engine; esh-ml1 sole backend; retire fv-ml1 seats

Prime, 2026-09-25: TEI serves embedding + reranking for esh-ml1 and the fleet
from now on; fv-ml1 retires both once esh-ml1 is up.

- stacks/embed-rerank: vLLM -> TEI 1.9.4 (89- Ada build), same ports
  8001/8013, fail-closed truncation (--auto-truncate false; embed
  --max-batch-tokens 32768).
- litellm: qwen3-embedding -> esh-ml1 only (hosted_vllm/, unchanged address);
  reranker -> huggingface/ provider at :8013 (hosted_vllm/ 422s on TEI's
  `texts` body). DB alias reranker-a3-bge-v2-m3 patched to the same target.
- Verified via the gateway against the retiring fv-ml1 seats: embed cosine
  median 0.999927 (n=203); rerank top-1/top-3 29/30.
- stacks/vllm: vllm-embed and vllm-rerank-a3 removed (containers retired on
  fv-ml1, GPU 1 freed ~6.1 GB); reward + coder unchanged.
- Bake-off record moved to docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md;
  CLAUDE.md gains the TEI convention.
This commit is contained in:
vh
2026-09-25 08:30:53 -07:00
parent 49b4bf0177
commit 7bdac80878
17 changed files with 271 additions and 418 deletions
+8 -18
View File
@@ -1,29 +1,19 @@
# embed-rerank tunables (esh-ml1). Copy to `.env` on the server.
#
# Everything model-shaped here MUST match fv-ml1's stacks/vllm .env: the same
# models, the same vLLM version, the same max-model-len. A drift in the
# embedding model or its version makes this seat's vectors incompatible with
# every index built against fv-ml1's.
# embed-rerank tunables (esh-ml1, TEI). Copy to `.env` on the server.
# PINNED to fv-ml1's version. Bump both sites together.
VLLM_VERSION=v0.24.0
# TEI image tag. `89-` = the Ada Lovelace (sm_89) build; a different GPU
# generation needs a different prefix (see the TEI README's image table).
# PINNED to an exact release.
TEI_TAG=89-1.9.4
# Same host ports as fv-ml1 (container listens on 8000).
# Kept from the vLLM era so the embedding gateway entry did not have to change.
EMBED_PORT=8001
RERANK_PORT=8013
# ⚠ Changing EMBED_MODEL invalidates every index built on it.
EMBED_MODEL=Qwen/Qwen3-Embedding-0.6B
RERANK_MODEL=BAAI/bge-reranker-v2-m3
# Fractions of the RTX 2000E Ada's 16,380 MiB. fv-ml1 runs 0.03 of a 96 GB
# card (~2.9 GB each); 0.20 here is ~3.2 GB each — the same budget plus a
# little, leaving ~9.5 GB free.
EMBED_GPU_MEM_UTIL=0.20
RERANK_GPU_MEM_UTIL=0.20
EMBED_MAX_MODEL_LEN=8192
RERANK_MAX_MODEL_LEN=8192
MAX_CLIENT_BATCH_SIZE=128
# fv-ml1's seats run with no API key (LiteLLM fronts them); match that.
API_KEY=
# Both models are public; no token needed.
HF_TOKEN=
+34 -15
View File
@@ -1,18 +1,40 @@
# embed-rerank
The fleet's embedding + reranking models on **esh-ml1** (CT 110 on esh-pve,
RTX 2000E Ada). The second site for `qwen3-embedding` and `reranker`; fv-ml1's
[`vllm`](../vllm/) stack is the first.
**The fleet's embedding + reranking service**, on **esh-ml1** (CT 110 on esh-pve,
RTX 2000E Ada), served by **Hugging Face Text Embeddings Inference (TEI)**.
Since 2026-09-25 it is the only backend behind the gateway names
`qwen3-embedding`, `reranker` and `reranker-a3-bge-v2-m3`.
| container | model | port |
|---|---|---|
| `vllm-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 |
| `vllm-rerank-bge` | `BAAI/bge-reranker-v2-m3` | 8013 |
**TEI is the fleet's embed/rerank engine** (Prime, 2026-09-25). New embedding or
reranking seats go on TEI, not vLLM. Why, and the measurements behind it:
[`docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md`](../../docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md).
**Keep it in lockstep with `stacks/vllm`:** same models, same `VLLM_VERSION`,
same `--max-model-len`. The LiteLLM groups treat the two sites as one model,
so any drift in the embedding model or its version silently mixes
incompatible vectors into consumers' indexes.
| container | model | port | endpoint |
|---|---|---|---|
| `tei-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 | `/v1/embeddings` (OpenAI), `/embed` |
| `tei-rerank` | `BAAI/bge-reranker-v2-m3` | 8013 | `/rerank` — body `{"query", "texts"}` |
## Gateway wiring (stacks/litellm/conf/config.yaml)
- `qwen3-embedding` → `hosted_vllm/Qwen/Qwen3-Embedding-0.6B`, `api_base: http://10.0.50.80:8001/v1`.
- `reranker` → **`huggingface/`**`BAAI/bge-reranker-v2-m3`, `api_base: http://10.0.50.80:8013` (**no `/v1`**).
⚠ `hosted_vllm/` sends `documents` and TEI answers **422 "missing field texts"**.
- `reranker-a3-bge-v2-m3` is a DB-only alias (not in config.yaml) with the same target.
## ⚠ Invariants
- **Changing `EMBED_MODEL` invalidates every index built on it** (Worldtree,
nevermore, Open WebUI). An engine change is allowed only with a parity
measurement against the current vectors. The TEI switch measured cosine median
0.999925 and old-index retrieval overlap 0.988.
- **Truncation is fail-closed** (`--auto-truncate false`). TEI's default silently
embeds a prefix of an over-length input and returns 200. Now: embed rejects
more than 32,768 tokens and rerank more than 8,192, both with 422.
- **The image tag is GPU-generation specific**: `89-` is Ada. A different card
needs a different prefix (TEI README image table).
- nevermore thresholds rerank scores at 0.3. TEI scores differ from the old vLLM
seat by up to 0.019, which produced 0 flips in 2,000 scores. Recheck
thresholding consumers after any engine or version change.
## Deploy
@@ -30,8 +52,5 @@ Host prerequisites (driver, LXC, docker, toolkit) are in
curl -s http://10.0.50.80:8001/v1/embeddings -H 'content-type: application/json' \
-d '{"model":"Qwen/Qwen3-Embedding-0.6B","input":"hello"}' | jq '.data[0].embedding | length' # 1024
curl -s http://10.0.50.80:8013/rerank -H 'content-type: application/json' \
-d '{"model":"BAAI/bge-reranker-v2-m3","query":"cat","documents":["a cat","a car"]}' | jq '.results[0]'
-d '{"query":"cat","texts":["a cat","a car"]}' # index 0 scores ~0.9986
```
Parity against fv-ml1 is recorded in the host README; re-measure after any
version bump on either side.
+66 -73
View File
@@ -1,63 +1,66 @@
# embed-rerank — the fleet's embedding + reranking models, served locally at ESH
# on esh-ml1 (CT 110 on esh-pve, RTX 2000E Ada, 16 GB).
# embed-rerank — THE fleet's embedding + reranking service, on esh-ml1 (CT 110 on
# esh-pve, RTX 2000E Ada, 16 GB). Served by Hugging Face Text Embeddings
# Inference (TEI).
#
# WHY: Prime's decision 2026-09-24 — the Ada card offloads embedding and
# reranking so ESH consumers (Open WebUI RAG, Paperless) are not tied to one
# seat on fv-ml1. It serves the SAME models as fv-ml1's `vllm` stack, with the
# SAME vLLM version and flags, because embedding vectors are model-specific:
# a different embedding model here would silently poison every index built
# against fv-ml1's. Mirror stacks/vllm/compose.yaml when that changes.
# Prime, 2026-09-25: "TEI is embed/reranker server for esh-ml1 and the FLEET in
# general, in future." It replaced vLLM here the same day, after a side-by-side
# bake-off on this card (docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md). TEI
# gives the same vectors as the old vLLM seats (no re-embedding), but it is
# ~1.3x slower on bulk work on this card. It is much lighter (2.6 GB VRAM for
# both, 8 GB image, ~4 s restart). fv-ml1's vLLM embed/rerank seats were
# retired after this went live.
#
# vllm-embed Qwen/Qwen3-Embedding-0.6B → /v1/embeddings :8001
# vllm-rerank-bge BAAI/bge-reranker-v2-m3 → /rerank, /score :8013
# tei-embed Qwen/Qwen3-Embedding-0.6B → /v1/embeddings (OpenAI), /embed :8001
# tei-rerank BAAI/bge-reranker-v2-m3 → /rerank (body: query + texts) :8013
#
# Ports match fv-ml1 on purpose, so a LiteLLM deployment for either site
# differs only in the host part of api_base.
# Ports kept from the vLLM era (and fv-ml1), so the embedding gateway entry did
# not change address. ⚠ The RERANK gateway entry must use LiteLLM's
# `huggingface/` provider: `hosted_vllm/` sends `documents` and TEI answers 422
# "missing field texts".
#
# `vllm-rerank-bge`, not fv-ml1's `vllm-rerank-a3`: that name carries R43
# bake-off provenance for THAT container; and plain `vllm-rerank` was the
# retired Qwen3-Reranker that measured harmful. Name the model instead.
# ⚠ Embedding vectors are model-specific. Never change EMBED_MODEL without a
# re-embedding plan for every index built on it (Worldtree, nevermore, Open WebUI).
#
# NO `tnet`/traefik-net: esh-ml1 runs no traefik, and an external network
# that does not exist would stop the stack from starting. Consumers reach the
# published ports directly.
# FAIL-CLOSED truncation (--auto-truncate false). TEI's default silently
# embedded the first 16,384 tokens of a ~40k-token input and returned 200.
# Turning it off requires --max-batch-tokens >= the model's max input (32,768
# for Qwen3-Embedding), or TEI refuses to start.
#
# Host setup (driver, LXC, docker, toolkit): playbooks/esh-pve-nvidia-host.yaml
# then playbooks/esh-ml1-lxc.yaml. Tunables live in .env — edit that, not this.
# NO `tnet`/traefik-net: esh-ml1 runs no traefik; consumers reach the published
# ports, and in practice only the LiteLLM gateway does (verified from the seats'
# logs 2026-09-25: every request matched a gateway spend-log row).
#
# Host setup: playbooks/esh-pve-nvidia-host.yaml, then playbooks/esh-ml1-lxc.yaml.
# Tunables live in .env.
name: embed-rerank
services:
vllm-embed:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-embed
tei-embed:
image: ghcr.io/huggingface/text-embeddings-inference:${TEI_TAG}
container_name: tei-embed
restart: unless-stopped
ipc: host
ports:
- "${EMBED_PORT}:8000"
- "${EMBED_PORT}:80"
volumes:
- /opt/aimodels/huggingface:/hfcache
- /opt/aimodels/tei-cache:/data
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
- HF_TOKEN=${HF_TOKEN:-}
command:
- --model-id
- ${EMBED_MODEL}
- --served-model-name
- ${EMBED_MODEL}
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${EMBED_GPU_MEM_UTIL}
- --max-model-len
- ${EMBED_MAX_MODEL_LEN}
# TEI on CUDA is float16-only; parity vs the bf16 vLLM vectors was measured.
- --dtype
- auto
- float16
# Default 32; vLLM had no cap and callers batch 64.
- --max-client-batch-size
- "${MAX_CLIENT_BATCH_SIZE}"
- --auto-truncate
- "false"
- --max-batch-tokens
- "32768"
deploy:
resources:
reservations:
@@ -66,49 +69,39 @@ services:
device_ids: ["0"]
capabilities: [gpu]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
test: ["CMD", "curl", "-fsS", "http://localhost:80/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
start_period: 60s
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Embed (esh-ml1)
- homepage.name=Embed — Qwen3 0.6B (TEI, esh-ml1)
- homepage.icon=mdi-vector-arrange-below
- homepage.description=Qwen3 Embedding 0.6B via vLLM (esh-ml1, RTX 2000E Ada)
- homepage.description=Fleet embeddings (qwen3-embedding) via TEI on esh-ml1
- homepage.href=http://10.0.50.80:${EMBED_PORT}/docs
vllm-rerank-bge:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-rerank-bge
tei-rerank:
image: ghcr.io/huggingface/text-embeddings-inference:${TEI_TAG}
container_name: tei-rerank
restart: unless-stopped
ipc: host
ports:
- "${RERANK_PORT}:8000"
- "${RERANK_PORT}:80"
volumes:
- /opt/aimodels/huggingface:/hfcache
- /opt/aimodels/tei-cache:/data
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
- HF_TOKEN=${HF_TOKEN:-}
command:
- --model-id
- ${RERANK_MODEL}
- --served-model-name
- ${RERANK_MODEL}
- --runner
- pooling
# bge-reranker-v2-m3 is natively a cross-encoder; no --hf-overrides.
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${RERANK_GPU_MEM_UTIL}
- --max-model-len
- ${RERANK_MAX_MODEL_LEN}
- --dtype
- auto
- float16
- --max-client-batch-size
- "${MAX_CLIENT_BATCH_SIZE}"
# Fail-closed; bge-reranker-v2-m3's max input (8,192) fits the default
# max-batch-tokens (16,384).
- --auto-truncate
- "false"
deploy:
resources:
reservations:
@@ -117,14 +110,14 @@ services:
device_ids: ["0"]
capabilities: [gpu]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
test: ["CMD", "curl", "-fsS", "http://localhost:80/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
start_period: 60s
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Rerank bge-v2-m3 (esh-ml1)
- homepage.name=Rerank — bge-v2-m3 (TEI, esh-ml1)
- homepage.icon=mdi-sort-variant
- homepage.description=bge-reranker-v2-m3 via vLLM (esh-ml1, RTX 2000E Ada)
- homepage.description=Fleet reranker (reranker) via TEI on esh-ml1
- homepage.href=http://10.0.50.80:${RERANK_PORT}/docs
+21 -37
View File
@@ -7,8 +7,8 @@
#
# Model-name → upstream mapping:
# phi4-mini → vLLM :8004 (generative chat)
# qwen3-embedding → vLLM :8001 (/v1/embeddings)
# qwen3-reranker → vLLM :8002 (/rerank)
# qwen3-embedding → TEI on esh-ml1 :8001 (/v1/embeddings)
# reranker → TEI on esh-ml1 :8013 (/rerank, huggingface/ provider)
# * (wildcard) → llama-swap :9292 (the swappable generative zoo)
#
# The wildcard fronts llama-swap so its whole model zoo logs through the
@@ -443,31 +443,21 @@ model_list:
# explicitly. Operator ruling 2026-08-23. ---
# --- Qwen3 embeddings ---
# TWO deployments, one name, ordered: fv-ml1 serves (order 1); esh-ml1 takes
# over only when fv-ml1 fails (order 2 — LiteLLM's order-based fallback, v1.97
# router.py). Same model, same vLLM version and flags at both sites, so the
# vectors are interchangeable: measured 2026-09-24, cosine FV-vs-ESH median
# 0.999908 (min 0.999772) over 11 texts, inside the FV-vs-FV self-noise floor
# (median 0.999927, min 0.999791); different-text negative control 0.07–0.38.
# ⚠ Failover cost, measured n=3 each: primary REFUSING → +0.15 s; primary
# HOST DOWN (no ARP) → ~18.7 s on EVERY call (no cooldown kicked in). Slow,
# not broken. ⚠ Never add a deployment here that serves a DIFFERENT embedding
# model — vectors are model-specific and a mixed group corrupts indexes.
# esh-ml1: stacks/embed-rerank, playbooks/esh-ml1-lxc.yaml.
- model_name: qwen3-embedding
litellm_params:
model: hosted_vllm/Qwen/Qwen3-Embedding-0.6B
api_base: http://10.251.50.54:8001/v1
api_key: os.environ/VLLM_API_KEY
order: 1
model_info:
mode: embedding
# ONE backend: esh-ml1 (RTX 2000E Ada), served by Hugging Face TEI — the fleet's
# embed/rerank engine from 2026-09-25 (Prime). fv-ml1's vLLM seat was retired the
# same day. TEI's /v1/embeddings is OpenAI-compatible, so the provider stays
# `hosted_vllm/`. Same model as before; TEI vectors match the old vLLM ones
# (cosine median 0.999925, inside vLLM's own noise; old-index retrieval overlap
# 0.988 vs 0.986 self-noise), so no index needed re-embedding.
# ⚠ Single backend until the second RTX 2000 lands (Prime, 2026-09-25): if
# esh-ml1, esh-pve or the ESH mesh route is down, fleet embeddings are down.
# ⚠ Never point this name at a DIFFERENT embedding model — vectors are
# model-specific. Engine bake-off: docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md.
- model_name: qwen3-embedding
litellm_params:
model: hosted_vllm/Qwen/Qwen3-Embedding-0.6B
api_base: http://10.0.50.80:8001/v1
api_key: os.environ/VLLM_API_KEY
order: 2
model_info:
mode: embedding
@@ -508,23 +498,17 @@ model_list:
model_info:
mode: chat
# reranker → generic capability name for rerank (bge-reranker-v2-m3 since the
# R43 cutover). fv-ml1 order 1, esh-ml1 order 2 — same failover shape as
# qwen3-embedding above. Parity measured 2026-09-24: |score| FV-vs-ESH max
# 0.000145 vs FV-vs-FV floor 0.000181, identical ranking over 11 documents.
# R43 cutover). ONE backend: TEI on esh-ml1 (fv-ml1's vLLM seat retired
# 2026-09-25). ⚠ TEI's /rerank takes `texts`, not `documents`: this MUST use the
# `huggingface/` provider. `hosted_vllm/` gets a 422 "missing field texts".
# api_base has NO /v1. api_key is a placeholder; TEI runs without auth.
# Parity vs the retired vLLM seat: top-1/top-3 identical (129/130 top-1 over
# two runs); scores move up to 0.019, reshuffling only near-zero tail docs.
- model_name: reranker
litellm_params:
model: hosted_vllm/BAAI/bge-reranker-v2-m3
api_base: http://10.251.50.54:8013/v1
api_key: os.environ/VLLM_API_KEY
order: 1
model_info:
mode: rerank
- model_name: reranker
litellm_params:
model: hosted_vllm/BAAI/bge-reranker-v2-m3
api_base: http://10.0.50.80:8013/v1
api_key: os.environ/VLLM_API_KEY
order: 2
model: huggingface/BAAI/bge-reranker-v2-m3
api_base: http://10.0.50.80:8013
api_key: "none"
model_info:
mode: rerank
-117
View File
@@ -1,117 +0,0 @@
# tei-bakeoff — TEI vs vLLM for embedding + reranking on esh-ml1
Prime, 2026-09-25: *"run the TEI bake-off."* Hugging Face **Text Embeddings
Inference 1.9.4** (`89-1.9.4`, the Ada build) serving the same two models as
[`embed-rerank`](../embed-rerank/) (vLLM v0.24.0), side by side on the same RTX
2000E Ada. Not wired into the gateway. This stack either becomes the
embed-rerank stack or is torn down.
## Results (2026-09-25 0744–0805 PT)
### Parity — TEI matches the existing vLLM vectors
Reference = vLLM on fv-ml1, the engine every existing index was built with.
306 texts (6 fixed incl. CJK/code/6k-char + 300 *The Stand* paragraphs);
retrieval = 2,000-paragraph corpus, 50 instruction-format queries.
| embedding check | result | noise floor / control |
|---|---|---|
| cosine TEI vs vLLM-FV | median 0.999925, min 0.999861 | vLLM-FV vs itself: 0.999916 / 0.999796 |
| overlap@10, TEI index + TEI queries | 0.976 | vLLM-FV rerun 0.986; vLLM-ESH 0.972 |
| overlap@10, **TEI queries vs the OLD vLLM index** (migration case) | **0.988** | positive control (MRL-256 dims) 0.648 |
| hit@1 own paragraph | 0.80 | vLLM-FV 0.80 |
| different-text negative control | cosine median 0.31 | — |
TEI is deterministic (TEI vs itself: min 0.999995). **Switching engines does not
require re-embedding existing indexes.**
| rerank check (100 queries × 20 docs, 9 same-chapter distractors) | TEI vs vLLM-FV | vLLM-FV vs itself |
|---|---|---|
| source paragraph ranked #1 | 0.95 (same as vLLM) | 0.95 |
| top-1 agreement | 1.00 | 1.00 |
| top-3 set agreement | 1.00 | 1.00 |
| top-5 exact order | 0.98 | 1.00 |
| order among docs scoring > 0.05 | 1.00 (n=37) | 1.00 |
| full 20-doc order | **0.55** | 0.99 |
| max score difference | **0.019** (p99 0.0034) | 0.0016 |
Every decision that matters agrees; the differences are shuffles among
near-zero-scoring tail documents. ⚠ A consumer that **thresholds** on the rerank
score could see a borderline document flip (scores move up to ~0.02). An
earlier 30-query run with random distractors had one top-1 disagreement (29/30);
across both runs, 129/130.
### Speed — TEI is NOT faster on this card
On-box, 2 interleaved runs × 3 reps each, medians:
| workload | vLLM | TEI |
|---|---|---|
| embed 1 short query, p50 | ~9.1 ms | **~6.9 ms** |
| embed 1 × ~512 tok, p50 | ~23 ms (bimodal 10–25) | ~28 ms |
| embed 1 × ~2k tok, p50 | **~86 ms** | ~108 ms |
| bulk embed, passages/s | **~50** | ~37 |
| whole novel (*The Stand*, 12,814 paragraphs), 64/request, 4 in flight | **31 s** | 40 s |
| rerank 20 docs, p50 | **~169 ms** | ~180 ms |
| rerank 20 docs, req/s at conc 8 | ~5.7 | ~5.6 |
TEI ran its fused `FlashQwen3` path. `--max-batch-tokens 32768` did not change
the novel time (39–40 s), so it stays at the default. **Through the gateway these
gaps mostly vanish**: LiteLLM is the bottleneck there (see
`servers/esh-ml1/README.md`).
### Footprint — TEI is much lighter
| | vLLM (both) | TEI (both) |
|---|---|---|
| VRAM (host `nvidia-smi`, per process) | 3,298 + 1,512 MiB | 1,352 + 1,256 MiB |
| image | 29.9 GB | 8.16 GB |
| warm restart to healthy (n=3) | ~24 s | ~4 s |
| container RAM just after start | 2.5 + 5.1 GiB | 0.7 + 1.7 GiB |
⚠ vLLM's VRAM figure is mostly a **setting** (`--gpu-memory-utilization 0.20`
each). A lower setting would shrink it; that has not been tested.
### External review — Dvalin (Grok research peer), 2026-09-25
**Agree with caveats: adopt TEI for these two seats; footprint is the right
reason.** His external evidence (not re-verified here): a Runpod 2026-09-14
engine comparison shows vLLM ahead of TEI on Qwen3-Embedding bulk throughput
(median ~2.6×, single unreplicated runs, near parity on smaller cards) and TEI
ahead on BERT-family models — consistent with our 1.3× on a 50 W card and the
rerank near-tie. Open TEI 1.9 issue #857 (tokio panic "No backend receiver"
under load, n=1) is a watch item, not a gate. A lower vLLM memory setting
would shrink only the embedder (the reranker at 1,512 MiB is already under
its cap); 0.12 is the only setting he'd expect to both boot and help — untested.
His caveats, and status: fail-closed truncation (**done**, above); recheck score-
threshold consumers (**done**, nevermore, above); two gateway providers must be
written down at cut-over (open).
## Gateway compatibility (LiteLLM v1.97, temp aliases, deleted after)
- Embeddings: **drop-in**. TEI's `/v1/embeddings` works with the existing
`hosted_vllm/` provider.
- Rerank: needs the **`huggingface/` provider** (scores identical to direct
TEI). `hosted_vllm/` fails with 422 (`missing field texts`); TEI's `/rerank`
body differs from vLLM's.
## Behaviour differences to carry into any cut-over
- **Truncation — now fail-closed (`--auto-truncate false`).** Measured with
TEI's default: a ~40k-token input returned **200 with a vector of its first
16,384 tokens**, silently, where vLLM returns 400. With truncation off, TEI
embed rejects > 32,768 tokens and TEI rerank rejects > 8,192 (both 422,
verified). Truncation off requires `--max-batch-tokens` at or above the model's max
input (32,768 for Qwen3-Embedding) or TEI refuses to start; VRAM unchanged.
Remaining difference: TEI embed **accepts** 8,193–32,767 tokens that vLLM
(`--max-model-len 8192`) rejects. It embeds the whole input, which is wider,
not lossy.
- **nevermore thresholds rerank scores** (`NEVERMORE_RERANK_THRESHOLD`, default
0.3, cluster-member filter in `digest.py`). Across the 2,000 bake-off scores,
**0 flips** at 0.3 or 0.5 for TEI vs vLLM — but only 1 score landed within
±0.02 of 0.3 (bge scores sit near 0 or 1), so this cannot exclude rare flips
on borderline headlines. Worst case: a same-story headline clusters
differently. Low stakes.
- TEI caps inputs per request at `--max-client-batch-size` (default 32; set to 128 here).
- TEI on CUDA runs float16 only (vLLM served Qwen3-Embedding in bf16); parity
above says it does not matter here.
-81
View File
@@ -1,81 +0,0 @@
# tei-bakeoff — Hugging Face Text Embeddings Inference (TEI) serving the SAME two
# models as stacks/embed-rerank (vLLM), side by side on esh-ml1, so the two
# engines can be compared on identical hardware. Prime, 2026-09-25: "run the
# TEI bake-off". Not wired into the gateway. Temporary: either this becomes the
# embed-rerank stack or it is torn down.
#
# tei-embed Qwen/Qwen3-Embedding-0.6B → /v1/embeddings, /embed :8081
# tei-rerank BAAI/bge-reranker-v2-m3 → /rerank :8083
#
# 89-* = the Ada Lovelace (sm_89) build. Pinned to an exact release.
# TEI on CUDA offers float16/float32 only (no bfloat16), while vLLM serves
# Qwen3-Embedding in bf16 — the parity measurement is what decides whether
# that matters.
# (--max-batch-tokens 32768 was tried on tei-embed: whole-novel time unchanged,
# 39.2–40.4 s vs 39.9–40.3 s, so it stays at the default.)
# --max-client-batch-size 128: the default is 32; vLLM has no such cap and the
# bake-off sends 64 per request.
# Separate HF cache from vLLM's, so neither engine can disturb the other's files.
name: tei-bakeoff
services:
tei-embed:
image: ghcr.io/huggingface/text-embeddings-inference:89-1.9.4
container_name: tei-embed
restart: unless-stopped
ports:
- "8081:80"
volumes:
- /opt/aimodels/tei-cache:/data
command:
- --model-id
- Qwen/Qwen3-Embedding-0.6B
- --served-model-name
- Qwen/Qwen3-Embedding-0.6B
- --dtype
- float16
- --max-client-batch-size
- "128"
# FAIL-CLOSED on over-length input (Dvalin review 2026-09-25; measured:
# TEI's default silently cut a ~40k-token input to 16,384 and returned
# 200, where vLLM returns 400). Truncation off requires max-batch-tokens
# >= the model's max input (Qwen3-Embedding: 32,768) or TEI refuses to start.
- --auto-truncate
- "false"
- --max-batch-tokens
- "32768"
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"]
capabilities: [gpu]
tei-rerank:
image: ghcr.io/huggingface/text-embeddings-inference:89-1.9.4
container_name: tei-rerank
restart: unless-stopped
ports:
- "8083:80"
volumes:
- /opt/aimodels/tei-cache:/data
command:
- --model-id
- BAAI/bge-reranker-v2-m3
- --dtype
- float16
- --max-client-batch-size
- "128"
# Fail-closed, as tei-embed. bge-reranker-v2-m3 max input (8,192) fits the
# default max-batch-tokens (16,384).
- --auto-truncate
- "false"
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"]
capabilities: [gpu]
+3 -15
View File
@@ -12,23 +12,15 @@
VLLM_VERSION=v0.24.0
# Host ports (container always listens on 8000 internally)
EMBED_PORT=8001
# 8013 — the reranker moved here 2026-08-20 when bge-v2-m3 (the R43 winner, which
# had been running as a throwaway `docker run` on this port) was promoted into
# this stack and the Qwen incumbent on :8002 was retired.
RERANK_PORT=8013
# (EMBED_PORT 8001 / RERANK_PORT 8013 retired 2026-09-25 with their seats —
# embedding + reranking moved to TEI on esh-ml1, stacks/embed-rerank.)
REWARD_PORT=8003
# GPU assignment — all services share this GPU
# (ana-ml2 has 0 and 1; default 1 keeps 0 free for heavy LLM work in llama-swap)
GPU_ID=1
# Models — reference by full repo name in API requests
EMBED_MODEL=Qwen/Qwen3-Embedding-0.6B
# bge-reranker-v2-m3 — the R43 bake-off winner, replacing Qwen3-Reranker-0.6B
# (measured HARMING 80/90 fleet queries). Multilingual cross-encoder; needs no
# --hf-overrides, unlike the Qwen reranker it displaced.
RERANK_MODEL=BAAI/bge-reranker-v2-m3
# Models
# Skywork is a local-path AWQ output, not from HF Hub. Bind-mounted into the
# reward container at /local-models — see compose.yaml. No env var here for
# the model path itself since it's hard-coded in the compose command.
@@ -62,16 +54,12 @@ RERANK_MODEL=BAAI/bge-reranker-v2-m3
# each at 0.05 (mostly util-reservation waste); 0.03 (~3.6 GB) fits weights +
# CUDA context with room, freeing ~4 GB back to granite. Recreate them ONE AT A
# TIME — concurrent recreate races the memory-profiling assertion.
EMBED_GPU_MEM_UTIL=0.03
RERANK_GPU_MEM_UTIL=0.03
REWARD_GPU_MEM_UTIL=0.10
# Context length caps — lower these if VRAM is tight.
# Qwen3-Embedding supports up to 32k; reranker up to 32k.
# Skywork capped at 16k server-side as defense-in-depth; JudgeClient also
# enforces the cap at dispatch time per spec.
EMBED_MAX_MODEL_LEN=8192
RERANK_MAX_MODEL_LEN=8192
REWARD_MAX_MODEL_LEN=16384
# Optional API key — leave blank for no auth (fine on the internal network).
+6
View File
@@ -1,5 +1,11 @@
# vllm
> ⚠ **2026-09-25: embedding + reranking LEFT this stack.** `vllm-embed` and
> `vllm-rerank-a3` were retired; the fleet's embed/rerank now runs on **TEI on
> esh-ml1** (`stacks/embed-rerank`), and TEI is the fleet engine for those from
> now on (Prime). What remains here is `vllm-reward` and `vllm-coder` on fv-ml1.
> The history below predates that.
Multi-service vLLM stack on ana-ml2. Started as Qwen3 embedding + rerank
(replacing the unmaintained Infinity stack); generalized to host any vLLM
model on the box, currently three services:
+10 -145
View File
@@ -1,162 +1,27 @@
# vLLM — Qwen3 Embedding + Reranker + Skywork Reward-V2 classifier.
# vLLM — utility seats on fv-ml1: Skywork Reward-V2 classifier + the coder FIM seat.
#
# Originally created to replace the unmaintained Infinity stack (embed +
# rerank); generalized 2026-05-13 to host any vLLM-served model on fv-ml1,
# starting with the Skywork-Reward-V2-Llama-3.1-8B reward classifier
# (AWQ-quantized locally, served from /tank/aimodels/llm/).
#
# vLLM runs one model per process, so this stack brings up three containers
# sharing a single GPU:
#
# vllm-embed — Qwen3-Embedding served as an OpenAI /v1/embeddings server
# vllm-rerank — Qwen3-Reranker served as a /rerank + /score server
# vllm-reward — Skywork-Reward-V2-Llama-3.1-8B-AWQ served as a /classify scorer
#
# The reranker is a causal-LM checkpoint; --hf-overrides re-maps it to
# Qwen3ForSequenceClassification so vLLM's reranking endpoints work and the
# model only emits two class logits (no/yes) instead of the full 151k vocab.
# ⚠ Embedding + reranking LEFT this stack 2026-09-25 (Prime): they now run on TEI
# on esh-ml1 (stacks/embed-rerank), and TEI is the fleet's embed/rerank engine
# from now on. Do not re-add them here.
#
# All tunables live in .env — edit that, not this file.
#
# Pre-download models to avoid first-run delay:
# scripts/elway fv-ml1 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Embedding-0.6B
# scripts/elway fv-ml1 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Reranker-0.6B
#
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on fv-ml1 and is
# bind-mounted into the reward service at /local-models. Not from HF Hub.
services:
vllm-embed:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-embed
restart: unless-stopped
ipc: host
ports:
- "${EMBED_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${EMBED_MODEL}
- --served-model-name
- ${EMBED_MODEL}
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${EMBED_GPU_MEM_UTIL}
- --max-model-len
- ${EMBED_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Embed (Qwen3)
- homepage.icon=mdi-vector-arrange-below
- homepage.description=Qwen3 Embedding via vLLM (fv-ml1)
- homepage.href=http://10.251.50.54:${EMBED_PORT}/docs
# THE fleet reranker. Backs the LiteLLM `reranker` alias, which is what every
# consumer should name — never the model, never a bake-off arm name.
#
# It won the Brokkr R43 bake-off (docs/pfi/reranker-selection-ledger.md) and
# replaced Qwen3-Reranker-0.6B, which was measured HARMING 80/90 fleet queries
# (no-reranker beat it 89/90 vs 56/90). The R42 v13 acceptance gate went
# 56/90 -> 90/90 on the cutover, its first-ever PASS.
#
# Promoted from a throwaway `docker run` to this service 2026-08-20 (the
# ledger's own open follow-up). It carried the bake-off's arm name from the
# start and KEEPS it: the ledger, persistent-memory and the R43 record all say
# `vllm-rerank-a3`, and renaming for tidiness would orphan every one of those
# references. The name carries its provenance.
#
# Retired alongside this promotion: `vllm-rerank` (Qwen3-Reranker-0.6B, :8002 —
# the rollback path, kept warm 13 days) and `vllm-rerank-a4`
# (gte-reranker-modernbert, :8014 — a documented throughput fallback that was
# never given a gateway alias, so it was unreachable the whole time).
vllm-rerank-a3:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-rerank-a3
restart: unless-stopped
ipc: host
ports:
- "${RERANK_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${RERANK_MODEL}
- --served-model-name
- ${RERANK_MODEL}
- --runner
- pooling
# No --hf-overrides here. The Qwen reranker needed one to be coerced into a
# sequence-classification head; bge-reranker-v2-m3 is natively a
# cross-encoder and vLLM resolves it directly.
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${RERANK_GPU_MEM_UTIL}
- --max-model-len
- ${RERANK_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Rerank (bge-v2-m3)
- homepage.icon=mdi-sort-variant
- homepage.description=BAAI bge-reranker-v2-m3 — the fleet reranker, backs the `reranker` alias (fv-ml1)
- homepage.href=http://10.251.50.54:${RERANK_PORT}/docs
# vllm-embed (Qwen3-Embedding-0.6B, :8001) and vllm-rerank-a3
# (bge-reranker-v2-m3, :8013) were RETIRED 2026-09-25 (Prime): embedding and
# reranking moved to TEI on esh-ml1 (stacks/embed-rerank), and TEI is now the
# fleet's embed/rerank engine. The gateway names `qwen3-embedding`, `reranker`
# and `reranker-a3-bge-v2-m3` did not change. Their definitions are in git
# history before this commit if a rollback is ever needed.
vllm-reward:
image: vllm/vllm-openai:${VLLM_VERSION}