feat(tei-bakeoff): TEI 1.9.4 vs vLLM on esh-ml1 — parity holds, not faster, much lighter

This commit is contained in:
vh
2026-09-25 08:05:17 -07:00
parent a1885718e9
commit 582b150131
4 changed files with 178 additions and 0 deletions
@@ -67,3 +67,14 @@ faster because it is network-closer). Bulk: ESH 3–10× slower (bulk embed ~50
~403 passages/s; rerank-20 ~170 vs ~51 ms p50, ~5.7 vs ~55 req/s at conc 8).
Noise floor 3–5% on throughput. ⇒ failover role confirmed; load-sharing would
roughly double rerank latency for half the calls.
**TEI bake-off (2026-09-25, `stacks/tei-bakeoff/README.md`).** Prime asked to
evaluate a purpose-built engine instead of vLLM → HF Text Embeddings Inference
1.9.4 (Infinity ruled out: last release 2025-08-22, dropped by us in May for no
Qwen3). Verdict: TEI **matches** (embed cosine vs vLLM-FV median 0.999925, inside
noise; TEI queries against the OLD vLLM index overlap@10 0.988 vs 0.986 self-noise
→ no re-embed needed; rerank decisions identical, tail scores move ≤0.019) but is
**not faster** on the Ada (novel 40 s vs 31 s; tiny queries ~25% faster). It is
**much lighter**: ~2.6 GB VRAM for both vs vLLM's 4.8 GB budget, 8 GB image vs 30,
~4 s restart vs ~24 s. Gateway: embed drop-in via `hosted_vllm/`; rerank needs the
`huggingface/` provider. Both engines left running side by side pending Prime's call.
+9
View File
@@ -117,6 +117,15 @@ no longer deployed sidecars here. See Recent decisions.)
_As of 2026-09-24 ~2245 PT._
### ⭐ esh-ml1: TEI vs vLLM decision pending (Prime), then "what else moves"
Prime's direction 2026-09-25: move the utility seats off fv-ml1 to esh-ml1
(accepting slower); **parakeet STAYS on fv-ml1** (real-time chat path). TEI
bake-off done — results `stacks/tei-bakeoff/README.md`; tei-embed :8081 +
tei-rerank :8083 running beside vLLM, NOT gateway-wired. Next: Prime picks the
engine, then the rest of the move (reward model is the open candidate; it needs
vLLM — TEI can't serve a Llama classifier).
### ✅ esh-ml1 built (2026-09-24) — two follow-ups
The RTX 2000E Ada now serves `qwen3-embedding` + `reranker` from **CT 110
+89
View File
@@ -0,0 +1,89 @@
# tei-bakeoff — TEI vs vLLM for embedding + reranking on esh-ml1
Prime, 2026-09-25: *"run the TEI bake-off."* Hugging Face **Text Embeddings
Inference 1.9.4** (`89-1.9.4`, the Ada build) serving the same two models as
[`embed-rerank`](../embed-rerank/) (vLLM v0.24.0), side by side on the same RTX
2000E Ada. Not wired into the gateway. This stack either becomes the
embed-rerank stack or is torn down.
## Results (2026-09-25 0744–0805 PT)
### Parity — TEI matches the existing vLLM vectors
Reference = vLLM on fv-ml1, the engine every existing index was built with.
306 texts (6 fixed incl. CJK/code/6k-char + 300 *The Stand* paragraphs);
retrieval = 2,000-paragraph corpus, 50 instruction-format queries.
| embedding check | result | noise floor / control |
|---|---|---|
| cosine TEI vs vLLM-FV | median 0.999925, min 0.999861 | vLLM-FV vs itself: 0.999916 / 0.999796 |
| overlap@10, TEI index + TEI queries | 0.976 | vLLM-FV rerun 0.986; vLLM-ESH 0.972 |
| overlap@10, **TEI queries vs the OLD vLLM index** (migration case) | **0.988** | positive control (MRL-256 dims) 0.648 |
| hit@1 own paragraph | 0.80 | vLLM-FV 0.80 |
| different-text negative control | cosine median 0.31 | — |
TEI is deterministic (TEI vs itself: min 0.999995). **Switching engines does not
require re-embedding existing indexes.**
| rerank check (100 queries × 20 docs, 9 same-chapter distractors) | TEI vs vLLM-FV | vLLM-FV vs itself |
|---|---|---|
| source paragraph ranked #1 | 0.95 (same as vLLM) | 0.95 |
| top-1 agreement | 1.00 | 1.00 |
| top-3 set agreement | 1.00 | 1.00 |
| top-5 exact order | 0.98 | 1.00 |
| order among docs scoring > 0.05 | 1.00 (n=37) | 1.00 |
| full 20-doc order | **0.55** | 0.99 |
| max score difference | **0.019** (p99 0.0034) | 0.0016 |
Every decision that matters agrees; the differences are shuffles among
near-zero-scoring tail documents. ⚠ A consumer that **thresholds** on the rerank
score could see a borderline document flip (scores move up to ~0.02). An
earlier 30-query run with random distractors had one top-1 disagreement (29/30);
across both runs, 129/130.
### Speed — TEI is NOT faster on this card
On-box, 2 interleaved runs × 3 reps each, medians:
| workload | vLLM | TEI |
|---|---|---|
| embed 1 short query, p50 | ~9.1 ms | **~6.9 ms** |
| embed 1 × ~512 tok, p50 | ~23 ms (bimodal 10–25) | ~28 ms |
| embed 1 × ~2k tok, p50 | **~86 ms** | ~108 ms |
| bulk embed, passages/s | **~50** | ~37 |
| whole novel (*The Stand*, 12,814 paragraphs), 64/request, 4 in flight | **31 s** | 40 s |
| rerank 20 docs, p50 | **~169 ms** | ~180 ms |
| rerank 20 docs, req/s at conc 8 | ~5.7 | ~5.6 |
TEI ran its fused `FlashQwen3` path. `--max-batch-tokens 32768` did not change
the novel time (39–40 s), so it stays at the default. **Through the gateway these
gaps mostly vanish**: LiteLLM is the bottleneck there (see
`servers/esh-ml1/README.md`).
### Footprint — TEI is much lighter
| | vLLM (both) | TEI (both) |
|---|---|---|
| VRAM (host `nvidia-smi`, per process) | 3,298 + 1,512 MiB | 1,352 + 1,256 MiB |
| image | 29.9 GB | 8.16 GB |
| warm restart to healthy (n=3) | ~24 s | ~4 s |
| container RAM just after start | 2.5 + 5.1 GiB | 0.7 + 1.7 GiB |
⚠ vLLM's VRAM figure is mostly a **setting** (`--gpu-memory-utilization 0.20`
each). A lower setting would shrink it; that has not been tested.
### Gateway compatibility (LiteLLM v1.97, temp aliases, deleted after)
- Embeddings: **drop-in**. TEI's `/v1/embeddings` works with the existing
`hosted_vllm/` provider.
- Rerank: needs the **`huggingface/` provider** (scores identical to direct
TEI). `hosted_vllm/` fails with 422 (`missing field texts`); TEI's `/rerank`
body differs from vLLM's.
## Behaviour differences to carry into any cut-over
- TEI **auto-truncates** over-length input by default; vLLM (at
`--max-model-len 8192`) rejects it with an error.
- TEI caps inputs per request at `--max-client-batch-size` (default 32; set to 128 here).
- TEI on CUDA runs float16 only (vLLM served Qwen3-Embedding in bf16); parity
above says it does not matter here.
+69
View File
@@ -0,0 +1,69 @@
# tei-bakeoff — Hugging Face Text Embeddings Inference (TEI) serving the SAME two
# models as stacks/embed-rerank (vLLM), side by side on esh-ml1, so the two
# engines can be compared on identical hardware. Prime, 2026-09-25: "run the
# TEI bake-off". Not wired into the gateway. Temporary: either this becomes the
# embed-rerank stack or it is torn down.
#
# tei-embed Qwen/Qwen3-Embedding-0.6B → /v1/embeddings, /embed :8081
# tei-rerank BAAI/bge-reranker-v2-m3 → /rerank :8083
#
# 89-* = the Ada Lovelace (sm_89) build. Pinned to an exact release.
# TEI on CUDA offers float16/float32 only (no bfloat16), while vLLM serves
# Qwen3-Embedding in bf16 — the parity measurement is what decides whether
# that matters.
# (--max-batch-tokens 32768 was tried on tei-embed: whole-novel time unchanged,
# 39.2–40.4 s vs 39.9–40.3 s, so it stays at the default.)
# --max-client-batch-size 128: the default is 32; vLLM has no such cap and the
# bake-off sends 64 per request.
# Separate HF cache from vLLM's, so neither engine can disturb the other's files.
name: tei-bakeoff
services:
tei-embed:
image: ghcr.io/huggingface/text-embeddings-inference:89-1.9.4
container_name: tei-embed
restart: unless-stopped
ports:
- "8081:80"
volumes:
- /opt/aimodels/tei-cache:/data
command:
- --model-id
- Qwen/Qwen3-Embedding-0.6B
- --served-model-name
- Qwen/Qwen3-Embedding-0.6B
- --dtype
- float16
- --max-client-batch-size
- "128"
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"]
capabilities: [gpu]
tei-rerank:
image: ghcr.io/huggingface/text-embeddings-inference:89-1.9.4
container_name: tei-rerank
restart: unless-stopped
ports:
- "8083:80"
volumes:
- /opt/aimodels/tei-cache:/data
command:
- --model-id
- BAAI/bge-reranker-v2-m3
- --dtype
- float16
- --max-client-batch-size
- "128"
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"]
capabilities: [gpu]