fix(tei-bakeoff): fail-closed truncation; record Dvalin review and nevermore threshold check
This commit is contained in:
@@ -72,7 +72,22 @@ gaps mostly vanish**: LiteLLM is the bottleneck there (see
|
|||||||
⚠ vLLM's VRAM figure is mostly a **setting** (`--gpu-memory-utilization 0.20`
|
⚠ vLLM's VRAM figure is mostly a **setting** (`--gpu-memory-utilization 0.20`
|
||||||
each). A lower setting would shrink it; that has not been tested.
|
each). A lower setting would shrink it; that has not been tested.
|
||||||
|
|
||||||
### Gateway compatibility (LiteLLM v1.97, temp aliases, deleted after)
|
### External review — Dvalin (Grok research peer), 2026-09-25
|
||||||
|
|
||||||
|
**Agree with caveats: adopt TEI for these two seats; footprint is the right
|
||||||
|
reason.** His external evidence (not re-verified here): a Runpod 2026-09-14
|
||||||
|
engine comparison shows vLLM ahead of TEI on Qwen3-Embedding bulk throughput
|
||||||
|
(median ~2.6×, single unreplicated runs, near parity on smaller cards) and TEI
|
||||||
|
ahead on BERT-family models — consistent with our 1.3× on a 50 W card and the
|
||||||
|
rerank near-tie. Open TEI 1.9 issue #857 (tokio panic "No backend receiver"
|
||||||
|
under load, n=1) is a watch item, not a gate. A lower vLLM memory setting
|
||||||
|
would shrink only the embedder (the reranker at 1,512 MiB is already under
|
||||||
|
its cap); 0.12 is the only setting he'd expect to both boot and help — untested.
|
||||||
|
His caveats, and status: fail-closed truncation (**done**, above); recheck score-
|
||||||
|
threshold consumers (**done**, nevermore, above); two gateway providers must be
|
||||||
|
written down at cut-over (open).
|
||||||
|
|
||||||
|
## Gateway compatibility (LiteLLM v1.97, temp aliases, deleted after)
|
||||||
|
|
||||||
- Embeddings: **drop-in**. TEI's `/v1/embeddings` works with the existing
|
- Embeddings: **drop-in**. TEI's `/v1/embeddings` works with the existing
|
||||||
`hosted_vllm/` provider.
|
`hosted_vllm/` provider.
|
||||||
@@ -82,8 +97,21 @@ each). A lower setting would shrink it; that has not been tested.
|
|||||||
|
|
||||||
## Behaviour differences to carry into any cut-over
|
## Behaviour differences to carry into any cut-over
|
||||||
|
|
||||||
- TEI **auto-truncates** over-length input by default; vLLM (at
|
- **Truncation — now fail-closed (`--auto-truncate false`).** Measured with
|
||||||
`--max-model-len 8192`) rejects it with an error.
|
TEI's default: a ~40k-token input returned **200 with a vector of its first
|
||||||
|
16,384 tokens**, silently, where vLLM returns 400. With truncation off, TEI
|
||||||
|
embed rejects > 32,768 tokens and TEI rerank rejects > 8,192 (both 422,
|
||||||
|
verified). Truncation off requires `--max-batch-tokens` at or above the model's max
|
||||||
|
input (32,768 for Qwen3-Embedding) or TEI refuses to start; VRAM unchanged.
|
||||||
|
Remaining difference: TEI embed **accepts** 8,193–32,767 tokens that vLLM
|
||||||
|
(`--max-model-len 8192`) rejects. It embeds the whole input, which is wider,
|
||||||
|
not lossy.
|
||||||
|
- **nevermore thresholds rerank scores** (`NEVERMORE_RERANK_THRESHOLD`, default
|
||||||
|
0.3, cluster-member filter in `digest.py`). Across the 2,000 bake-off scores,
|
||||||
|
**0 flips** at 0.3 or 0.5 for TEI vs vLLM — but only 1 score landed within
|
||||||
|
±0.02 of 0.3 (bge scores sit near 0 or 1), so this cannot exclude rare flips
|
||||||
|
on borderline headlines. Worst case: a same-story headline clusters
|
||||||
|
differently. Low stakes.
|
||||||
- TEI caps inputs per request at `--max-client-batch-size` (default 32; set to 128 here).
|
- TEI caps inputs per request at `--max-client-batch-size` (default 32; set to 128 here).
|
||||||
- TEI on CUDA runs float16 only (vLLM served Qwen3-Embedding in bf16); parity
|
- TEI on CUDA runs float16 only (vLLM served Qwen3-Embedding in bf16); parity
|
||||||
above says it does not matter here.
|
above says it does not matter here.
|
||||||
|
|||||||
@@ -37,6 +37,14 @@ services:
|
|||||||
- float16
|
- float16
|
||||||
- --max-client-batch-size
|
- --max-client-batch-size
|
||||||
- "128"
|
- "128"
|
||||||
|
# FAIL-CLOSED on over-length input (Dvalin review 2026-09-25; measured:
|
||||||
|
# TEI's default silently cut a ~40k-token input to 16,384 and returned
|
||||||
|
# 200, where vLLM returns 400). Truncation off requires max-batch-tokens
|
||||||
|
# >= the model's max input (Qwen3-Embedding: 32,768) or TEI refuses to start.
|
||||||
|
- --auto-truncate
|
||||||
|
- "false"
|
||||||
|
- --max-batch-tokens
|
||||||
|
- "32768"
|
||||||
deploy:
|
deploy:
|
||||||
resources:
|
resources:
|
||||||
reservations:
|
reservations:
|
||||||
@@ -60,6 +68,10 @@ services:
|
|||||||
- float16
|
- float16
|
||||||
- --max-client-batch-size
|
- --max-client-batch-size
|
||||||
- "128"
|
- "128"
|
||||||
|
# Fail-closed, as tei-embed. bge-reranker-v2-m3 max input (8,192) fits the
|
||||||
|
# default max-batch-tokens (16,384).
|
||||||
|
- --auto-truncate
|
||||||
|
- "false"
|
||||||
deploy:
|
deploy:
|
||||||
resources:
|
resources:
|
||||||
reservations:
|
reservations:
|
||||||
|
|||||||
Reference in New Issue
Block a user