Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md
T

81 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `[2026-09-24]` esh-ml1 built: RTX 2000E Ada serving embed + rerank, LiteLLM order-2 failover behind fv-ml1
Executes Prime's 2026-09-24 decision (LXC + host driver, not a VFIO VM). Done
2204–2240 PT the same night, **with no esh-pve reboot**.
**Host (esh-pve).** NVIDIA **580.178.04** open modules via DKMS from NVIDIA's
`-no-compat32.run` (sha256 checked against NVIDIA's published sum), applied
live by `playbooks/esh-pve-nvidia-host.yaml`: nouveau was never loaded and the
T400's vfio ids matched nothing, so nothing held the card. Our own
`nvidia-persistenced.service` creates `/dev/nvidia{0,ctl,-uvm,-uvm-tools}`
**before `pve-guests`** (else `pct start` refuses the `devN:` entries). Retired
`vfio.conf` (moved to `/root/nvidia/`) and `blacklist nvidia`. Initramfs NOT
rebuilt on purpose (boot path untouched). 580 because the fleet's vLLM
`v0.24.0` image is CUDA 13.0 (driver ≥ 580); fv-ml1 runs 580.65.06. `.run`
not an apt repo: one file on both sides makes the host-module/container-lib
version match true by construction; apt would let an upgrade move one side.
**CT 110 `esh-ml1`** (`playbooks/esh-ml1-lxc.yaml`): unprivileged Debian 12,
nesting+keyctl, 6c/16G/80G `local-lvm`, `10.0.50.80` VLAN 50 (static; UDM pool
is `.150–.250`), `startup order=30` after esh-scale/VMs. NVIDIA userspace from
the same `.run` with `--no-kernel-modules`; docker-ce; nvidia-container-toolkit
with `no-cgroups=true`. **Docker-in-unprivileged-LXC with the GPU worked first
try on lxc-pve 6.0.0-2** (no AppArmor workaround needed). Fleet ids per
`docs/pfi/fleet-conventions.md` (infra-ops 850, docker 851, vh 1000 — vh has
no key/password yet). **Not in the vzdump job, on purpose** (rebuildable).
DNS `esh-ml1.esh.internal` synced to all three resolvers.
**Stack `stacks/embed-rerank`:** `vllm-embed` :8001 + `vllm-rerank-bge` :8013,
same models/flags/vLLM digest (`251eba5cc7c1`) as fv-ml1, 0.20 mem-util each →
4,823 MiB used of 16,380.
**Parity (the load-bearing measurement; embeddings are model-specific).** 11
texts × 2 runs per site. Embed cosine FV-vs-ESH median 0.999908 / min 0.999772,
inside both self-noise floors (FV-vs-FV 0.999927 / 0.999791; ESH-vs-ESH
0.999911 / 0.999809). Negative control (different texts) 0.07–0.38. Sensitivity
floor ~2×10⁻⁴ cosine. Rerank |score| max 0.000145 vs FV-FV floor 0.000181,
identical ranking.
**Gateway.** `qwen3-embedding` and `reranker` each got a second deployment
(`order: 2`, esh-ml1); fv-ml1 is `order: 1`. LiteLLM v1.97 order-fallback
proven with throwaway groups (created + deleted): primary refusing → ESH in
+0.15 s (3/3); primary host-down (no ARP) → ESH but **~18.7 s on every call,
no cooldown** (3/3); dead-only negative control → 500 (3/3). Post-restart
production traffic is served by fv-ml1 (header `x-litellm-model-api-base`).
`/health?model=` returned 503 with empty lists for both groups — not
investigated. Gateway restart: liveliness back in ~52 s.
**Found + fixed:** DB-only alias `reranker-a3-bge-v2-m3` pointed at ana-ml2's
pre-relocation IP `10.250.50.54` — dead since 2026-09-12 (500 after ~23 s), no
callers in 7-day spend logs. PATCHed to `10.251.50.54:8013`. Sweep of
`/model/info` found no other `10.250.50.54` targets.
**Who actually uses these names (spend logs, 7 days to 2026-09-25 05:35Z):**
`qwen3-embedding` — worldtree-gateway 481 (nh3-dev + corviduo-dev), nevermore
20; `reranker` — nevermore 13. **No ESH-side consumer called either in 7 days**
(Open WebUI's RAG embeds only on document upload). So today esh-ml1's value is
an FV-outage failover for Worldtree + nevermore, not local service for ESH.
**Open:** a direct ESH→esh-ml1 path (survives a mesh outage) would help only
ESH consumers, which are idle — recommended parked. esh-ml1 not yet in
Homepage docker.yaml or Beszel. The ~18.7 s host-down failover penalty is
untuned (router retries/cooldown).
**Speed A/B vs fv-ml1 (2026-09-25 0628–0640 PT, table in `servers/esh-ml1/README.md`).**
Single queries: a wash (on-box ESH ~9 ms vs FV ~12 ms; from ana-docker ESH is
faster because it is network-closer). Bulk: ESH 3–10× slower (bulk embed ~50 vs
~403 passages/s; rerank-20 ~170 vs ~51 ms p50, ~5.7 vs ~55 req/s at conc 8).
Noise floor 3–5% on throughput. ⇒ failover role confirmed; load-sharing would
roughly double rerank latency for half the calls.
**TEI bake-off (2026-09-25, `stacks/tei-bakeoff/README.md`).** Prime asked to
evaluate a purpose-built engine instead of vLLM → HF Text Embeddings Inference
1.9.4 (Infinity ruled out: last release 2025-08-22, dropped by us in May for no
Qwen3). Verdict: TEI **matches** (embed cosine vs vLLM-FV median 0.999925, inside
noise; TEI queries against the OLD vLLM index overlap@10 0.988 vs 0.986 self-noise
→ no re-embed needed; rerank decisions identical, tail scores move ≤0.019) but is
**not faster** on the Ada (novel 40 s vs 31 s; tiny queries ~25% faster). It is
**much lighter**: ~2.6 GB VRAM for both vs vLLM's 4.8 GB budget, 8 GB image vs 30,
~4 s restart vs ~24 s. Gateway: embed drop-in via `hosted_vllm/`; rerank needs the
`huggingface/` provider. Both engines left running side by side pending Prime's call.