81 lines
5.1 KiB
Markdown
81 lines
5.1 KiB
Markdown
# `[2026-09-24]` esh-ml1 built: RTX 2000E Ada serving embed + rerank, LiteLLM order-2 failover behind fv-ml1
|
||
|
||
Executes Prime's 2026-09-24 decision (LXC + host driver, not a VFIO VM). Done
|
||
2204–2240 PT the same night, **with no esh-pve reboot**.
|
||
|
||
**Host (esh-pve).** NVIDIA **580.178.04** open modules via DKMS from NVIDIA's
|
||
`-no-compat32.run` (sha256 checked against NVIDIA's published sum), applied
|
||
live by `playbooks/esh-pve-nvidia-host.yaml`: nouveau was never loaded and the
|
||
T400's vfio ids matched nothing, so nothing held the card. Our own
|
||
`nvidia-persistenced.service` creates `/dev/nvidia{0,ctl,-uvm,-uvm-tools}`
|
||
**before `pve-guests`** (else `pct start` refuses the `devN:` entries). Retired
|
||
`vfio.conf` (moved to `/root/nvidia/`) and `blacklist nvidia`. Initramfs NOT
|
||
rebuilt on purpose (boot path untouched). 580 because the fleet's vLLM
|
||
`v0.24.0` image is CUDA 13.0 (driver ≥ 580); fv-ml1 runs 580.65.06. `.run`
|
||
not an apt repo: one file on both sides makes the host-module/container-lib
|
||
version match true by construction; apt would let an upgrade move one side.
|
||
|
||
**CT 110 `esh-ml1`** (`playbooks/esh-ml1-lxc.yaml`): unprivileged Debian 12,
|
||
nesting+keyctl, 6c/16G/80G `local-lvm`, `10.0.50.80` VLAN 50 (static; UDM pool
|
||
is `.150–.250`), `startup order=30` after esh-scale/VMs. NVIDIA userspace from
|
||
the same `.run` with `--no-kernel-modules`; docker-ce; nvidia-container-toolkit
|
||
with `no-cgroups=true`. **Docker-in-unprivileged-LXC with the GPU worked first
|
||
try on lxc-pve 6.0.0-2** (no AppArmor workaround needed). Fleet ids per
|
||
`docs/pfi/fleet-conventions.md` (infra-ops 850, docker 851, vh 1000 — vh has
|
||
no key/password yet). **Not in the vzdump job, on purpose** (rebuildable).
|
||
DNS `esh-ml1.esh.internal` synced to all three resolvers.
|
||
|
||
**Stack `stacks/embed-rerank`:** `vllm-embed` :8001 + `vllm-rerank-bge` :8013,
|
||
same models/flags/vLLM digest (`251eba5cc7c1`) as fv-ml1, 0.20 mem-util each →
|
||
4,823 MiB used of 16,380.
|
||
|
||
**Parity (the load-bearing measurement; embeddings are model-specific).** 11
|
||
texts × 2 runs per site. Embed cosine FV-vs-ESH median 0.999908 / min 0.999772,
|
||
inside both self-noise floors (FV-vs-FV 0.999927 / 0.999791; ESH-vs-ESH
|
||
0.999911 / 0.999809). Negative control (different texts) 0.07–0.38. Sensitivity
|
||
floor ~2×10⁻⁴ cosine. Rerank |score| max 0.000145 vs FV-FV floor 0.000181,
|
||
identical ranking.
|
||
|
||
**Gateway.** `qwen3-embedding` and `reranker` each got a second deployment
|
||
(`order: 2`, esh-ml1); fv-ml1 is `order: 1`. LiteLLM v1.97 order-fallback
|
||
proven with throwaway groups (created + deleted): primary refusing → ESH in
|
||
+0.15 s (3/3); primary host-down (no ARP) → ESH but **~18.7 s on every call,
|
||
no cooldown** (3/3); dead-only negative control → 500 (3/3). Post-restart
|
||
production traffic is served by fv-ml1 (header `x-litellm-model-api-base`).
|
||
`/health?model=` returned 503 with empty lists for both groups — not
|
||
investigated. Gateway restart: liveliness back in ~52 s.
|
||
|
||
**Found + fixed:** DB-only alias `reranker-a3-bge-v2-m3` pointed at ana-ml2's
|
||
pre-relocation IP `10.250.50.54` — dead since 2026-09-12 (500 after ~23 s), no
|
||
callers in 7-day spend logs. PATCHed to `10.251.50.54:8013`. Sweep of
|
||
`/model/info` found no other `10.250.50.54` targets.
|
||
|
||
**Who actually uses these names (spend logs, 7 days to 2026-09-25 05:35Z):**
|
||
`qwen3-embedding` — worldtree-gateway 481 (nh3-dev + corviduo-dev), nevermore
|
||
20; `reranker` — nevermore 13. **No ESH-side consumer called either in 7 days**
|
||
(Open WebUI's RAG embeds only on document upload). So today esh-ml1's value is
|
||
an FV-outage failover for Worldtree + nevermore, not local service for ESH.
|
||
|
||
**Open:** a direct ESH→esh-ml1 path (survives a mesh outage) would help only
|
||
ESH consumers, which are idle — recommended parked. esh-ml1 not yet in
|
||
Homepage docker.yaml or Beszel. The ~18.7 s host-down failover penalty is
|
||
untuned (router retries/cooldown).
|
||
|
||
**Speed A/B vs fv-ml1 (2026-09-25 0628–0640 PT, table in `servers/esh-ml1/README.md`).**
|
||
Single queries: a wash (on-box ESH ~9 ms vs FV ~12 ms; from ana-docker ESH is
|
||
faster because it is network-closer). Bulk: ESH 3–10× slower (bulk embed ~50 vs
|
||
~403 passages/s; rerank-20 ~170 vs ~51 ms p50, ~5.7 vs ~55 req/s at conc 8).
|
||
Noise floor 3–5% on throughput. ⇒ failover role confirmed; load-sharing would
|
||
roughly double rerank latency for half the calls.
|
||
|
||
**TEI bake-off (2026-09-25, `stacks/tei-bakeoff/README.md`).** Prime asked to
|
||
evaluate a purpose-built engine instead of vLLM → HF Text Embeddings Inference
|
||
1.9.4 (Infinity ruled out: last release 2025-08-22, dropped by us in May for no
|
||
Qwen3). Verdict: TEI **matches** (embed cosine vs vLLM-FV median 0.999925, inside
|
||
noise; TEI queries against the OLD vLLM index overlap@10 0.988 vs 0.986 self-noise
|
||
→ no re-embed needed; rerank decisions identical, tail scores move ≤0.019) but is
|
||
**not faster** on the Ada (novel 40 s vs 31 s; tiny queries ~25% faster). It is
|
||
**much lighter**: ~2.6 GB VRAM for both vs vLLM's 4.8 GB budget, 8 GB image vs 30,
|
||
~4 s restart vs ~24 s. Gateway: embed drop-in via `hosted_vllm/`; rerank needs the
|
||
`huggingface/` provider. Both engines left running side by side pending Prime's call.
|