Files
esh-pfi-infrastructure/servers/esh-ml1/README.md
T

123 lines
6.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# esh-ml1
GPU LXC for the ESH home lab: **CT 110 on esh-pve**, holding the **NVIDIA RTX
2000E Ada** (16 GB, 50 W, `01:00.0`). It serves the fleet's embedding and
reranking models locally at ESH. Built 2026-09-24.
| | |
|---|---|
| **IP** | `10.0.50.80/24`, VLAN 50, gateway `10.0.50.1` (static, outside the UDM's `.150–.250` DHCP pool) |
| **DNS** | `esh-ml1.esh.internal` |
| **SSH** | `ssh esh-ml1` → `infra-ops@10.0.50.80` (NOPASSWD sudo) · from the host: `pct enter 110` |
| **OS** | Debian 12, unprivileged, `nesting=1,keyctl=1` |
| **Size** | 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on `local-lvm` |
| **Boot** | `onboot: 1`, `startup: order=30` — after esh-scale (1), esh-vm-db (10) and esh-vm-docker (20), so a GPU fault never delays ESH's DNS or mesh route |
| **Backups** | **None, on purpose.** esh-pve's vzdump job lists vmids explicitly and 110 is not one. Everything is rebuilt from the playbooks and the stack; models re-download. |
## What it serves
`stacks/embed-rerank` (`/opt/docker/compose/embed-rerank`):
| container | model | port | gateway name |
|---|---|---|---|
| `vllm-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 | `qwen3-embedding` (order 2) |
| `vllm-rerank-bge` | `BAAI/bge-reranker-v2-m3` | 8013 | `reranker` (order 2) |
The same models, vLLM version (`v0.24.0`, digest `251eba5cc7c1`) and flags as
fv-ml1's `vllm` stack, so the two sites are interchangeable. In LiteLLM they
are the **order-2 failover** behind fv-ml1: fv-ml1 serves every request while
it is up.
**Parity, measured 2026-09-24** (11 texts incl. CJK, code, a 6k-char passage;
2 runs per site):
| | median | min |
|---|---|---|
| embed cosine FV vs ESH, same text | 0.999908 | 0.999772 |
| noise floor FV vs FV | 0.999927 | 0.999791 |
| noise floor ESH vs ESH | 0.999911 | 0.999809 |
| negative control, different texts | 0.232 | 0.071 |
The cross-site difference is inside each site's own run-to-run noise; this
method cannot resolve a cosine gap below ~2×10⁻⁴. Reranker scores differed by
at most 0.000145 (FV-vs-FV floor 0.000181), with identical ranking.
**Speed vs fv-ml1, measured 2026-09-25 0628–0640 PT** (same vLLM v0.24.0 and
flags; fv-ml1 GPU 1 is shared with other seats but read 0% util before and
after; synthetic text; each cell = median of 6 reps = 2 interleaved runs × 3):
| workload | on-box FV | on-box ESH | from ana-docker FV | from ana-docker ESH |
|---|---|---|---|---|
| embed 1 query, p50 | ~12 ms | ~9 ms | ~29 ms | ~23 ms |
| embed 1 × ~512 tok, p50 | ~15 ms | ~23 ms | ~31 ms | ~31 ms |
| embed 1 × ~2k tok, p50 | ~25 ms | ~86 ms | ~48 ms | ~90 ms |
| bulk embed, passages/s (64×512 tok, conc 4) | ~403 | ~50 | ~374 | ~51 |
| rerank 20 docs, p50 | ~51 ms | ~170 ms | ~82 ms | ~175 ms |
| rerank 20 docs, req/s at conc 8 | ~55 | ~5.7 | ~52 | ~5.9 |
Read it as: **single queries are a wash** (ESH is ~6 ms faster from the gateway
because it is closer — 6 ms RTT vs 17.6 ms), **anything bulk is 3–10× slower**
on the 50 W Ada. Noise floor (run vs run, same site): ~3–5% on throughput, ±3 ms
on small-query latency on fv-ml1, so the small-query rows cannot rank the sites.
Positive control: 2k-token inputs were slower than 512-token ones on both
sites. One fv-ml1 rep of the gateway-vantage rerank burst stalled to 1.0 req/s
(64 requests in ~64 s); it did not recur in 11 other reps — unexplained, n=1.
**Consequence: keep esh-ml1 as failover, don't load-share** — sharing would
roughly double rerank latency for half the fleet's calls.
**Whole-novel embedding, measured 2026-09-25 0653–0717 PT** — *The Stand*
(uncut, ~1,150 pages): 470,783 words, 12,814 `<p>` paragraphs (median 22 words),
617,833 tokens. Client on nh3-dev (where worldtree-gateway runs); 3 reps each:
| how | fv-ml1 | esh-ml1 |
|---|---|---|
| via gateway, 1 paragraph per request (1,000-para sample, scaled) | 59.8 ms/para → ~12.8 min | 55.3 ms/para → ~11.8 min |
| via gateway, 64 per request, 1 at a time | 100–123 s | 91–100 s |
| via gateway, 64 per request, 4 in flight | 46–54 s | 45–50 s |
| direct to the seat, 64 per request, 4 in flight | 13.4–14.1 s | 28.8–29.6 s |
Through the gateway the GPU is not the bottleneck: one-per-request is network +
LiteLLM overhead, and batched runs pinned LiteLLM at ~76% CPU (64×1024-float
JSON per response). Only the direct path shows the card: esh-ml1 is ~2.2×
slower, ~21k vs ~45k tokens/s.
**VRAM:** 0.20 × 16,380 MiB each; 4,823 MiB in use with both loaded, ~11 GB
free.
## How it is built
1. [`playbooks/esh-pve-nvidia-host.yaml`](../../playbooks/esh-pve-nvidia-host.yaml)
— driver **580.178.04** (open modules, DKMS) on the **hypervisor**, the
`nvidia-persistenced` unit that creates the device nodes before
`pve-guests`, and removal of the old VFIO/blacklist config.
2. [`playbooks/esh-ml1-lxc.yaml`](../../playbooks/esh-ml1-lxc.yaml) — the CT,
`dev0–3` GPU nodes, the NVIDIA userspace from the **same `.run`** with
`--no-kernel-modules`, fleet ids (infra-ops 850, docker 851, vh 1000),
docker-ce and nvidia-container-toolkit with `no-cgroups = true`.
3. `scripts/deploy-stack.sh esh-ml1 embed-rerank`, then `docker compose up -d`.
Both playbooks are idempotent; re-run them to repair.
## ⚠ Driver version lock
The kernel module lives on esh-pve; the libraries live in this container. They
**must be the same version**, or every CUDA call fails with *"driver/library
version mismatch"*. To upgrade: bump `driver_version` + `driver_sha256` in the
host playbook and `driver_version` in the LXC playbook, run the host one, then
the LXC one, then restart the stack.
A PVE kernel update is handled by DKMS (`proxmox-headers-6.8` pulls headers
for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in)
needs that series' headers meta-package installed first, or the module will not
build and this CT will fail to start at the next boot.
## Not yet wired
- **Homepage**: the compose carries labels, but esh-ml1 is not in
`stacks/homepage/conf/docker.yaml` (it would need dockerd on tcp/2375 like
the other hosts).
- **Beszel**: no agent yet.
- **ESH consumers** (Open WebUI RAG, Paperless) still go through the gateway
at ana-docker, so they do not survive a mesh outage. See persistent-memory
for the open decision.