Files
esh-pfi-infrastructure/servers/esh-ml1/README.md
T
vh 5402568b76 feat(esh-ml1): RTX 2000E Ada on esh-pve serves embed + rerank as a LiteLLM failover
esh-pve: NVIDIA 580.178.04 (open modules, DKMS) installed on the host from
NVIDIA's .run and loaded live, no reboot. nvidia-persistenced unit creates the
device nodes before pve-guests; the T400's vfio-pci ids and `blacklist nvidia`
retired. playbooks/esh-pve-nvidia-host.yaml.

esh-ml1: CT 110, unprivileged Debian 12, 10.0.50.80, GPU nodes via devN,
NVIDIA userspace from the same .run (--no-kernel-modules), docker-ce +
nvidia-container-toolkit (no-cgroups). playbooks/esh-ml1-lxc.yaml. Not in the
vzdump job on purpose. DNS esh-ml1.esh.internal.

stacks/embed-rerank: Qwen3-Embedding-0.6B :8001 + bge-reranker-v2-m3 :8013 on
the same vLLM v0.24.0 digest and flags as fv-ml1. Measured parity: embed
cosine FV-vs-ESH median 0.999908 (min 0.999772), inside both self-noise
floors; rerank max |delta| 0.000145 vs floor 0.000181, identical ranking.

litellm: qwen3-embedding and reranker gain an esh-ml1 deployment at order 2
behind fv-ml1 (order 1). Order fallback proven with throwaway groups:
refused primary +0.15 s, host-down primary ~18.7 s per call, dead-only 500.

Also: repaired the DB-only alias reranker-a3-bge-v2-m3, dead since the
fv-ml1 relocation (still named 10.250.50.54); documented the third
unkillable homepage wedge on esh-docker-vm.
2026-09-24 22:38:55 -07:00

84 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# esh-ml1
GPU LXC for the ESH home lab: **CT 110 on esh-pve**, holding the **NVIDIA RTX
2000E Ada** (16 GB, 50 W, `01:00.0`). It serves the fleet's embedding and
reranking models locally at ESH. Built 2026-09-24.
| | |
|---|---|
| **IP** | `10.0.50.80/24`, VLAN 50, gateway `10.0.50.1` (static, outside the UDM's `.150–.250` DHCP pool) |
| **DNS** | `esh-ml1.esh.internal` |
| **SSH** | `ssh esh-ml1` → `infra-ops@10.0.50.80` (NOPASSWD sudo) · from the host: `pct enter 110` |
| **OS** | Debian 12, unprivileged, `nesting=1,keyctl=1` |
| **Size** | 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on `local-lvm` |
| **Boot** | `onboot: 1`, `startup: order=30` — after esh-scale (1), esh-vm-db (10) and esh-vm-docker (20), so a GPU fault never delays ESH's DNS or mesh route |
| **Backups** | **None, on purpose.** esh-pve's vzdump job lists vmids explicitly and 110 is not one. Everything is rebuilt from the playbooks and the stack; models re-download. |
## What it serves
`stacks/embed-rerank` (`/opt/docker/compose/embed-rerank`):
| container | model | port | gateway name |
|---|---|---|---|
| `vllm-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 | `qwen3-embedding` (order 2) |
| `vllm-rerank-bge` | `BAAI/bge-reranker-v2-m3` | 8013 | `reranker` (order 2) |
The same models, vLLM version (`v0.24.0`, digest `251eba5cc7c1`) and flags as
fv-ml1's `vllm` stack, so the two sites are interchangeable. In LiteLLM they
are the **order-2 failover** behind fv-ml1: fv-ml1 serves every request while
it is up.
**Parity, measured 2026-09-24** (11 texts incl. CJK, code, a 6k-char passage;
2 runs per site):
| | median | min |
|---|---|---|
| embed cosine FV vs ESH, same text | 0.999908 | 0.999772 |
| noise floor FV vs FV | 0.999927 | 0.999791 |
| noise floor ESH vs ESH | 0.999911 | 0.999809 |
| negative control, different texts | 0.232 | 0.071 |
The cross-site difference is inside each site's own run-to-run noise; this
method cannot resolve a cosine gap below ~2×10⁻⁴. Reranker scores differed by
at most 0.000145 (FV-vs-FV floor 0.000181), with identical ranking.
**VRAM:** 0.20 × 16,380 MiB each; 4,823 MiB in use with both loaded, ~11 GB
free.
## How it is built
1. [`playbooks/esh-pve-nvidia-host.yaml`](../../playbooks/esh-pve-nvidia-host.yaml)
— driver **580.178.04** (open modules, DKMS) on the **hypervisor**, the
`nvidia-persistenced` unit that creates the device nodes before
`pve-guests`, and removal of the old VFIO/blacklist config.
2. [`playbooks/esh-ml1-lxc.yaml`](../../playbooks/esh-ml1-lxc.yaml) — the CT,
`dev0–3` GPU nodes, the NVIDIA userspace from the **same `.run`** with
`--no-kernel-modules`, fleet ids (infra-ops 850, docker 851, vh 1000),
docker-ce and nvidia-container-toolkit with `no-cgroups = true`.
3. `scripts/deploy-stack.sh esh-ml1 embed-rerank`, then `docker compose up -d`.
Both playbooks are idempotent; re-run them to repair.
## ⚠ Driver version lock
The kernel module lives on esh-pve; the libraries live in this container. They
**must be the same version**, or every CUDA call fails with *"driver/library
version mismatch"*. To upgrade: bump `driver_version` + `driver_sha256` in the
host playbook and `driver_version` in the LXC playbook, run the host one, then
the LXC one, then restart the stack.
A PVE kernel update is handled by DKMS (`proxmox-headers-6.8` pulls headers
for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in)
needs that series' headers meta-package installed first, or the module will not
build and this CT will fail to start at the next boot.
## Not yet wired
- **Homepage**: the compose carries labels, but esh-ml1 is not in
`stacks/homepage/conf/docker.yaml` (it would need dockerd on tcp/2375 like
the other hosts).
- **Beszel**: no agent yet.
- **ESH consumers** (Open WebUI RAG, Paperless) still go through the gateway
at ana-docker, so they do not survive a mesh outage. See persistent-memory
for the open decision.