Files
esh-pfi-infrastructure/servers/esh-ml1/README.md
T
vh 5402568b76 feat(esh-ml1): RTX 2000E Ada on esh-pve serves embed + rerank as a LiteLLM failover
esh-pve: NVIDIA 580.178.04 (open modules, DKMS) installed on the host from
NVIDIA's .run and loaded live, no reboot. nvidia-persistenced unit creates the
device nodes before pve-guests; the T400's vfio-pci ids and `blacklist nvidia`
retired. playbooks/esh-pve-nvidia-host.yaml.

esh-ml1: CT 110, unprivileged Debian 12, 10.0.50.80, GPU nodes via devN,
NVIDIA userspace from the same .run (--no-kernel-modules), docker-ce +
nvidia-container-toolkit (no-cgroups). playbooks/esh-ml1-lxc.yaml. Not in the
vzdump job on purpose. DNS esh-ml1.esh.internal.

stacks/embed-rerank: Qwen3-Embedding-0.6B :8001 + bge-reranker-v2-m3 :8013 on
the same vLLM v0.24.0 digest and flags as fv-ml1. Measured parity: embed
cosine FV-vs-ESH median 0.999908 (min 0.999772), inside both self-noise
floors; rerank max |delta| 0.000145 vs floor 0.000181, identical ranking.

litellm: qwen3-embedding and reranker gain an esh-ml1 deployment at order 2
behind fv-ml1 (order 1). Order fallback proven with throwaway groups:
refused primary +0.15 s, host-down primary ~18.7 s per call, dead-only 500.

Also: repaired the DB-only alias reranker-a3-bge-v2-m3, dead since the
fv-ml1 relocation (still named 10.250.50.54); documented the third
unkillable homepage wedge on esh-docker-vm.
2026-09-24 22:38:55 -07:00

3.9 KiB
Raw Blame History

esh-ml1

GPU LXC for the ESH home lab: CT 110 on esh-pve, holding the NVIDIA RTX 2000E Ada (16 GB, 50 W, 01:00.0). It serves the fleet's embedding and reranking models locally at ESH. Built 2026-09-24.

IP 10.0.50.80/24, VLAN 50, gateway 10.0.50.1 (static, outside the UDM's .150–.250 DHCP pool)
DNS esh-ml1.esh.internal
SSH ssh esh-ml1 → infra-ops@10.0.50.80 (NOPASSWD sudo) · from the host: pct enter 110
OS Debian 12, unprivileged, nesting=1,keyctl=1
Size 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on local-lvm
Boot onboot: 1, startup: order=30 — after esh-scale (1), esh-vm-db (10) and esh-vm-docker (20), so a GPU fault never delays ESH's DNS or mesh route
Backups None, on purpose. esh-pve's vzdump job lists vmids explicitly and 110 is not one. Everything is rebuilt from the playbooks and the stack; models re-download.

What it serves

stacks/embed-rerank (/opt/docker/compose/embed-rerank):

container model port gateway name
vllm-embed Qwen/Qwen3-Embedding-0.6B 8001 qwen3-embedding (order 2)
vllm-rerank-bge BAAI/bge-reranker-v2-m3 8013 reranker (order 2)

The same models, vLLM version (v0.24.0, digest 251eba5cc7c1) and flags as fv-ml1's vllm stack, so the two sites are interchangeable. In LiteLLM they are the order-2 failover behind fv-ml1: fv-ml1 serves every request while it is up.

Parity, measured 2026-09-24 (11 texts incl. CJK, code, a 6k-char passage; 2 runs per site):

median min
embed cosine FV vs ESH, same text 0.999908 0.999772
noise floor FV vs FV 0.999927 0.999791
noise floor ESH vs ESH 0.999911 0.999809
negative control, different texts 0.232 0.071

The cross-site difference is inside each site's own run-to-run noise; this method cannot resolve a cosine gap below ~2×10⁻⁴. Reranker scores differed by at most 0.000145 (FV-vs-FV floor 0.000181), with identical ranking.

VRAM: 0.20 × 16,380 MiB each; 4,823 MiB in use with both loaded, ~11 GB free.

How it is built

  1. playbooks/esh-pve-nvidia-host.yaml — driver 580.178.04 (open modules, DKMS) on the hypervisor, the nvidia-persistenced unit that creates the device nodes before pve-guests, and removal of the old VFIO/blacklist config.
  2. playbooks/esh-ml1-lxc.yaml — the CT, dev0–3 GPU nodes, the NVIDIA userspace from the same .run with --no-kernel-modules, fleet ids (infra-ops 850, docker 851, vh 1000), docker-ce and nvidia-container-toolkit with no-cgroups = true.
  3. scripts/deploy-stack.sh esh-ml1 embed-rerank, then docker compose up -d.

Both playbooks are idempotent; re-run them to repair.

⚠ Driver version lock

The kernel module lives on esh-pve; the libraries live in this container. They must be the same version, or every CUDA call fails with "driver/library version mismatch". To upgrade: bump driver_version + driver_sha256 in the host playbook and driver_version in the LXC playbook, run the host one, then the LXC one, then restart the stack.

A PVE kernel update is handled by DKMS (proxmox-headers-6.8 pulls headers for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in) needs that series' headers meta-package installed first, or the module will not build and this CT will fail to start at the next boot.

Not yet wired

  • Homepage: the compose carries labels, but esh-ml1 is not in stacks/homepage/conf/docker.yaml (it would need dockerd on tcp/2375 like the other hosts).
  • Beszel: no agent yet.
  • ESH consumers (Open WebUI RAG, Paperless) still go through the gateway at ana-docker, so they do not survive a mesh outage. See persistent-memory for the open decision.