# esh-ml1 GPU LXC for the ESH home lab: **CT 110 on esh-pve**, holding the **NVIDIA RTX 2000E Ada** (16 GB, 50 W, `01:00.0`). It serves the fleet's embedding and reranking models locally at ESH. Built 2026-09-24. | | | |---|---| | **IP** | `10.0.50.80/24`, VLAN 50, gateway `10.0.50.1` (static, outside the UDM's `.150–.250` DHCP pool) | | **DNS** | `esh-ml1.esh.internal` | | **SSH** | `ssh esh-ml1` → `infra-ops@10.0.50.80` (NOPASSWD sudo) · from the host: `pct enter 110` | | **OS** | Debian 12, unprivileged, `nesting=1,keyctl=1` | | **Size** | 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on `local-lvm` | | **Boot** | `onboot: 1`, `startup: order=30` — after esh-scale (1), esh-vm-db (10) and esh-vm-docker (20), so a GPU fault never delays ESH's DNS or mesh route | | **Backups** | **None, on purpose.** esh-pve's vzdump job lists vmids explicitly and 110 is not one. Everything is rebuilt from the playbooks and the stack; models re-download. | ## What it serves `stacks/embed-rerank` (`/opt/docker/compose/embed-rerank`): | container | model | port | gateway name | |---|---|---|---| | `vllm-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 | `qwen3-embedding` (order 2) | | `vllm-rerank-bge` | `BAAI/bge-reranker-v2-m3` | 8013 | `reranker` (order 2) | The same models, vLLM version (`v0.24.0`, digest `251eba5cc7c1`) and flags as fv-ml1's `vllm` stack, so the two sites are interchangeable. In LiteLLM they are the **order-2 failover** behind fv-ml1: fv-ml1 serves every request while it is up. **Parity, measured 2026-09-24** (11 texts incl. CJK, code, a 6k-char passage; 2 runs per site): | | median | min | |---|---|---| | embed cosine FV vs ESH, same text | 0.999908 | 0.999772 | | noise floor FV vs FV | 0.999927 | 0.999791 | | noise floor ESH vs ESH | 0.999911 | 0.999809 | | negative control, different texts | 0.232 | 0.071 | The cross-site difference is inside each site's own run-to-run noise; this method cannot resolve a cosine gap below ~2×10⁻⁴. Reranker scores differed by at most 0.000145 (FV-vs-FV floor 0.000181), with identical ranking. **VRAM:** 0.20 × 16,380 MiB each; 4,823 MiB in use with both loaded, ~11 GB free. ## How it is built 1. [`playbooks/esh-pve-nvidia-host.yaml`](../../playbooks/esh-pve-nvidia-host.yaml) — driver **580.178.04** (open modules, DKMS) on the **hypervisor**, the `nvidia-persistenced` unit that creates the device nodes before `pve-guests`, and removal of the old VFIO/blacklist config. 2. [`playbooks/esh-ml1-lxc.yaml`](../../playbooks/esh-ml1-lxc.yaml) — the CT, `dev0–3` GPU nodes, the NVIDIA userspace from the **same `.run`** with `--no-kernel-modules`, fleet ids (infra-ops 850, docker 851, vh 1000), docker-ce and nvidia-container-toolkit with `no-cgroups = true`. 3. `scripts/deploy-stack.sh esh-ml1 embed-rerank`, then `docker compose up -d`. Both playbooks are idempotent; re-run them to repair. ## ⚠ Driver version lock The kernel module lives on esh-pve; the libraries live in this container. They **must be the same version**, or every CUDA call fails with *"driver/library version mismatch"*. To upgrade: bump `driver_version` + `driver_sha256` in the host playbook and `driver_version` in the LXC playbook, run the host one, then the LXC one, then restart the stack. A PVE kernel update is handled by DKMS (`proxmox-headers-6.8` pulls headers for each new kernel). Moving esh-pve to a different kernel series (6.14 opt-in) needs that series' headers meta-package installed first, or the module will not build and this CT will fail to start at the next boot. ## Not yet wired - **Homepage**: the compose carries labels, but esh-ml1 is not in `stacks/homepage/conf/docker.yaml` (it would need dockerd on tcp/2375 like the other hosts). - **Beszel**: no agent yet. - **ESH consumers** (Open WebUI RAG, Paperless) still go through the gateway at ana-docker, so they do not survive a mesh outage. See persistent-memory for the open decision.