feat(esh-ml1): RTX 2000E Ada on esh-pve serves embed + rerank as a LiteLLM failover

esh-pve: NVIDIA 580.178.04 (open modules, DKMS) installed on the host from
NVIDIA's .run and loaded live, no reboot. nvidia-persistenced unit creates the
device nodes before pve-guests; the T400's vfio-pci ids and `blacklist nvidia`
retired. playbooks/esh-pve-nvidia-host.yaml.

esh-ml1: CT 110, unprivileged Debian 12, 10.0.50.80, GPU nodes via devN,
NVIDIA userspace from the same .run (--no-kernel-modules), docker-ce +
nvidia-container-toolkit (no-cgroups). playbooks/esh-ml1-lxc.yaml. Not in the
vzdump job on purpose. DNS esh-ml1.esh.internal.

stacks/embed-rerank: Qwen3-Embedding-0.6B :8001 + bge-reranker-v2-m3 :8013 on
the same vLLM v0.24.0 digest and flags as fv-ml1. Measured parity: embed
cosine FV-vs-ESH median 0.999908 (min 0.999772), inside both self-noise
floors; rerank max |delta| 0.000145 vs floor 0.000181, identical ranking.

litellm: qwen3-embedding and reranker gain an esh-ml1 deployment at order 2
behind fv-ml1 (order 1). Order fallback proven with throwaway groups:
refused primary +0.15 s, host-down primary ~18.7 s per call, dead-only 500.

Also: repaired the DB-only alias reranker-a3-bge-v2-m3, dead since the
fv-ml1 relocation (still named 10.250.50.54); documented the third
unkillable homepage wedge on esh-docker-vm.
This commit is contained in:
vh
2026-09-24 22:38:55 -07:00
parent eb46973051
commit 5402568b76
16 changed files with 1053 additions and 30 deletions
+35 -5
View File
@@ -13,7 +13,8 @@ Proxmox VE hypervisor for the ESH home lab (`pve.esteban.net`). Non-PFI scope.
- **CPU:** 13th Gen Intel Core i9-13900H
- **RAM:** 62.5 GB
- **Kernel:** `6.8.12-16-pve` (Proxmox 8.x)
- **Kernel:** `6.8.12-42-pve`, PVE 8.4.20 (as of 2026-09-24)
- **GPU:** NVIDIA RTX 2000E Ada, 16 GB (`01:00.0`) — driven by the **host** (580.178.04, DKMS), used by the esh-ml1 LXC. See below.
- **Storage:** local `pve-root` at 94G (21% used) + two large NFS mounts from `10.0.50.50`:
- `/mnt/pve/esh-nas` — 92 TB available on `/mnt/pvestore`
- `/mnt/pve/tank-vmbu` — 93 TB, VM backup target (~1.4 TB used)
@@ -47,10 +48,39 @@ fully up. esh-vm-docker, esh-vm-db and esh-scale all started by themselves.
esh-vm-docker's NFS (`/mnt/books`, `/mnt/backup`) mounted, traefik served with no
403s, and AdGuard resolved.
The new card has **no driver bound**. `/etc/modprobe.d` still lists the T400's IDs
for vfio-pci (`10de:1ff2,10de:10fa`), which no longer match anything; harmless, but
clean it up when the card's use is decided. The intended use is a local
embedding endpoint.
The new card got its driver the same night — see the next section.
## GPU — host driver + the esh-ml1 LXC (2026-09-24)
**Prime's decision: LXC + host driver, NOT a VFIO VM.** Both hangs on this box
this year came from VFIO passthrough. The card now runs under the host's own
NVIDIA driver and is shared into **CT 110 `esh-ml1`**, which serves embedding
and reranking ([`servers/esh-ml1/README.md`](../esh-ml1/README.md)).
- **Driver 580.178.04, open kernel modules, DKMS**, from NVIDIA's `.run`
(`/root/nvidia/`). Applied by
[`playbooks/esh-pve-nvidia-host.yaml`](../../playbooks/esh-pve-nvidia-host.yaml)
**live, with no reboot**: nouveau was never loaded and nothing held the card.
- **`nvidia-persistenced.service`** (ours, in `/etc/systemd/system`) runs
`nvidia-modprobe -c0 -u` and the persistence daemon **before
`pve-guests`**, so `/dev/nvidia0`, `nvidiactl` and `nvidia-uvm{,-tools}`
exist when CT 110 starts. Without them `pct start` refuses the `devN:`
entries. If the unit fails, only CT 110 fails; no other guest depends on it.
- `/etc/modules-load.d/nvidia.conf` loads `nvidia` + `nvidia_uvm`.
- **Retired:** the T400's vfio-pci id binding (`/etc/modprobe.d/vfio.conf`, moved
to `/root/nvidia/vfio.conf.retired-2026-09-24`) and `blacklist nvidia` in
`/etc/modprobe.d/blacklist.conf` (`blacklist nouveau` stays). The `vfio*`
lines in `/etc/modules` were left alone; they bind nothing. The initramfs was
not rebuilt, to keep the boot path untouched; the next kernel update will.
- The installer also enabled NVIDIA's `nvidia-{suspend,hibernate,resume}`
units. They only act on a suspend and are harmless on a server.
- ⚠ **Kernel updates:** DKMS rebuilds the module when a new
`proxmox-headers-6.8.*` arrives (the `proxmox-headers-6.8` meta-package is
installed). Before opting into a different kernel series, install that
series' headers meta-package, or the next boot comes up with no GPU and CT
110 will not start.
- ⚠ **The host module and CT 110's libraries must stay the same version.**
Upgrade both playbooks together (host first).
## Watchdog — hardware, not software