feat(esh-ml1): RTX 2000E Ada on esh-pve serves embed + rerank as a LiteLLM failover

esh-pve: NVIDIA 580.178.04 (open modules, DKMS) installed on the host from
NVIDIA's .run and loaded live, no reboot. nvidia-persistenced unit creates the
device nodes before pve-guests; the T400's vfio-pci ids and `blacklist nvidia`
retired. playbooks/esh-pve-nvidia-host.yaml.

esh-ml1: CT 110, unprivileged Debian 12, 10.0.50.80, GPU nodes via devN,
NVIDIA userspace from the same .run (--no-kernel-modules), docker-ce +
nvidia-container-toolkit (no-cgroups). playbooks/esh-ml1-lxc.yaml. Not in the
vzdump job on purpose. DNS esh-ml1.esh.internal.

stacks/embed-rerank: Qwen3-Embedding-0.6B :8001 + bge-reranker-v2-m3 :8013 on
the same vLLM v0.24.0 digest and flags as fv-ml1. Measured parity: embed
cosine FV-vs-ESH median 0.999908 (min 0.999772), inside both self-noise
floors; rerank max |delta| 0.000145 vs floor 0.000181, identical ranking.

litellm: qwen3-embedding and reranker gain an esh-ml1 deployment at order 2
behind fv-ml1 (order 1). Order fallback proven with throwaway groups:
refused primary +0.15 s, host-down primary ~18.7 s per call, dead-only 500.

Also: repaired the DB-only alias reranker-a3-bge-v2-m3, dead since the
fv-ml1 relocation (still named 10.250.50.54); documented the third
unkillable homepage wedge on esh-docker-vm.
This commit is contained in:
vh
2026-09-24 22:38:55 -07:00
parent eb46973051
commit 5402568b76
16 changed files with 1053 additions and 30 deletions
+18 -24
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-24 ~2150 PT (NH3 power outage recovered: pbs-nh3 onboot set, NFS → automount. esh-pve: VM 102 retired, a 14-day hung VFIO process cleared, T400 → RTX 2000 Ada; ⭐ NEXT = provision it as LXC + host driver for embed/rerank. pfi-gx10 AC-restore patch applied but UNVALIDATED — box OFF until Prime's AC pull 09-25. elway root:root fix + fleet ownership audit. Miranda standing order + Prime callsign in CLAUDE.md. task-board mothballed.)_
_Last updated: 2026-09-24 ~2245 PT (⭐ esh-ml1 BUILT: RTX 2000E Ada as CT 110 on esh-pve, host driver 580.178.04 live-loaded with NO reboot, vLLM embed+rerank at parity with fv-ml1, LiteLLM order-2 failover. Dead `reranker-a3-bge-v2-m3` alias repaired. ⚠ homepage wedged unkillably on esh-docker-vm (3rd time) — needs a VM reboot, Prime's call. pfi-gx10 AC-restore still UNVALIDATED — AC pull 09-25.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,30 +115,26 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-09-24 ~2150 PT._
_As of 2026-09-24 ~2245 PT._
### ⭐ NEXT: provision the RTX 2000 Ada on esh-pve as an LXC + host NVIDIA driver
### ✅ esh-ml1 built (2026-09-24) — two follow-ups
Prime's decision (2026-09-24, over a VFIO VM): the card serves **embedding +
reranking offload**. Installed tonight in esh-pve's single slot (`01:00.0`,
`10de:28b0`), no driver bound. Plan shape, not yet started:
The RTX 2000E Ada now serves `qwen3-embedding` + `reranker` from **CT 110
`esh-ml1`** (10.0.50.80) as the **order-2 failover** behind fv-ml1. Host driver
went on live, no esh-pve reboot. Full record → Recent decisions.
- **Parked unless Prime wants it:** a direct ESH→esh-ml1 path for ESH consumers
(survives a mesh outage). Spend logs show NO ESH-side caller in 7 days; the
real users are worldtree-gateway + nevermore.
- **Not wired yet:** esh-ml1 in Homepage `docker.yaml` (needs dockerd tcp/2375)
and Beszel. The ~18.7 s host-down failover penalty is untuned.
1. NVIDIA driver on the PVE host (headers for the running `6.8.12-*-pve`
kernel; DKMS). Blacklist nouveau. Remove the stale T400 vfio ids from
`/etc/modprobe.d` (`10de:1ff2,10de:10fa`).
2. An LXC (Debian 12 template — PVE here rejects Debian 13) with the GPU
device nodes bound in and the SAME userspace driver version as the host;
docker + nvidia-container-toolkit inside.
3. Serve the SAME models the fleet already uses so vectors stay compatible:
`qwen3-embedding` (currently fv-ml1:8001 via LiteLLM) and the rerankers
(`reranker`, `reranker-a3-bge-v2-m3`). Confirm sizes from the live seats
before choosing an engine (TEI vs vLLM).
4. Wire as a LiteLLM failover/local deployment under the SAME model names; ESH
consumers (Open WebUI, Paperless) keep working if FV or the mesh is down.
### ⚠ homepage wedged on esh-docker-vm — needs a VM reboot (Prime's call)
⚠ esh-pve is ESH's only DNS and its only mesh route (esh-scale CT 108 lives
there). Driver work means reboots: do them when an ESH outage is acceptable,
and confirm power-off by the light, not by ping (my path in dies with esh-scale).
3rd unkillable wedge (06-03, 09-18, now 2026-09-24 ~2224 PT): node thread D in
`vm_mmap_pgoff`, then `exit_mmap` after a kill; `docker restart` failed. Rest
of the VM healthy. The reboot blips ESH DNS. Signature + safe diagnosis
commands: `servers/esh-docker-vm/README.md`. Option to move Homepage to
ana-docker surfaced to Prime.
### pfi-gx10 is OFF — AC-pull test 2026-09-25 (Prime)
@@ -163,9 +159,6 @@ a shutdown (stays off by design). Outcomes and revert in
- **infra-hermes owns a daily 0110 job** proving the first High Seat report
(`~/.high-seat/reports/*.jsonl`) lands in a nh3-dev restic snapshot; it
replies to svos-dev (thread `01M37P61Q85KWDVYN0A00P8856`) and pings me.
- **Credentials in auto-memory:** the new global rule says never write one into
a memory file (memory is copied off-box hourly to `vh/claude-memory`). Four
of my memory files still carry the shared LiteLLM key literal — scrub them.
- **Auto-memory `MEMORY.md` is over its 24.4 KB load limit** (tail truncated at
load) — shorten index lines.
- ESH has a single outside route (esh-scale on esh-pve) — noted, untracked.
@@ -173,6 +166,7 @@ a shutdown (stays off by design). Outcomes and revert in
## Recent decisions
- `[2026-09-24]` **esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM `order: 2` failover; dead DB alias `reranker-a3-bge-v2-m3` (pre-relocation IP) repaired.** → `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`
- `[2026-09-24]` **esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank** (implementation deferred to next session, tracked here + `servers/esh-pve/README.md`). → `persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md`
- `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
- `[2026-09-24]` **NH3 power outage recovered** — pbs-nh3 had no `onboot` (set), NFS boot race fixed with automount (`1cbde50`), every other Claude session on nh3-dev died. → `persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md`