Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md
T

2.4 KiB
Raw Blame History

[2026-09-24] esh-pve: VM 102 retired, a 14-day hung VFIO process found, T400 → RTX 2000 Ada

VM 102 esh-vm-workstation retired permanently (Prime: "not the machine I want running a virtual windows machine"). Final backup pbs-ana:backup/vm/102/2026-09-10T10:32:20Z (verification ok) set protected so prune cannot take it; qm destroy 102 --purge.

⚠⚠ Found: kvm -id 102 had been hung in D state since the 2026-09-10 0332 vzdump. vzdump starts a stopped VM (here with the T400 passed through, x-vga=1) to read its disks; that process never exited. It held 16 GB of locked RAM, the GPU, and two LVs, while qm status 102 said stopped the whole time. Same class as the 2026-08-19 hard freeze: VFIO + this box. Only a reboot cleared it.

GPU swap, Prime on site, 2119–2137. Guests shut down cleanly; host powered off; the T400 was removed and an NVIDIA RTX 2000 / 2000E Ada installed (01:00.0, 10de:28b0, subsystem 10de:1871). Host back 2136:49; leftover LVs removed; all guests auto-started; esh-docker-vm NFS, traefik (0×403) and AdGuard verified. The new card has no driver bound; /etc/modprobe.d still lists the T400's vfio ids (10de:1ff2,10de:10fa), now matching nothing. Commits 719e3fa, 7cbd4c3.

Decision (Prime, 2026-09-24): the card serves embedding + reranking as an LXC with the NVIDIA driver on the host — NOT a VFIO VM. Rationale: both esh-pve hangs this year came from VFIO passthrough; LXC avoids it and pins no VM RAM. Accepted cost: a driver on the hypervisor, rebuilt on PVE kernel updates, host/container versions must match. Implementation deferred to the next session (tracking: this entry + servers/esh-pve/README.md).

ESH single-route observation: esh-scale (CT 108, ESH's only mesh subnet router) lives on esh-pve, so while esh-pve was down all of ESH looked dark from outside, including esh-pve-nas and vm-esh-nas, which were up. Untracked by operator choice so far — a second ESH route is an idea, not a decision.

vPro: the copper cable is in the AMT-capable port (I226-LM, enp89s0, MAC 58:47:ca:76:99:32), verified at 2500 Mb/s; the I226-V (enp88s0) is not AMT-capable. Both links set back admin-down on the host. MEBx provisioning deferred by Prime ("some other time"): Ctrl+P at boot, enable manageability, static IP, KVM on with User Opt-in None, activate network access; MeshCentral should live on ana-docker/nh3-docker, never on esh-pve.