2.4 KiB
[2026-09-24] esh-pve: VM 102 retired, a 14-day hung VFIO process found, T400 → RTX 2000 Ada
VM 102 esh-vm-workstation retired permanently (Prime: "not the machine I
want running a virtual windows machine"). Final backup
pbs-ana:backup/vm/102/2026-09-10T10:32:20Z (verification ok) set
protected so prune cannot take it; qm destroy 102 --purge.
⚠⚠ Found: kvm -id 102 had been hung in D state since the 2026-09-10 0332
vzdump. vzdump starts a stopped VM (here with the T400 passed through,
x-vga=1) to read its disks; that process never exited. It held 16 GB of
locked RAM, the GPU, and two LVs, while qm status 102 said stopped the
whole time. Same class as the 2026-08-19 hard freeze: VFIO + this box. Only a
reboot cleared it.
GPU swap, Prime on site, 2119–2137. Guests shut down cleanly; host powered
off; the T400 was removed and an NVIDIA RTX 2000 / 2000E Ada installed
(01:00.0, 10de:28b0, subsystem 10de:1871). Host back 2136:49; leftover
LVs removed; all guests auto-started; esh-docker-vm NFS, traefik (0×403) and
AdGuard verified. The new card has no driver bound; /etc/modprobe.d still
lists the T400's vfio ids (10de:1ff2,10de:10fa), now matching nothing.
Commits 719e3fa, 7cbd4c3.
Decision (Prime, 2026-09-24): the card serves embedding + reranking as an
LXC with the NVIDIA driver on the host — NOT a VFIO VM. Rationale: both
esh-pve hangs this year came from VFIO passthrough; LXC avoids it and pins no
VM RAM. Accepted cost: a driver on the hypervisor, rebuilt on PVE kernel
updates, host/container versions must match. Implementation deferred to the
next session (tracking: this entry + servers/esh-pve/README.md).
ESH single-route observation: esh-scale (CT 108, ESH's only mesh subnet router) lives on esh-pve, so while esh-pve was down all of ESH looked dark from outside, including esh-pve-nas and vm-esh-nas, which were up. Untracked by operator choice so far — a second ESH route is an idea, not a decision.
vPro: the copper cable is in the AMT-capable port (I226-LM, enp89s0,
MAC 58:47:ca:76:99:32), verified at 2500 Mb/s; the I226-V (enp88s0) is
not AMT-capable. Both links set back admin-down on the host. MEBx
provisioning deferred by Prime ("some other time"): Ctrl+P at boot, enable
manageability, static IP, KVM on with User Opt-in None, activate network
access; MeshCentral should live on ana-docker/nh3-docker, never on esh-pve.