Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md
T

41 lines
2.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `[2026-09-24]` esh-pve: VM 102 retired, a 14-day hung VFIO process found, T400 → RTX 2000 Ada
**VM 102 `esh-vm-workstation` retired permanently** (Prime: "not the machine I
want running a virtual windows machine"). Final backup
`pbs-ana:backup/vm/102/2026-09-10T10:32:20Z` (verification ok) set
**protected** so prune cannot take it; `qm destroy 102 --purge`.
⚠⚠ **Found: `kvm -id 102` had been hung in D state since the 2026-09-10 0332
vzdump.** vzdump starts a stopped VM (here with the T400 passed through,
`x-vga=1`) to read its disks; that process never exited. It held **16 GB of
locked RAM**, the GPU, and two LVs, while `qm status 102` said `stopped` the
whole time. Same class as the 2026-08-19 hard freeze: VFIO + this box. Only a
reboot cleared it.
**GPU swap, Prime on site, 2119–2137.** Guests shut down cleanly; host powered
off; the T400 was removed and an **NVIDIA RTX 2000 / 2000E Ada** installed
(`01:00.0`, `10de:28b0`, subsystem `10de:1871`). Host back 2136:49; leftover
LVs removed; all guests auto-started; esh-docker-vm NFS, traefik (0×403) and
AdGuard verified. The new card has **no driver bound**; `/etc/modprobe.d` still
lists the T400's vfio ids (`10de:1ff2,10de:10fa`), now matching nothing.
Commits `719e3fa`, `7cbd4c3`.
**Decision (Prime, 2026-09-24): the card serves embedding + reranking as an
LXC with the NVIDIA driver on the host — NOT a VFIO VM.** Rationale: both
esh-pve hangs this year came from VFIO passthrough; LXC avoids it and pins no
VM RAM. Accepted cost: a driver on the hypervisor, rebuilt on PVE kernel
updates, host/container versions must match. **Implementation deferred to the
next session** (tracking: this entry + `servers/esh-pve/README.md`).
**ESH single-route observation:** esh-scale (CT 108, ESH's only mesh subnet
router) lives on esh-pve, so while esh-pve was down *all* of ESH looked dark
from outside, including esh-pve-nas and vm-esh-nas, which were up. Untracked by
operator choice so far — a second ESH route is an idea, not a decision.
**vPro:** the copper cable is in the AMT-capable port (**I226-LM**, `enp89s0`,
MAC `58:47:ca:76:99:32`), verified at 2500 Mb/s; the I226-V (`enp88s0`) is
not AMT-capable. Both links set back admin-down on the host. **MEBx
provisioning deferred by Prime** ("some other time"): Ctrl+P at boot, enable
manageability, static IP, KVM on with User Opt-in None, activate network
access; MeshCentral should live on ana-docker/nh3-docker, never on esh-pve.