41 lines
2.4 KiB
Markdown
41 lines
2.4 KiB
Markdown
# `[2026-09-24]` esh-pve: VM 102 retired, a 14-day hung VFIO process found, T400 → RTX 2000 Ada
|
||
|
||
**VM 102 `esh-vm-workstation` retired permanently** (Prime: "not the machine I
|
||
want running a virtual windows machine"). Final backup
|
||
`pbs-ana:backup/vm/102/2026-09-10T10:32:20Z` (verification ok) set
|
||
**protected** so prune cannot take it; `qm destroy 102 --purge`.
|
||
|
||
⚠⚠ **Found: `kvm -id 102` had been hung in D state since the 2026-09-10 0332
|
||
vzdump.** vzdump starts a stopped VM (here with the T400 passed through,
|
||
`x-vga=1`) to read its disks; that process never exited. It held **16 GB of
|
||
locked RAM**, the GPU, and two LVs, while `qm status 102` said `stopped` the
|
||
whole time. Same class as the 2026-08-19 hard freeze: VFIO + this box. Only a
|
||
reboot cleared it.
|
||
|
||
**GPU swap, Prime on site, 2119–2137.** Guests shut down cleanly; host powered
|
||
off; the T400 was removed and an **NVIDIA RTX 2000 / 2000E Ada** installed
|
||
(`01:00.0`, `10de:28b0`, subsystem `10de:1871`). Host back 2136:49; leftover
|
||
LVs removed; all guests auto-started; esh-docker-vm NFS, traefik (0×403) and
|
||
AdGuard verified. The new card has **no driver bound**; `/etc/modprobe.d` still
|
||
lists the T400's vfio ids (`10de:1ff2,10de:10fa`), now matching nothing.
|
||
Commits `719e3fa`, `7cbd4c3`.
|
||
|
||
**Decision (Prime, 2026-09-24): the card serves embedding + reranking as an
|
||
LXC with the NVIDIA driver on the host — NOT a VFIO VM.** Rationale: both
|
||
esh-pve hangs this year came from VFIO passthrough; LXC avoids it and pins no
|
||
VM RAM. Accepted cost: a driver on the hypervisor, rebuilt on PVE kernel
|
||
updates, host/container versions must match. **Implementation deferred to the
|
||
next session** (tracking: this entry + `servers/esh-pve/README.md`).
|
||
|
||
**ESH single-route observation:** esh-scale (CT 108, ESH's only mesh subnet
|
||
router) lives on esh-pve, so while esh-pve was down *all* of ESH looked dark
|
||
from outside, including esh-pve-nas and vm-esh-nas, which were up. Untracked by
|
||
operator choice so far — a second ESH route is an idea, not a decision.
|
||
|
||
**vPro:** the copper cable is in the AMT-capable port (**I226-LM**, `enp89s0`,
|
||
MAC `58:47:ca:76:99:32`), verified at 2500 Mb/s; the I226-V (`enp88s0`) is
|
||
not AMT-capable. Both links set back admin-down on the host. **MEBx
|
||
provisioning deferred by Prime** ("some other time"): Ctrl+P at boot, enable
|
||
manageability, static IP, KVM on with User Opt-in None, activate network
|
||
access; MeshCentral should live on ana-docker/nh3-docker, never on esh-pve.
|