refactor(playbooks): host-generic GPU host + GPU LXC playbooks for nh3-ml1

- esh-pve-nvidia-host -> pve-nvidia-host: headers/dkms/build-essential step,
  nouveau blacklist + guarded unload (refuses if nouveau bound a device)
- esh-ml1-lxc -> gpu-lxc: host vars have no defaults (elway aborts on undefined),
  rootfs storage/startup order parameterized, CT kept out of all-guests vzdump jobs
- embed-rerank: Homepage labels take HOST_NAME/HOST_IP, defaults = esh-ml1
This commit is contained in:
vh
2026-09-25 14:17:31 -07:00
parent cdd7605e89
commit bc278d4ba8
10 changed files with 170 additions and 72 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ and reranking ([`servers/esh-ml1/README.md`](../esh-ml1/README.md)).
- **Driver 580.178.04, open kernel modules, DKMS**, from NVIDIA's `.run`
(`/root/nvidia/`). Applied by
[`playbooks/esh-pve-nvidia-host.yaml`](../../playbooks/esh-pve-nvidia-host.yaml)
[`playbooks/pve-nvidia-host.yaml`](../../playbooks/pve-nvidia-host.yaml)
**live, with no reboot**: nouveau was never loaded and nothing held the card.
- **`nvidia-persistenced.service`** (ours, in `/etc/systemd/system`) runs
`nvidia-modprobe -c0 -u` and the persistence daemon **before