memory: snapshot — nh3-pve shutting down for the NH3 site visit (post-boot checklist), nh3-ml1 blocked on Secure Boot, gx10 AC-restore validated
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
# `[2026-09-25]` nh3-ml1 — second embed/rerank GPU LXC on nh3-pve (DEFERRED: Secure Boot)
|
||||
|
||||
**Ask (Prime, ~1410 PT):** "stand up an nh3-ml1 that takes advantage of that GPU"
|
||||
(the RTX 2000E Ada installed in nh3-pve that morning, `01:00.0`).
|
||||
|
||||
## Decisions taken on Prime's behalf (he said "clear to proceed" at ~1418)
|
||||
|
||||
- **Purpose: the second TEI embed/rerank backend.** esh-ml1's README had been
|
||||
waiting on "the second RTX 2000" for failover. No reward-seat copy is planned,
|
||||
because that seat has no working consumer.
|
||||
- **Placement: CT 109, `10.100.50.80/24`, VLAN 50, gw `.1`, rootfs on `local-zfs`,
|
||||
6 cores / 16 GB / 80 GB, startup order 30.** The address came from the UDM
|
||||
(`nh3-servers` DHCP pool `.150–.249`; in use `.40 .42 .45 .46 .50 .60 .71 .90`).
|
||||
`.80` mirrors esh-ml1. nh3-pve had 40 of 62 GB RAM free.
|
||||
- **Not backed up**, like esh-ml1. nh3-pve's vzdump job is `all 1` (exclude 107),
|
||||
so `gpu-lxc.yaml` adds the CT to that job's `exclude` list. The guard and the
|
||||
step were tested against sample configs: one positive case, plus negatives for
|
||||
already excluded, a vmid-list job, a non-vzdump `all 1` stanza and a missing
|
||||
file.
|
||||
- **Playbooks made host-generic (`bc278d4`):** `esh-pve-nvidia-host` →
|
||||
`pve-nvidia-host`, which gained headers/dkms/build-essential, a nouveau
|
||||
blacklist and an unload that refuses if nouveau has bound a device.
|
||||
`esh-ml1-lxc` → `gpu-lxc`, whose host vars have **no defaults**; elway aborts on
|
||||
an undefined `{{ var }}` (tested). Per-host `--var` lines are in its header.
|
||||
- **`stacks/embed-rerank` Homepage labels** take `HOST_NAME` / `HOST_IP`. The
|
||||
defaults render byte-identical to esh-ml1's live labels, so nothing is
|
||||
recreated there.
|
||||
|
||||
## What happened
|
||||
|
||||
1. The first `pve-nvidia-host` run died on host-key verification for
|
||||
`nh3-pve.nh3.internal`. Every `when:` got ssh rc 255 and SKIPPED, and verify
|
||||
failed. I pinned the name after confirming its ed25519 key equals the known
|
||||
IP entry.
|
||||
2. The second run installed headers 6.8.12-11 + `proxmox-headers-6.8` + dkms +
|
||||
build-essential, blacklisted and unloaded nouveau (it had failed its probe
|
||||
with `gsp ctor failed: -2` and was bound to nothing), then **the DKMS module
|
||||
was refused at load: Secure Boot is ON** (`mokutil --sb-state` enabled;
|
||||
lockdown `integrity`; the only MOK is the Proxmox Secure Boot CA). esh-pve,
|
||||
the same MS-01 and BIOS, has Secure Boot OFF. The installer rolled back: no
|
||||
dkms entry, no `nvidia-smi`. All guests were unaffected.
|
||||
3. `7ddd116` added a pre-flight that refuses when Secure Boot is on without an
|
||||
enrolled `/var/lib/dkms/mok.pub`. Tested: it refuses on nh3-pve and passes on
|
||||
esh-pve.
|
||||
|
||||
## Open calls (Prime's)
|
||||
|
||||
- **Secure Boot:** (1) turn it off in the BIOS on the NH3 visit (recommended; one
|
||||
more toggle in the menu he is already in for IGFX, and it matches esh-pve);
|
||||
(2) keep it and enroll a DKMS MOK (prep remotely, then answer MokManager at
|
||||
boot; DKMS 3.x then signs rebuilds itself); (3) a VFIO VM (works without a
|
||||
visit, but reverses his 09-24 "LXC + host driver, not VFIO" rule, and the RTX is
|
||||
currently the boot VGA). Options 1 and 2 need the console at boot, which is
|
||||
blind until IGFX is set, so the visit carries both.
|
||||
- **Gateway routing** once parity passes: recommendation load-share. The two
|
||||
cards are identical, and the old "don't load-share" rule came from pairing
|
||||
esh-ml1 with the much faster fv-ml1.
|
||||
|
||||
## Resume
|
||||
|
||||
The Current-state checklist, step 9. Nothing host-side exists yet for CT 109.
|
||||
Reference in New Issue
Block a user