feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1

NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled.

- nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml.
- nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4)
  deployed with HOST_NAME/HOST_IP labels.
- Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min
  0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control
  0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed
  identical within rep spread.
- gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006).
  With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl
  reopen; applied on nh3-pve (one package).
- pve-nvidia-host.yaml: document that the headers meta drags in the newest
  kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot).
- Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage
  nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal.
- nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels,
  AMT cabled but unreachable on the network.

Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
vh
2026-09-25 15:55:33 -07:00
parent 6fa8213c20
commit 5960526c3f
17 changed files with 815 additions and 112 deletions
+19
View File
@@ -65,6 +65,25 @@ vars:
infra_ops_pubkey: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIN+1HBwfXrkfTYWdcnWCjLJ6VLAGC87gxH5h5vKaaA3c infra-ops@pfi-fleet"
steps:
# runc >= 1.2.8 (the CVE-2025-52881 fix; the CT gets 1.5.x from docker-ce)
# re-opens /proc/sys files, and lxc-pve 6.0.0-1's AppArmor profile denies it
# in a nesting CT: EVERY `docker run` fails with "open sysctl
# net.ipv4.ip_unprivileged_port_start file: reopen fd 8: permission denied".
# lxc-pve 6.0.0-2 (Proxmox fix #7006) lifts the restriction when nesting is on.
# Measured 2026-09-25: nh3-pve (PVE 8.4.1, lxc-pve 6.0.0-1) failed the GPU
# verify below; esh-pve (8.4.20, 6.0.0-2) never did. One package, no deps,
# running guests unaffected (the change only relaxes the profile). A CT that
# was already running needs a `pct reboot` to pick it up.
- name: "lxc-pve carries fix #7006 (Docker's runc works in a nesting CT)"
shell: |
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
apt-get install -y -qq --only-upgrade lxc-pve
v=$(dpkg-query -W -f='${Version}' lxc-pve)
dpkg --compare-versions "$v" ge 6.0.0-2 || { echo "lxc-pve is still $v after upgrade" >&2; exit 1; }
when: "dpkg --compare-versions \"$(dpkg-query -W -f='${Version}' lxc-pve)\" lt 6.0.0-2"
- name: Fetch the Debian 12 template
shell: pveam update >/dev/null && pveam download local {{ template }}
creates: /var/lib/vz/template/cache/{{ template }}