feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1
NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled. - nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml. - nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4) deployed with HOST_NAME/HOST_IP labels. - Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min 0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control 0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed identical within rep spread. - gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006). With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl reopen; applied on nh3-pve (one package). - pve-nvidia-host.yaml: document that the headers meta drags in the newest kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot). - Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal. - nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels, AMT cabled but unreachable on the network. Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
@@ -65,6 +65,25 @@ vars:
|
||||
infra_ops_pubkey: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIN+1HBwfXrkfTYWdcnWCjLJ6VLAGC87gxH5h5vKaaA3c infra-ops@pfi-fleet"
|
||||
|
||||
steps:
|
||||
# runc >= 1.2.8 (the CVE-2025-52881 fix; the CT gets 1.5.x from docker-ce)
|
||||
# re-opens /proc/sys files, and lxc-pve 6.0.0-1's AppArmor profile denies it
|
||||
# in a nesting CT: EVERY `docker run` fails with "open sysctl
|
||||
# net.ipv4.ip_unprivileged_port_start file: reopen fd 8: permission denied".
|
||||
# lxc-pve 6.0.0-2 (Proxmox fix #7006) lifts the restriction when nesting is on.
|
||||
# Measured 2026-09-25: nh3-pve (PVE 8.4.1, lxc-pve 6.0.0-1) failed the GPU
|
||||
# verify below; esh-pve (8.4.20, 6.0.0-2) never did. One package, no deps,
|
||||
# running guests unaffected (the change only relaxes the profile). A CT that
|
||||
# was already running needs a `pct reboot` to pick it up.
|
||||
- name: "lxc-pve carries fix #7006 (Docker's runc works in a nesting CT)"
|
||||
shell: |
|
||||
set -e
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq --only-upgrade lxc-pve
|
||||
v=$(dpkg-query -W -f='${Version}' lxc-pve)
|
||||
dpkg --compare-versions "$v" ge 6.0.0-2 || { echo "lxc-pve is still $v after upgrade" >&2; exit 1; }
|
||||
when: "dpkg --compare-versions \"$(dpkg-query -W -f='${Version}' lxc-pve)\" lt 6.0.0-2"
|
||||
|
||||
- name: Fetch the Debian 12 template
|
||||
shell: pveam update >/dev/null && pveam download local {{ template }}
|
||||
creates: /var/lib/vz/template/cache/{{ template }}
|
||||
|
||||
@@ -72,6 +72,12 @@ steps:
|
||||
# DKMS needs the headers for the RUNNING kernel (built now) and the series
|
||||
# meta-package (so each future kernel's headers arrive with it and DKMS
|
||||
# rebuilds on upgrade). nh3-pve had neither, nor dkms or a compiler.
|
||||
# ⚠ The series meta depends on the NEWEST headers, which drags the
|
||||
# proxmox-kernel meta and the newest kernel image in with it. On nh3-pve
|
||||
# (2026-09-25) this installed 6.8.12-43 next to the running -11, and the next
|
||||
# reboot silently booted -43. Expect a kernel change on the next reboot of any
|
||||
# host that was behind, and before that reboot check that `dkms status` lists
|
||||
# nvidia for the kernel that will boot, not only for `uname -r`.
|
||||
- name: Kernel headers (running kernel + series meta), dkms, build tools
|
||||
shell: |
|
||||
set -e
|
||||
|
||||
Reference in New Issue
Block a user