feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1

NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled.

- nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml.
- nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4)
  deployed with HOST_NAME/HOST_IP labels.
- Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min
  0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control
  0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed
  identical within rep spread.
- gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006).
  With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl
  reopen; applied on nh3-pve (one package).
- pve-nvidia-host.yaml: document that the headers meta drags in the newest
  kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot).
- Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage
  nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal.
- nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels,
  AMT cabled but unreachable on the network.

Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
vh
2026-09-25 15:55:33 -07:00
parent 6fa8213c20
commit 5960526c3f
17 changed files with 815 additions and 112 deletions
+6 -4
View File
@@ -5,10 +5,12 @@ GPU LXC for the ESH home lab: **CT 110 on esh-pve**, holding the **NVIDIA RTX
reranking service** — the only backend behind the gateway's `qwen3-embedding`,
`reranker` and `reranker-a3-bge-v2-m3` since 2026-09-25. Built 2026-09-24.
⚠ **Single backend until the second RTX 2000 arrives** (Prime, 2026-09-25). If
esh-ml1, esh-pve, or ESH's mesh route (esh-scale, CT 108) is down, fleet
embeddings and reranking are down: Worldtree recall, nevermore clustering, Open
WebUI RAG.
⚠ **Still the single gateway backend.** The second RTX 2000 is live as
[nh3-ml1](../nh3-ml1/README.md) (2026-09-25, parity-verified: the two hosts cannot
be told apart), but LiteLLM does not route to it yet; that is Prime's call. Until
it does, if esh-ml1, esh-pve, or ESH's mesh route (esh-scale, CT 108) is down,
fleet embeddings and reranking are down: Worldtree recall, nevermore clustering,
Open WebUI RAG.
| | |
|---|---|
+105
View File
@@ -0,0 +1,105 @@
# nh3-ml1
GPU LXC for the NH3 site: **CT 109 on nh3-pve**, holding the **NVIDIA RTX 2000E
Ada** (16 GB, 50 W, `01:00.0`, PCIe gen4 x8). **It is the second embedding and
reranking backend**, the twin of [esh-ml1](../esh-ml1/README.md): same card, same
driver, same TEI image and models. Built 2026-09-25, after the NH3 site visit
turned Secure Boot off on nh3-pve.
⚠ **Not behind the gateway yet.** It serves on its own ports and is monitored, but
LiteLLM still routes `qwen3-embedding` and `reranker` to esh-ml1 alone. Routing is
Prime's call (recommendation: load-share; see below).
| | |
|---|---|
| **IP** | `10.100.50.80/24`, VLAN 50 (`nh3-servers`), gateway `10.100.50.1` (static, outside the UDM's `.150–.249` DHCP pool) |
| **DNS** | `nh3-ml1.nh3.internal` |
| **SSH** | `ssh nh3-ml1` → `infra-ops@10.100.50.80` (NOPASSWD sudo) · from the host: `pct enter 109` |
| **OS** | Debian 12, unprivileged, `nesting=1,keyctl=1` |
| **Size** | 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on `local-zfs` |
| **Boot** | `onboot: 1`, `startup: order=30`, after the site's core guests, so a GPU fault never delays NH3's DNS or mesh route |
| **Backups** | **None, on purpose.** nh3-pve's vzdump job is `all 1`; the playbook added 109 to its `exclude` list. Everything is rebuilt from the playbooks and the stack; models re-download. |
## What it serves
`stacks/embed-rerank` (`/opt/docker/compose/embed-rerank`), **TEI 1.9.4**
(`89-1.9.4`). The live `.env` differs from `.env.example` only in
`HOST_NAME=nh3-ml1` and `HOST_IP=10.100.50.80`, which feed the Homepage labels.
| container | model | port |
|---|---|---|
| `tei-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 (`/v1/embeddings`, `/embed`) |
| `tei-rerank` | `BAAI/bge-reranker-v2-m3` | 8013 (`/rerank`, body `query` + `texts`) |
VRAM ~2.7 GB for both, so ~13 GB is free.
## Parity and speed vs esh-ml1 (2026-09-25 ~1540 PT)
Client on nh3-dev, calling both seats directly. Corpus: 1,120 paragraphs from
this repo's docs plus 6 fixed texts (CJK, code, emoji, a 1-char input, a 6k-char
passage); 50 instruction-format queries; each host embedded everything twice.
| embedding check | result | noise floor / control |
|---|---|---|
| per-text cosine, esh vs nh3 | median 0.999998, min 0.999993 | esh vs esh 0.999998 / 0.999995; nh3 vs nh3 0.999998 / 0.999993 |
| overlap@10, nh3 queries on the esh index | 1.000 | esh rerun 1.000; **positive control** MRL-256 truncation 0.684 |
| overlap@10, nh3 index + nh3 queries | 1.000 | — |
| hit@1 own paragraph | 0.76 (both hosts) | — |
| **negative control**, different texts | cosine median 0.50 | — |
| rerank check (100 queries × 20 docs) | esh vs nh3 | esh vs esh | **positive control** (query cut to 4 words) |
|---|---|---|---|
| top-1 agreement | 1.00 | 1.00 | 0.94 |
| top-5 exact order | 0.98 | 0.97 | 0.09 |
| max score difference | 0.0014 | 0.0020 | 0.97 |
**The two hosts cannot be told apart.** Every esh-vs-nh3 figure sits inside the
esh-vs-esh noise. Sensitivity floor: this method cannot resolve an embedding
cosine gap below ~5×10⁻⁶ or a rerank score difference below ~0.002. An index
built on either host serves queries from the other.
**Speed, on-box**, 3 interleaved reps per host (range across reps):
| workload | nh3-ml1 | esh-ml1 |
|---|---|---|
| embed 1 short query, p50 | 6.60–6.74 ms | 6.97–7.02 ms |
| bulk embed, 64 per request, passages/s | 103.7–105.0 | 105.5–108.1 |
| rerank 20 docs, p50 | 155.8–159.2 ms | 151.8–160.5 ms |
The two hosts run at the same speed. esh-ml1 was also holding the idle reward seat
(~8 GB VRAM, 0% util) during these runs.
**Gateway routing (Prime's call).** The recommendation is **load-share**. The
2026-09-25 rule against load-sharing came from pairing esh-ml1 with the much
faster fv-ml1. These two cards are identical, and a second site removes the
single-host outage the esh-ml1 README warns about.
## How it is built
1. [`playbooks/pve-nvidia-host.yaml`](../../playbooks/pve-nvidia-host.yaml) on
nh3-pve: driver **580.178.04** (open modules, DKMS) plus the
`nvidia-persistenced` unit. It needs **Secure Boot off**, which was turned off
in the BIOS on the 2026-09-25 visit; the pre-flight refuses otherwise.
2. [`playbooks/gpu-lxc.yaml`](../../playbooks/gpu-lxc.yaml) with the
"Run (nh3-ml1)" `--var` line from its header.
⚠ nh3-pve was on **lxc-pve 6.0.0-1**. With it, every `docker run` in the CT
failed with *"open sysctl net.ipv4.ip_unprivileged_port_start file: reopen
fd 8: permission denied"* (runc 1.5 against the old AppArmor profile).
Upgrading lxc-pve alone to 6.0.0-2 (Proxmox fix #7006) and then running
`pct reboot 109` fixed it. The playbook now does the upgrade as its first step.
3. `scripts/deploy-stack.sh nh3-ml1 embed-rerank`, then set `HOST_NAME` /
`HOST_IP` in `.env`, then `docker compose up -d`.
The **driver version lock** and the DKMS/kernel notes in
[esh-ml1's README](../esh-ml1/README.md#-driver-version-lock) apply here unchanged.
nh3-pve runs kernel 6.8.12-43, which the 2026-09-25 headers install pulled in
(see `servers/nh3-pve/README.md`).
## Monitoring and telemetry (wired 2026-09-25)
| layer | what | where |
|---|---|---|
| **Beszel** | NVIDIA agent `henrygd/beszel-agent-nvidia:0.18.7`, `stacks/beszel` + `hosts/nh3-ml1.yaml`, hub system `1feeeq61g4mkqre`; GPU util, VRAM and power are sampled | alerts → infra-ops: Status down 2 m, Disk >85% 5 m, CPU >95% 15 m, Memory >90% 10 m, Temperature >85 °C 5 m |
| **Uptime Kuma** | `Embed — Qwen3 0.6B (TEI, nh3-ml1)` → `:8001/health` (#29); `Rerank — bge-v2-m3 (TEI, nh3-ml1)` → `:8013/health` (#30) | `stacks/uptimekuma/monitors.yaml` |
| **Homepage** | two cards under *AI - Eval & Retrieval*; dockerd on tcp/2375 bound to `10.100.50.80` | `stacks/homepage/conf/docker.yaml` → `nh3-ml1-docker` |
| **Dozzle** | agent `v10.4.1` on `10.100.50.80:7007`, compose dir `dozzle-agent`; added to the hub's `DOZZLE_REMOTE_AGENT` | hub on ana-docker :8088 |
+1
View File
@@ -0,0 +1 @@
infra-ops@10.100.50.80
File diff suppressed because one or more lines are too long
+51 -49
View File
@@ -13,7 +13,7 @@ Proxmox VE hypervisor for the NH3 site (`nh3-vmhost.phasefinal.com`).
- **CPU:** 13th Gen Intel Core i9-13900H
- **RAM:** 62.5 GB
- **Kernel:** `6.8.12-11-pve` (Proxmox 8.x)
- **Kernel:** `6.8.12-43-pve` since 2026-09-25 (was `6.8.12-11`; see the kernel bullet below). PVE `8.4.1`, well behind esh-pve's `8.4.20` (177 packages pending)
- **Storage:** mostly networked — `/mnt/pve/pfi-nh3-nas` (42 TB) mounted from the Synology at `10.100.50.50:/volume1/VMStorage`; ~27 TB used
## What it runs
@@ -33,6 +33,7 @@ after a power loss** is the column that matters in a recovery: a guest listed as
| 103 | nh3-wg | CT | 1 | up |
| 106 | nh3-headscale | CT | 1 | up |
| 107 | nh3-scale | CT | 1 | up (mesh subnet router + fleet egress proxy) |
| 109 | nh3-ml1 (`10.100.50.80`) | CT | 1 (order 30) | up. GPU LXC, second embed/rerank backend (`servers/nh3-ml1/README.md`); needs the NVIDIA module, so it is the one guest a driver fault can stop |
**Power-loss recovery (2026-09-24 outage).** Every guest boots at once, and
nh3-nas is the slowest to serve NFS. NFS clients now mount nh3-nas shares on
@@ -55,12 +56,19 @@ not power back on by itself.
vmbr0 members, so either cage works. **Never drop either port from the bridge**
without checking which one has carrier (`ip -br link`). The bridge carries the
I226-V's MAC `…:96:0d` because it is the first port listed.
- **AMT: NOT wired.** The AMT-capable I226-LM (`enp88s0`, MAC `58:47:ca:76:96:0e`)
has no cable. The ME is present (`/dev/mei0`, "AMT SOL Redirection" 00:16.3), but
MEBx provisioning status is unknown. To wire it: cable the LM port, then at boot
press Ctrl+P → set the MEBx password, enable manageability, set network (static
or DHCP), KVM on, User Opt-in = None, activate network access. MEBx can be driven
remotely through the NanoKVM below.
- **AMT: cabled and enabled, but NOT reachable (2026-09-25 visit).** The I226-LM
(`enp88s0`, MAC `58:47:ca:76:96:0e`) now has a 1 Gb link (measured at 1552 by
bringing the port up unbridged for a few seconds). Prime reports AMT enabled in
MEBx. **Nothing answers on the network, though.** The UDM has no client or lease
for `…:96:0e`, and no host on `10.100.{0,10,50,250}.0/24` has 16992 or 16993
open. The sweep was checked against known-open ports and does see them. The
cable is not on the USW Pro 24, whose up ports are 19, 22, 23 and 26, all
accounted for. So it is on another switch, most likely nh3-sw1
(`10.100.250.2`, no infra-ops access). Likely causes: MEBx "Activate Network
Access" was not done, a static IP outside those subnets, or a switch port on a
VLAN that gets no DHCP. MEBx menu: Ctrl+P at boot → password, manageability on,
network (static or DHCP), KVM on, User Opt-in = None, activate network access.
**Keep the NanoKVM here until AMT KVM is confirmed.**
- **Console OOB exists: a Sipeed NanoKVM** is attached (USB `3346:1009` on the host;
web UI **`https://10.100.250.171`**, switch port 23, nh3-mgmt). It gives video and
keyboard, so BIOS, MEBx and a host that booted without network are all reachable
@@ -80,48 +88,42 @@ not power back on by itself.
- **`enp88s0` (the AMT port) is no longer a vmbr0 bridge port** (file edited
2026-09-25, effective next boot). STP is off, so bridging a second cabled uplink
into the same L2 would loop the site LAN.
- **GPU installed 2026-09-25: RTX 2000E Ada at `01:00.0`** (`10de:28b0`). The pins
held: the X710 moved to bus 03 and every NIC kept its name. No NVIDIA driver yet,
so `nouveau` binds it (`gsp ctor failed: -2` is expected without GSP firmware).
⚠ **After the install the iGPU is gone from the PCI bus**: `00:02.0` enumerated on
the 2026-09-24 boot and is absent now, `/dev/dri` does not exist, and the RTX is
`boot_vga=1`. The BIOS's Auto primary display picked the PCIe card and hid the
iGPU, so the NanoKVM's iGPU-HDMI capture has no source. Fix: BIOS → Primary
Display = IGFX (or enable iGPU Multi-Monitor). Reaching the BIOS now needs a
display on the RTX's mini-DP, for example the NanoKVM through a mini-DP→HDMI
adapter. Verify with `lspci | grep 00:02.0` and `boot_vga` on `00:02.0`.
- ⚠ **Secure Boot is ON here** (`mokutil --sb-state`, lockdown `integrity`;
enrolled MOK = the Proxmox Secure Boot CA). esh-pve, the same model, has it
OFF. Any DKMS-built module, the NVIDIA driver included, is refused until its
key is enrolled through MokManager at boot or Secure Boot is turned off in the
BIOS. Both need the console, and the console is blind (below). The 2026-09-25
NVIDIA install failed on this and rolled itself back. What it left, all
harmless: headers `6.8.12-11` plus the series meta, dkms and build-essential;
nouveau blacklisted and unloaded; the `.run` staged in `/root/nvidia`.
- **OOB decision (Prime via Miranda, 2026-09-25): HOLD. nh3-pve stays console-blind
until the next NH3 site visit.** No BIOS change and no reboot until then. The
accepted risk: a boot without network means a site trip. **Target end state:**
the NanoKVM moves to pfi-gx10, and this MS-01 gets out-of-band access through its
own vPro/AMT (cable the I226-LM `enp88s0` and provision MEBx; see the AMT bullet
above).
⚠ **Do the BIOS fix on that same visit.** AMT's KVM redirection captures only
the Intel iGPU's framebuffer, so with the iGPU hidden as it is now, AMT KVM would
be just as blind as the NanoKVM. Order: set Primary Display = IGFX, provision
MEBx, confirm AMT KVM shows the console, and only then move the NanoKVM to the
gx10.
**Reaching the BIOS** (2026-09-25 facts). The NanoKVM input is HDMI, the RTX has
mini-DP only, and the adapters on hand are mini-DP→DP only. The paths:
(1) an **active** mini-DP→HDMI adapter to feed the NanoKVM, since a passive one
depends on DP++; (2) a DisplayPort monitor plus a USB keyboard at the rack;
(3) pull the card, boot on the iGPU, set IGFX explicitly (not Auto), then refit
the card. The NIC pins make a card-out boot safe.
`systemctl reboot --firmware-setup` works here (`OsIndicationsSupported` bit 0
set), so nobody has to catch the Del key at POST.
**An OS-side patch is not possible:** the AMI `Setup`/`SaSetup` variables are not
runtime-visible here, unlike on the gx10.
**Known-good reference: esh-pve.** Same MS-01, same BIOS `AHWSA.1.17`, same RTX
2000E, and its `00:02.0` is present with `boot_vga=1`. The target state works on
this hardware.
- **GPU: RTX 2000E Ada at `01:00.0`** (`10de:28b0`), installed 2026-09-25. The
NIC pins held: the X710 moved to bus 03 and every NIC kept its name. **NVIDIA
580.178.04** (open modules, DKMS) has been on the host since 2026-09-25 at 1527
(`playbooks/pve-nvidia-host.yaml`), with `nvidia-persistenced` ordered before
`pve-guests`. It serves CT 109 nh3-ml1.
- **iGPU restored (2026-09-25 visit).** With the card in, the BIOS's Auto primary
display had hidden the iGPU. That left the NanoKVM (iGPU HDMI) blind and would
have blinded AMT KVM too, since AMT captures only the iGPU. It was set on the
visit, and since the 1523 boot `00:02.0` is back with `boot_vga=1` and i915
loaded. The NanoKVM should have video again (not checked from here).
`systemctl reboot --firmware-setup` works (`OsIndicationsSupported` bit 0), so
nobody has to catch Del at POST. The AMI `Setup` variables are not
runtime-visible, so there is no OS-side BIOS patch. esh-pve, the same MS-01 and
BIOS `AHWSA.1.17`, is the known-good reference.
- **Secure Boot: OFF since the 2026-09-25 visit** (`mokutil --sb-state`: disabled),
which matches esh-pve. While it was ON (lockdown `integrity`), the DKMS NVIDIA
module was refused and the first install rolled itself back. The playbook's
pre-flight refuses if it is ever turned back on without an enrolled DKMS MOK.
- ⚠ **Kernel jumped `6.8.12-11` → `6.8.12-43` at the visit reboot, pulled in by
our own playbook.** At 1419 the headers step ran `apt-get install
proxmox-headers-6.8`. That upgraded the `proxmox-kernel-6.8` meta and installed
`proxmox-kernel-6.8.12-43-pve-signed`. Nobody chose the new kernel; it booted
because it was the newest.
Side effect: **every -43 boot oopses in Bluetooth** (`btmtk_usb_hci_wmt_sync` →
NULL deref in `hci_power_on`, the MS-01's MediaTek BT; taint `D`). It hit on all
three -43 boots and on none of the -11 boots. esh-pve on `-42` shows the same
oops and has run fine, so it is benign so far: only the BT worker dies. The fix
is to blacklist `btusb` on both hypervisors. That is not done, because it only
takes effect at the next boot.
- **lxc-pve 6.0.0-1 → 6.0.0-2** (2026-09-25 1533, that one package only). This is
Proxmox fix #7006. Without it, runc 1.5 inside a nesting CT fails every
`docker run`. `playbooks/gpu-lxc.yaml` now upgrades it first.
- **OOB plan status.** Prime ruled on 2026-09-25 via Miranda to HOLD console-blind
until the site visit. Target: the NanoKVM moves to pfi-gx10, and this MS-01 uses
its own AMT. The visit did IGFX, turned SB off and cabled plus enabled AMT.
**Still open:** AMT is not on the network (above), so the NanoKVM stays here.
## Refresh state