feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1
NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled. - nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml. - nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4) deployed with HOST_NAME/HOST_IP labels. - Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min 0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control 0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed identical within rep spread. - gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006). With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl reopen; applied on nh3-pve (one package). - pve-nvidia-host.yaml: document that the headers meta drags in the newest kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot). - Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal. - nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels, AMT cabled but unreachable on the network. Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
@@ -5,10 +5,12 @@ GPU LXC for the ESH home lab: **CT 110 on esh-pve**, holding the **NVIDIA RTX
|
||||
reranking service** — the only backend behind the gateway's `qwen3-embedding`,
|
||||
`reranker` and `reranker-a3-bge-v2-m3` since 2026-09-25. Built 2026-09-24.
|
||||
|
||||
⚠ **Single backend until the second RTX 2000 arrives** (Prime, 2026-09-25). If
|
||||
esh-ml1, esh-pve, or ESH's mesh route (esh-scale, CT 108) is down, fleet
|
||||
embeddings and reranking are down: Worldtree recall, nevermore clustering, Open
|
||||
WebUI RAG.
|
||||
⚠ **Still the single gateway backend.** The second RTX 2000 is live as
|
||||
[nh3-ml1](../nh3-ml1/README.md) (2026-09-25, parity-verified: the two hosts cannot
|
||||
be told apart), but LiteLLM does not route to it yet; that is Prime's call. Until
|
||||
it does, if esh-ml1, esh-pve, or ESH's mesh route (esh-scale, CT 108) is down,
|
||||
fleet embeddings and reranking are down: Worldtree recall, nevermore clustering,
|
||||
Open WebUI RAG.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
# nh3-ml1
|
||||
|
||||
GPU LXC for the NH3 site: **CT 109 on nh3-pve**, holding the **NVIDIA RTX 2000E
|
||||
Ada** (16 GB, 50 W, `01:00.0`, PCIe gen4 x8). **It is the second embedding and
|
||||
reranking backend**, the twin of [esh-ml1](../esh-ml1/README.md): same card, same
|
||||
driver, same TEI image and models. Built 2026-09-25, after the NH3 site visit
|
||||
turned Secure Boot off on nh3-pve.
|
||||
|
||||
⚠ **Not behind the gateway yet.** It serves on its own ports and is monitored, but
|
||||
LiteLLM still routes `qwen3-embedding` and `reranker` to esh-ml1 alone. Routing is
|
||||
Prime's call (recommendation: load-share; see below).
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **IP** | `10.100.50.80/24`, VLAN 50 (`nh3-servers`), gateway `10.100.50.1` (static, outside the UDM's `.150–.249` DHCP pool) |
|
||||
| **DNS** | `nh3-ml1.nh3.internal` |
|
||||
| **SSH** | `ssh nh3-ml1` → `infra-ops@10.100.50.80` (NOPASSWD sudo) · from the host: `pct enter 109` |
|
||||
| **OS** | Debian 12, unprivileged, `nesting=1,keyctl=1` |
|
||||
| **Size** | 6 cores, 16 GB RAM + 2 GB swap, 80 GB rootfs on `local-zfs` |
|
||||
| **Boot** | `onboot: 1`, `startup: order=30`, after the site's core guests, so a GPU fault never delays NH3's DNS or mesh route |
|
||||
| **Backups** | **None, on purpose.** nh3-pve's vzdump job is `all 1`; the playbook added 109 to its `exclude` list. Everything is rebuilt from the playbooks and the stack; models re-download. |
|
||||
|
||||
## What it serves
|
||||
|
||||
`stacks/embed-rerank` (`/opt/docker/compose/embed-rerank`), **TEI 1.9.4**
|
||||
(`89-1.9.4`). The live `.env` differs from `.env.example` only in
|
||||
`HOST_NAME=nh3-ml1` and `HOST_IP=10.100.50.80`, which feed the Homepage labels.
|
||||
|
||||
| container | model | port |
|
||||
|---|---|---|
|
||||
| `tei-embed` | `Qwen/Qwen3-Embedding-0.6B` | 8001 (`/v1/embeddings`, `/embed`) |
|
||||
| `tei-rerank` | `BAAI/bge-reranker-v2-m3` | 8013 (`/rerank`, body `query` + `texts`) |
|
||||
|
||||
VRAM ~2.7 GB for both, so ~13 GB is free.
|
||||
|
||||
## Parity and speed vs esh-ml1 (2026-09-25 ~1540 PT)
|
||||
|
||||
Client on nh3-dev, calling both seats directly. Corpus: 1,120 paragraphs from
|
||||
this repo's docs plus 6 fixed texts (CJK, code, emoji, a 1-char input, a 6k-char
|
||||
passage); 50 instruction-format queries; each host embedded everything twice.
|
||||
|
||||
| embedding check | result | noise floor / control |
|
||||
|---|---|---|
|
||||
| per-text cosine, esh vs nh3 | median 0.999998, min 0.999993 | esh vs esh 0.999998 / 0.999995; nh3 vs nh3 0.999998 / 0.999993 |
|
||||
| overlap@10, nh3 queries on the esh index | 1.000 | esh rerun 1.000; **positive control** MRL-256 truncation 0.684 |
|
||||
| overlap@10, nh3 index + nh3 queries | 1.000 | — |
|
||||
| hit@1 own paragraph | 0.76 (both hosts) | — |
|
||||
| **negative control**, different texts | cosine median 0.50 | — |
|
||||
|
||||
| rerank check (100 queries × 20 docs) | esh vs nh3 | esh vs esh | **positive control** (query cut to 4 words) |
|
||||
|---|---|---|---|
|
||||
| top-1 agreement | 1.00 | 1.00 | 0.94 |
|
||||
| top-5 exact order | 0.98 | 0.97 | 0.09 |
|
||||
| max score difference | 0.0014 | 0.0020 | 0.97 |
|
||||
|
||||
**The two hosts cannot be told apart.** Every esh-vs-nh3 figure sits inside the
|
||||
esh-vs-esh noise. Sensitivity floor: this method cannot resolve an embedding
|
||||
cosine gap below ~5×10⁻⁶ or a rerank score difference below ~0.002. An index
|
||||
built on either host serves queries from the other.
|
||||
|
||||
**Speed, on-box**, 3 interleaved reps per host (range across reps):
|
||||
|
||||
| workload | nh3-ml1 | esh-ml1 |
|
||||
|---|---|---|
|
||||
| embed 1 short query, p50 | 6.60–6.74 ms | 6.97–7.02 ms |
|
||||
| bulk embed, 64 per request, passages/s | 103.7–105.0 | 105.5–108.1 |
|
||||
| rerank 20 docs, p50 | 155.8–159.2 ms | 151.8–160.5 ms |
|
||||
|
||||
The two hosts run at the same speed. esh-ml1 was also holding the idle reward seat
|
||||
(~8 GB VRAM, 0% util) during these runs.
|
||||
|
||||
**Gateway routing (Prime's call).** The recommendation is **load-share**. The
|
||||
2026-09-25 rule against load-sharing came from pairing esh-ml1 with the much
|
||||
faster fv-ml1. These two cards are identical, and a second site removes the
|
||||
single-host outage the esh-ml1 README warns about.
|
||||
|
||||
## How it is built
|
||||
|
||||
1. [`playbooks/pve-nvidia-host.yaml`](../../playbooks/pve-nvidia-host.yaml) on
|
||||
nh3-pve: driver **580.178.04** (open modules, DKMS) plus the
|
||||
`nvidia-persistenced` unit. It needs **Secure Boot off**, which was turned off
|
||||
in the BIOS on the 2026-09-25 visit; the pre-flight refuses otherwise.
|
||||
2. [`playbooks/gpu-lxc.yaml`](../../playbooks/gpu-lxc.yaml) with the
|
||||
"Run (nh3-ml1)" `--var` line from its header.
|
||||
⚠ nh3-pve was on **lxc-pve 6.0.0-1**. With it, every `docker run` in the CT
|
||||
failed with *"open sysctl net.ipv4.ip_unprivileged_port_start file: reopen
|
||||
fd 8: permission denied"* (runc 1.5 against the old AppArmor profile).
|
||||
Upgrading lxc-pve alone to 6.0.0-2 (Proxmox fix #7006) and then running
|
||||
`pct reboot 109` fixed it. The playbook now does the upgrade as its first step.
|
||||
3. `scripts/deploy-stack.sh nh3-ml1 embed-rerank`, then set `HOST_NAME` /
|
||||
`HOST_IP` in `.env`, then `docker compose up -d`.
|
||||
|
||||
The **driver version lock** and the DKMS/kernel notes in
|
||||
[esh-ml1's README](../esh-ml1/README.md#-driver-version-lock) apply here unchanged.
|
||||
nh3-pve runs kernel 6.8.12-43, which the 2026-09-25 headers install pulled in
|
||||
(see `servers/nh3-pve/README.md`).
|
||||
|
||||
## Monitoring and telemetry (wired 2026-09-25)
|
||||
|
||||
| layer | what | where |
|
||||
|---|---|---|
|
||||
| **Beszel** | NVIDIA agent `henrygd/beszel-agent-nvidia:0.18.7`, `stacks/beszel` + `hosts/nh3-ml1.yaml`, hub system `1feeeq61g4mkqre`; GPU util, VRAM and power are sampled | alerts → infra-ops: Status down 2 m, Disk >85% 5 m, CPU >95% 15 m, Memory >90% 10 m, Temperature >85 °C 5 m |
|
||||
| **Uptime Kuma** | `Embed — Qwen3 0.6B (TEI, nh3-ml1)` → `:8001/health` (#29); `Rerank — bge-v2-m3 (TEI, nh3-ml1)` → `:8013/health` (#30) | `stacks/uptimekuma/monitors.yaml` |
|
||||
| **Homepage** | two cards under *AI - Eval & Retrieval*; dockerd on tcp/2375 bound to `10.100.50.80` | `stacks/homepage/conf/docker.yaml` → `nh3-ml1-docker` |
|
||||
| **Dozzle** | agent `v10.4.1` on `10.100.50.80:7007`, compose dir `dozzle-agent`; added to the hub's `DOZZLE_REMOTE_AGENT` | hub on ana-docker :8088 |
|
||||
@@ -0,0 +1 @@
|
||||
infra-ops@10.100.50.80
|
||||
File diff suppressed because one or more lines are too long
+51
-49
@@ -13,7 +13,7 @@ Proxmox VE hypervisor for the NH3 site (`nh3-vmhost.phasefinal.com`).
|
||||
|
||||
- **CPU:** 13th Gen Intel Core i9-13900H
|
||||
- **RAM:** 62.5 GB
|
||||
- **Kernel:** `6.8.12-11-pve` (Proxmox 8.x)
|
||||
- **Kernel:** `6.8.12-43-pve` since 2026-09-25 (was `6.8.12-11`; see the kernel bullet below). PVE `8.4.1`, well behind esh-pve's `8.4.20` (177 packages pending)
|
||||
- **Storage:** mostly networked — `/mnt/pve/pfi-nh3-nas` (42 TB) mounted from the Synology at `10.100.50.50:/volume1/VMStorage`; ~27 TB used
|
||||
|
||||
## What it runs
|
||||
@@ -33,6 +33,7 @@ after a power loss** is the column that matters in a recovery: a guest listed as
|
||||
| 103 | nh3-wg | CT | 1 | up |
|
||||
| 106 | nh3-headscale | CT | 1 | up |
|
||||
| 107 | nh3-scale | CT | 1 | up (mesh subnet router + fleet egress proxy) |
|
||||
| 109 | nh3-ml1 (`10.100.50.80`) | CT | 1 (order 30) | up. GPU LXC, second embed/rerank backend (`servers/nh3-ml1/README.md`); needs the NVIDIA module, so it is the one guest a driver fault can stop |
|
||||
|
||||
**Power-loss recovery (2026-09-24 outage).** Every guest boots at once, and
|
||||
nh3-nas is the slowest to serve NFS. NFS clients now mount nh3-nas shares on
|
||||
@@ -55,12 +56,19 @@ not power back on by itself.
|
||||
vmbr0 members, so either cage works. **Never drop either port from the bridge**
|
||||
without checking which one has carrier (`ip -br link`). The bridge carries the
|
||||
I226-V's MAC `…:96:0d` because it is the first port listed.
|
||||
- **AMT: NOT wired.** The AMT-capable I226-LM (`enp88s0`, MAC `58:47:ca:76:96:0e`)
|
||||
has no cable. The ME is present (`/dev/mei0`, "AMT SOL Redirection" 00:16.3), but
|
||||
MEBx provisioning status is unknown. To wire it: cable the LM port, then at boot
|
||||
press Ctrl+P → set the MEBx password, enable manageability, set network (static
|
||||
or DHCP), KVM on, User Opt-in = None, activate network access. MEBx can be driven
|
||||
remotely through the NanoKVM below.
|
||||
- **AMT: cabled and enabled, but NOT reachable (2026-09-25 visit).** The I226-LM
|
||||
(`enp88s0`, MAC `58:47:ca:76:96:0e`) now has a 1 Gb link (measured at 1552 by
|
||||
bringing the port up unbridged for a few seconds). Prime reports AMT enabled in
|
||||
MEBx. **Nothing answers on the network, though.** The UDM has no client or lease
|
||||
for `…:96:0e`, and no host on `10.100.{0,10,50,250}.0/24` has 16992 or 16993
|
||||
open. The sweep was checked against known-open ports and does see them. The
|
||||
cable is not on the USW Pro 24, whose up ports are 19, 22, 23 and 26, all
|
||||
accounted for. So it is on another switch, most likely nh3-sw1
|
||||
(`10.100.250.2`, no infra-ops access). Likely causes: MEBx "Activate Network
|
||||
Access" was not done, a static IP outside those subnets, or a switch port on a
|
||||
VLAN that gets no DHCP. MEBx menu: Ctrl+P at boot → password, manageability on,
|
||||
network (static or DHCP), KVM on, User Opt-in = None, activate network access.
|
||||
**Keep the NanoKVM here until AMT KVM is confirmed.**
|
||||
- **Console OOB exists: a Sipeed NanoKVM** is attached (USB `3346:1009` on the host;
|
||||
web UI **`https://10.100.250.171`**, switch port 23, nh3-mgmt). It gives video and
|
||||
keyboard, so BIOS, MEBx and a host that booted without network are all reachable
|
||||
@@ -80,48 +88,42 @@ not power back on by itself.
|
||||
- **`enp88s0` (the AMT port) is no longer a vmbr0 bridge port** (file edited
|
||||
2026-09-25, effective next boot). STP is off, so bridging a second cabled uplink
|
||||
into the same L2 would loop the site LAN.
|
||||
- **GPU installed 2026-09-25: RTX 2000E Ada at `01:00.0`** (`10de:28b0`). The pins
|
||||
held: the X710 moved to bus 03 and every NIC kept its name. No NVIDIA driver yet,
|
||||
so `nouveau` binds it (`gsp ctor failed: -2` is expected without GSP firmware).
|
||||
⚠ **After the install the iGPU is gone from the PCI bus**: `00:02.0` enumerated on
|
||||
the 2026-09-24 boot and is absent now, `/dev/dri` does not exist, and the RTX is
|
||||
`boot_vga=1`. The BIOS's Auto primary display picked the PCIe card and hid the
|
||||
iGPU, so the NanoKVM's iGPU-HDMI capture has no source. Fix: BIOS → Primary
|
||||
Display = IGFX (or enable iGPU Multi-Monitor). Reaching the BIOS now needs a
|
||||
display on the RTX's mini-DP, for example the NanoKVM through a mini-DP→HDMI
|
||||
adapter. Verify with `lspci | grep 00:02.0` and `boot_vga` on `00:02.0`.
|
||||
- ⚠ **Secure Boot is ON here** (`mokutil --sb-state`, lockdown `integrity`;
|
||||
enrolled MOK = the Proxmox Secure Boot CA). esh-pve, the same model, has it
|
||||
OFF. Any DKMS-built module, the NVIDIA driver included, is refused until its
|
||||
key is enrolled through MokManager at boot or Secure Boot is turned off in the
|
||||
BIOS. Both need the console, and the console is blind (below). The 2026-09-25
|
||||
NVIDIA install failed on this and rolled itself back. What it left, all
|
||||
harmless: headers `6.8.12-11` plus the series meta, dkms and build-essential;
|
||||
nouveau blacklisted and unloaded; the `.run` staged in `/root/nvidia`.
|
||||
- **OOB decision (Prime via Miranda, 2026-09-25): HOLD. nh3-pve stays console-blind
|
||||
until the next NH3 site visit.** No BIOS change and no reboot until then. The
|
||||
accepted risk: a boot without network means a site trip. **Target end state:**
|
||||
the NanoKVM moves to pfi-gx10, and this MS-01 gets out-of-band access through its
|
||||
own vPro/AMT (cable the I226-LM `enp88s0` and provision MEBx; see the AMT bullet
|
||||
above).
|
||||
⚠ **Do the BIOS fix on that same visit.** AMT's KVM redirection captures only
|
||||
the Intel iGPU's framebuffer, so with the iGPU hidden as it is now, AMT KVM would
|
||||
be just as blind as the NanoKVM. Order: set Primary Display = IGFX, provision
|
||||
MEBx, confirm AMT KVM shows the console, and only then move the NanoKVM to the
|
||||
gx10.
|
||||
**Reaching the BIOS** (2026-09-25 facts). The NanoKVM input is HDMI, the RTX has
|
||||
mini-DP only, and the adapters on hand are mini-DP→DP only. The paths:
|
||||
(1) an **active** mini-DP→HDMI adapter to feed the NanoKVM, since a passive one
|
||||
depends on DP++; (2) a DisplayPort monitor plus a USB keyboard at the rack;
|
||||
(3) pull the card, boot on the iGPU, set IGFX explicitly (not Auto), then refit
|
||||
the card. The NIC pins make a card-out boot safe.
|
||||
`systemctl reboot --firmware-setup` works here (`OsIndicationsSupported` bit 0
|
||||
set), so nobody has to catch the Del key at POST.
|
||||
**An OS-side patch is not possible:** the AMI `Setup`/`SaSetup` variables are not
|
||||
runtime-visible here, unlike on the gx10.
|
||||
**Known-good reference: esh-pve.** Same MS-01, same BIOS `AHWSA.1.17`, same RTX
|
||||
2000E, and its `00:02.0` is present with `boot_vga=1`. The target state works on
|
||||
this hardware.
|
||||
- **GPU: RTX 2000E Ada at `01:00.0`** (`10de:28b0`), installed 2026-09-25. The
|
||||
NIC pins held: the X710 moved to bus 03 and every NIC kept its name. **NVIDIA
|
||||
580.178.04** (open modules, DKMS) has been on the host since 2026-09-25 at 1527
|
||||
(`playbooks/pve-nvidia-host.yaml`), with `nvidia-persistenced` ordered before
|
||||
`pve-guests`. It serves CT 109 nh3-ml1.
|
||||
- **iGPU restored (2026-09-25 visit).** With the card in, the BIOS's Auto primary
|
||||
display had hidden the iGPU. That left the NanoKVM (iGPU HDMI) blind and would
|
||||
have blinded AMT KVM too, since AMT captures only the iGPU. It was set on the
|
||||
visit, and since the 1523 boot `00:02.0` is back with `boot_vga=1` and i915
|
||||
loaded. The NanoKVM should have video again (not checked from here).
|
||||
`systemctl reboot --firmware-setup` works (`OsIndicationsSupported` bit 0), so
|
||||
nobody has to catch Del at POST. The AMI `Setup` variables are not
|
||||
runtime-visible, so there is no OS-side BIOS patch. esh-pve, the same MS-01 and
|
||||
BIOS `AHWSA.1.17`, is the known-good reference.
|
||||
- **Secure Boot: OFF since the 2026-09-25 visit** (`mokutil --sb-state`: disabled),
|
||||
which matches esh-pve. While it was ON (lockdown `integrity`), the DKMS NVIDIA
|
||||
module was refused and the first install rolled itself back. The playbook's
|
||||
pre-flight refuses if it is ever turned back on without an enrolled DKMS MOK.
|
||||
- ⚠ **Kernel jumped `6.8.12-11` → `6.8.12-43` at the visit reboot, pulled in by
|
||||
our own playbook.** At 1419 the headers step ran `apt-get install
|
||||
proxmox-headers-6.8`. That upgraded the `proxmox-kernel-6.8` meta and installed
|
||||
`proxmox-kernel-6.8.12-43-pve-signed`. Nobody chose the new kernel; it booted
|
||||
because it was the newest.
|
||||
Side effect: **every -43 boot oopses in Bluetooth** (`btmtk_usb_hci_wmt_sync` →
|
||||
NULL deref in `hci_power_on`, the MS-01's MediaTek BT; taint `D`). It hit on all
|
||||
three -43 boots and on none of the -11 boots. esh-pve on `-42` shows the same
|
||||
oops and has run fine, so it is benign so far: only the BT worker dies. The fix
|
||||
is to blacklist `btusb` on both hypervisors. That is not done, because it only
|
||||
takes effect at the next boot.
|
||||
- **lxc-pve 6.0.0-1 → 6.0.0-2** (2026-09-25 1533, that one package only). This is
|
||||
Proxmox fix #7006. Without it, runc 1.5 inside a nesting CT fails every
|
||||
`docker run`. `playbooks/gpu-lxc.yaml` now upgrades it first.
|
||||
- **OOB plan status.** Prime ruled on 2026-09-25 via Miranda to HOLD console-blind
|
||||
until the site visit. Target: the NanoKVM moves to pfi-gx10, and this MS-01 uses
|
||||
its own AMT. The visit did IGFX, turned SB off and cabled plus enabled AMT.
|
||||
**Still open:** AMT is not on the network (above), so the NanoKVM stays here.
|
||||
|
||||
## Refresh state
|
||||
|
||||
|
||||
Reference in New Issue
Block a user