feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1

NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled.

- nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml.
- nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4)
  deployed with HOST_NAME/HOST_IP labels.
- Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min
  0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control
  0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed
  identical within rep spread.
- gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006).
  With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl
  reopen; applied on nh3-pve (one package).
- pve-nvidia-host.yaml: document that the headers meta drags in the newest
  kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot).
- Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage
  nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal.
- nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels,
  AMT cabled but unreachable on the network.

Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
vh
2026-09-25 15:55:33 -07:00
parent 6fa8213c20
commit 5960526c3f
17 changed files with 815 additions and 112 deletions
+1
View File
@@ -24,6 +24,7 @@ filesystem samples verified; fleet 13/14 up with known fv-ml1 outage.
| vm-esh-nas | beszel-agent-esh-nas | /mnt/books, /mnt/share, /mnt/music, /mnt/media |
| nh3-dev | beszel | none since 2026-09-25 (was /mnt/backup, /mnt/smithy — see below) |
| esh-ml1 (added 2026-09-25) | beszel | none — NVIDIA image (`hosts/esh-ml1.yaml`) for the RTX 2000E Ada |
| nh3-ml1 (added 2026-09-25) | beszel | none — NVIDIA image (`hosts/nh3-ml1.yaml`) for the RTX 2000E Ada; hub system `1feeeq61g4mkqre`, same 5 alerts as esh-ml1 |
⚠ **nh3-dev's agent was DOWN from the 2026-09-24 NH3 power recovery until
2026-09-25.** Docker could not bind `/mnt/smithy` at boot ("no such device"): since
+15
View File
@@ -0,0 +1,15 @@
# nh3-ml1 (CT 109 on nh3-pve) — Beszel agent with NVIDIA GPU telemetry for the
# RTX 2000E Ada (utilization, VRAM, temperature, power). Twin of hosts/esh-ml1.yaml:
# Docker-in-LXC, NVIDIA container toolkit with no-cgroups=true
# (playbooks/gpu-lxc.yaml). No extra filesystems: the root filesystem holds
# everything, models included.
services:
beszel-agent:
image: henrygd/beszel-agent-nvidia:0.18.7
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [utility]
+1
View File
@@ -8,6 +8,7 @@ Container log viewer. One UI on **ana-docker** aggregates logs from every Docker
- **esh-docker-vm** / **vm-esh-nas** (agents) — `10.0.50.45:7007`, `10.0.50.154:7007`
- **irv-ml1** (agent) — `10.6.110.50:7007` (mesh address)
- **esh-ml1** (agent, added 2026-09-25) — `10.0.50.80:7007`, pinned `v10.4.1` = the hub's version; compose dir `dozzle-agent`. The host needed an (empty) `traefik-net` network because the compose declares it external.
- **nh3-ml1** (agent, added 2026-09-25) — `10.100.50.80:7007`, same shape as esh-ml1 (`v10.4.1`, `dozzle-agent`, empty `traefik-net`). Hub `clients` went 7 → 8.
- **nh3-docker** (agent, cross-site) — `10.100.50.40:7007`. ⚠ **Stopped by hand ~2026-04 (Exited 0) and left that way**; the hub logs a refused connection for it. Revive it or drop it from the list deliberately.
⚠ **The hub's agent list silently rots when a host moves.** Until 2026-09-25 it
+7 -1
View File
@@ -3,7 +3,10 @@
**The fleet's embedding + reranking service**, on **esh-ml1** (CT 110 on esh-pve,
RTX 2000E Ada), served by **Hugging Face Text Embeddings Inference (TEI)**.
Since 2026-09-25 it is the only backend behind the gateway names
`qwen3-embedding`, `reranker` and `reranker-a3-bge-v2-m3`.
`qwen3-embedding`, `reranker` and `reranker-a3-bge-v2-m3`. A second instance
runs on **nh3-ml1** (CT 109 on nh3-pve, the same card), parity-verified against
esh-ml1 on 2026-09-25 (`servers/nh3-ml1/README.md`). It is not in the gateway
yet; that is Prime's call.
**TEI is the fleet's embed/rerank engine** (Prime, 2026-09-25). New embedding or
reranking seats go on TEI, not vLLM. Why, and the measurements behind it:
@@ -43,6 +46,9 @@ scripts/deploy-stack.sh esh-ml1 embed-rerank
ssh esh-ml1 'cd /opt/docker/compose/embed-rerank && cp -n .env.example .env && docker compose config -q && docker compose up -d'
```
nh3-ml1 is the same, plus `HOST_NAME=nh3-ml1` and `HOST_IP=10.100.50.80` in its
`.env` (they only feed the Homepage labels).
Host prerequisites (driver, LXC, docker, toolkit) are in
[`servers/esh-ml1/README.md`](../../servers/esh-ml1/README.md).
+6
View File
@@ -37,6 +37,12 @@ esh-ml1-docker:
host: 10.0.50.80
port: 2375
# nh3-ml1 — CT 109 on nh3-pve: the second embed/rerank (TEI) backend.
# dockerd listens only on its own address (playbooks/gpu-lxc.yaml), 2026-09-25.
nh3-ml1-docker:
host: 10.100.50.80
port: 2375
# Example TLS socket (if/when a host moves off plaintext 2375):
# ana-pfi-docker:
# host: 10.250.50.70
+17 -8
View File
@@ -93,19 +93,28 @@ monitors:
url: http://10.250.50.70:8200/api/v1/services
# ---- fleet embed/rerank: the one EXCEPTION to "seats are OUT" ----
# Since 2026-09-25 these are the SOLE backends behind the gateway's
# `qwen3-embedding` and `reranker` (TEI on esh-ml1, no failover until the second
# RTX 2000 arrives). They are not come-and-go seats: dead = Worldtree recall,
# nevermore clustering and Open WebUI RAG all fail. TEI's /health runs the
# backend, so a loaded-but-broken model reads DOWN, not UP. The reward seat on the
# same box stays OUT: it has no working consumer (stacks/reward-seat/README.md).
# Since 2026-09-25 these back the gateway's `qwen3-embedding` and `reranker`
# (TEI on esh-ml1; nh3-ml1 is the second RTX 2000, parity-verified the same day,
# gateway routing pending Prime). They are not come-and-go seats: dead =
# Worldtree recall, nevermore clustering and Open WebUI RAG all fail. TEI's
# /health runs the backend, so a loaded-but-broken model reads DOWN, not UP. The
# reward seat on esh-ml1 stays OUT: it has no working consumer
# (stacks/reward-seat/README.md).
- name: Embed — Qwen3 0.6B (TEI, esh-ml1)
url: http://10.0.50.80:8001/health
description: sole backend for gateway `qwen3-embedding`
description: gateway `qwen3-embedding` backend (esh-ml1)
- name: Rerank — bge-v2-m3 (TEI, esh-ml1)
url: http://10.0.50.80:8013/health
description: sole backend for gateway `reranker`
description: gateway `reranker` backend (esh-ml1)
- name: Embed — Qwen3 0.6B (TEI, nh3-ml1)
url: http://10.100.50.80:8001/health
description: second `qwen3-embedding` backend (nh3-ml1)
- name: Rerank — bge-v2-m3 (TEI, nh3-ml1)
url: http://10.100.50.80:8013/health
description: second `reranker` backend (nh3-ml1)
- name: talk
url: https://talk.nh3.phasefinal.com:8092/