feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1

NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled.

- nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml.
- nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4)
  deployed with HOST_NAME/HOST_IP labels.
- Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min
  0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control
  0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed
  identical within rep spread.
- gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006).
  With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl
  reopen; applied on nh3-pve (one package).
- pve-nvidia-host.yaml: document that the headers meta drags in the newest
  kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot).
- Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage
  nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal.
- nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels,
  AMT cabled but unreachable on the network.

Gateway routing to nh3-ml1 is not changed.
This commit is contained in:
vh
2026-09-25 15:55:33 -07:00
parent 6fa8213c20
commit 5960526c3f
17 changed files with 815 additions and 112 deletions
+17 -8
View File
@@ -93,19 +93,28 @@ monitors:
url: http://10.250.50.70:8200/api/v1/services
# ---- fleet embed/rerank: the one EXCEPTION to "seats are OUT" ----
# Since 2026-09-25 these are the SOLE backends behind the gateway's
# `qwen3-embedding` and `reranker` (TEI on esh-ml1, no failover until the second
# RTX 2000 arrives). They are not come-and-go seats: dead = Worldtree recall,
# nevermore clustering and Open WebUI RAG all fail. TEI's /health runs the
# backend, so a loaded-but-broken model reads DOWN, not UP. The reward seat on the
# same box stays OUT: it has no working consumer (stacks/reward-seat/README.md).
# Since 2026-09-25 these back the gateway's `qwen3-embedding` and `reranker`
# (TEI on esh-ml1; nh3-ml1 is the second RTX 2000, parity-verified the same day,
# gateway routing pending Prime). They are not come-and-go seats: dead =
# Worldtree recall, nevermore clustering and Open WebUI RAG all fail. TEI's
# /health runs the backend, so a loaded-but-broken model reads DOWN, not UP. The
# reward seat on esh-ml1 stays OUT: it has no working consumer
# (stacks/reward-seat/README.md).
- name: Embed — Qwen3 0.6B (TEI, esh-ml1)
url: http://10.0.50.80:8001/health
description: sole backend for gateway `qwen3-embedding`
description: gateway `qwen3-embedding` backend (esh-ml1)
- name: Rerank — bge-v2-m3 (TEI, esh-ml1)
url: http://10.0.50.80:8013/health
description: sole backend for gateway `reranker`
description: gateway `reranker` backend (esh-ml1)
- name: Embed — Qwen3 0.6B (TEI, nh3-ml1)
url: http://10.100.50.80:8001/health
description: second `qwen3-embedding` backend (nh3-ml1)
- name: Rerank — bge-v2-m3 (TEI, nh3-ml1)
url: http://10.100.50.80:8013/health
description: second `reranker` backend (nh3-ml1)
- name: talk
url: https://talk.nh3.phasefinal.com:8092/