NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled. - nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml. - nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4) deployed with HOST_NAME/HOST_IP labels. - Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min 0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control 0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed identical within rep spread. - gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006). With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl reopen; applied on nh3-pve (one package). - pve-nvidia-host.yaml: document that the headers meta drags in the newest kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot). - Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal. - nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels, AMT cabled but unreachable on the network. Gateway routing to nh3-ml1 is not changed.
4.6 KiB
4.6 KiB
[2026-09-25] nh3-ml1 LIVE after the NH3 site visit
Prime, ~1524 PT: "secure boot disabled, amt enabled, card plugged." This is the
resume of 2026-09-25-nh3-ml1-standup.md. nh3-pve came up at 1523.
Post-boot checklist: all green
- Uplink is
enp2s0f1np1. Guests 100, 101, 102, 105 and CTs 103, 106, 107 are up; 104 and 108 are off as expected. - NFS automounts are good on nh3-docker, nh3-dev and pbs-nh3. The post office is push/reachable. The 14 queued althing messages were all Beszel/Kuma "UP" recoveries, so nothing needed acting on.
- Secure Boot is disabled.
00:02.0is back withboot_vga=1and i915 loaded.
Build (elway, both playbooks idempotent)
pve-nvidia-host.yaml: the driver installed (DKMS, 36 s), 6/6 verify.gpu-lxc.yaml: CT 109 built, then verify 3 FAILED.docker rundied withopen sysctl net.ipv4.ip_unprivileged_port_start file: reopen fd 8: permission denied.- The cause is the lxc-pve version, not runc. Both CTs run runc 1.5.1. nh3-pve (PVE 8.4.1) had lxc-pve 6.0.0-1, and esh-pve (8.4.20) has 6.0.0-2, whose changelog is Proxmox "fix #7006: do not restrict /proc or /sys if nesting is enabled".
- Decided autonomously: upgrade that one package (a dry-run showed
1 upgraded, 0 new, 0 removed) and
pct reboot 109. GPU-in-docker then passed. ops-log recorded. - The playbook gained a first step that upgrades lxc-pve when it is below
6.0.0-2, with a quoted name, because an unquoted
#7006truncated it as a YAML comment. The re-run was 5/5 verify OK. The guard skips on esh-pve (tested).
embed-rerankdeploy.deploy-stack.sh --yeswas denied by the auto-mode classifier as a blind apply, so the pattern now is: preview withecho n | deploy-stack.sh …(it shows the diff and aborts), thenecho y |..env= example +HOST_NAME=nh3-ml1,HOST_IP=10.100.50.80. Healthy in about 120 s including the image pull; 2.7 GB VRAM.
Parity and speed (vs esh-ml1)
Script: session scratchpad, not committed. Corpus: 1,120 repo-doc paragraphs plus 6 fixed texts, 2 runs per host, 50 queries.
- Embedding cosine esh↔nh3: median 0.999998, min 0.999993. The floors are esh 0.999995 and nh3 0.999993.
- Negative control (different texts): median 0.50.
- overlap@10 = 1.000 both ways. Positive control, MRL-256: 0.684.
- Rerank, 100×20: top-1 1.00, max score diff 0.0014 (the esh-esh floor is 0.0020). Positive control, a 4-word query: top-5 order 0.09, max diff 0.97.
- Speed on-box, 3 interleaved reps: short embed p50 6.6–6.7 vs 7.0 ms; bulk 104–105 vs 106–108 passages/s; rerank-20 p50 156–159 vs 152–161 ms. The same.
Monitoring and wiring
- Beszel: agent (
hosts/nh3-ml1.yaml, same hub key as esh-ml1), system1feeeq61g4mkqre, 5 alerts, GPU samples present.- ⚠ zsh gotcha:
for a in "X 1 2"; do set -- $adoes NOT word-split in zsh, so the first batch of alert POSTs failed silently (empty fields). Use a function with explicit args.
- ⚠ zsh gotcha:
- Kuma #29/#30 (the seed also refreshed the esh-ml1 descriptions). Homepage
nh3-ml1-dockershows 2 cards with no restart. Dozzle agent added, and the hub.envbacked up as.env.bak-20260925-nh3ml1before the edit. - DNS
nh3-ml1.nh3.internalon all 3 resolvers.~/.ssh/configHost nh3-ml1;servers/nh3-ml1/has ssh-target, system-details and README.
Found along the way
- Kernel 6.8.12-43 came from our own playbook. The headers step's
proxmox-headers-6.8install upgraded theproxmox-kernel-6.8meta (apt history 14:19:23). The visit reboot booted it. Documented inpve-nvidia-host.yaml. - btmtk Bluetooth oops on every -4x boot (2/2 lines on each of 3 -43 boots,
0 on the -11 boots; esh-pve -42 is the same). It taints
D. Only the BT worker dies. Proposed fix: blacklist btusb on both PVE hosts. Not done. - AMT is not reachable.
- The LM port has a 1 Gb carrier (brief unbridged
ip link upwith IPv6 disabled, then restored; ops-log recorded). - No UDM client or lease for
58:47:ca:76:96:0e. - A /dev/tcp sweep of 16992/16993 over 10.100.{0,10,50,250}.0/24 found nothing. Its positive controls, :8390 and :8006, did answer.
- The USW Pro 24's up ports (19, 22, 23, 26) are all other devices. So the cable
is probably on nh3-sw1 (
10.100.250.2), where SSH times out and no creds are held. - Asked Prime.
- The LM port has a 1 Gb carrier (brief unbridged
Open (Prime's)
- Gateway routing for
qwen3-embedding/reranker. Recommendation: load-share (identical cards, and a second site). - AMT network settings and switch port. The NanoKVM stays on nh3-pve until AMT KVM shows the console.
- btusb blacklist on nh3-pve and esh-pve. Low priority.