memory: snapshot — nh3-pve shutting down for the NH3 site visit (post-boot checklist), nh3-ml1 blocked on Secure Boot, gx10 AC-restore validated
This commit is contained in:
+50
-68
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-25 ~1336 PT (⭐⭐ nh3-pve back from the GPU install, checklist passed; ⚠ iGPU hidden by BIOS → NanoKVM blind, HELD to next NH3 visit (Prime); gx10 AC-restore VALIDATED. Today: TEI = fleet embed/rerank engine, esh-ml1 sole backend + reward seat moved there, fv-ml1 GPU 1 freed ~15 GB; esh-ml1 monitoring wired; homepage wedge cleared; NIC names pinned on nh3-pve. 34 old entries archived.)_
|
||||
_Last updated: 2026-09-25 ~1430 PT (⭐⭐ nh3-pve SHUT DOWN by Prime for the NH3 site visit — post-boot checklist in Current state. nh3-ml1 blocked on nh3-pve SECURE BOOT; Prime's SB call pending. NanoKVM blind until IGFX fix on the visit; gx10 AC-restore VALIDATED; GPU playbooks host-generic.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -115,74 +115,52 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-09-24 ~2245 PT._
|
||||
_As of 2026-09-25 ~1430 PT._
|
||||
|
||||
### ⭐⭐ nh3-pve GPU install — host back 2026-09-25 1116, checklist run 1122
|
||||
### ⭐⭐ nh3-pve SHUT DOWN 2026-09-25 ~1430 PT (Prime's order) for the NH3 site visit
|
||||
|
||||
**Result:** checklist items 1 and 3–6 pass, and nothing needed restarting. Guests,
|
||||
NFS, the post office (push), svos, hermes-gateway, Booth, talk, ttyd, zellij,
|
||||
backup timers, pbs-nh3 datastore, the mesh (irv-ml1 direct), the egress proxy and
|
||||
NH3 DNS are all fine. Open items:
|
||||
- ⚠ **iGPU gone → NanoKVM has no picture. Prime's call (1333, via Miranda): HOLD,
|
||||
and stay blind until the next NH3 visit.** No BIOS change and no reboot. Target:
|
||||
the NanoKVM moves to pfi-gx10, and nh3-pve gets its own vPro/AMT. ⚠ AMT KVM
|
||||
captures only the iGPU, so the IGFX BIOS fix has to happen on that same visit,
|
||||
before the NanoKVM leaves. The order is in `servers/nh3-pve/README.md`.
|
||||
- The uplink cable moved from `enp2s0f0np0` to `enp2s0f1np1`. It is harmless
|
||||
because both ports are bridge members. Recorded in the README.
|
||||
- ✅ **pfi-gx10 AC-restore patch VALIDATED.** Prime confirmed he replugged AC and
|
||||
it came up by itself at 1108:54; the README is updated.
|
||||
- Pre-existing and not caused by the reboot: `althing-pump-infra-hermes` and
|
||||
`althing-seat-daemon-jekyll` are both disabled (stopped 09-19 and 09-24). The
|
||||
`infra-hermes` handle shows `pull`.
|
||||
This session runs ON nh3-dev and dies with it. nh3-dev wipes /tmp at boot, so there
|
||||
is NO /tmp handoff after this. **Successor: work this list first.**
|
||||
|
||||
**nh3-ml1 (Prime, 2026-09-25 ~1410): IN PROGRESS.** CT 109 on nh3-pve at
|
||||
10.100.50.80 (VLAN 50, outside the UDM pool .150–.249), local-zfs, a second
|
||||
embed/rerank (TEI) backend on the new RTX. Playbooks are host-generic as of
|
||||
`bc278d4`. Prime cleared the driver install at ~1418. **It then failed: nh3-pve has
|
||||
SECURE BOOT ON** (esh-pve has it off), and the unsigned DKMS module was refused.
|
||||
The installer rolled back. The playbook now refuses on this up front. Unblocking
|
||||
needs the console at boot (MokManager enroll, or SB off in the BIOS), which ties
|
||||
it to the IGFX/NanoKVM visit, unless Prime picks a VFIO VM. Waiting on his call.
|
||||
After it: `gpu-lxc.yaml` with nh3 vars → embed-rerank with
|
||||
`HOST_NAME/HOST_IP` → parity vs esh-ml1 → monitoring and DNS → gateway routing
|
||||
(Prime's call).
|
||||
|
||||
The pre-shutdown record follows, kept for the next GPU visit:
|
||||
|
||||
Pre-flight done before shutdown:
|
||||
- NIC names pinned by MAC (a GPU renames every NIC on this MS-01), initramfs
|
||||
rebuilt and ESP synced.
|
||||
- AMT port `enp88s0` removed from vmbr0 in the file.
|
||||
- nh3-dev's Beszel agent no longer binds NAS shares.
|
||||
- Details in `servers/nh3-pve/README.md`.
|
||||
**The site-visit plan.** Prime has mini-DP→DP adapters, so a DisplayPort monitor on
|
||||
the RTX plus a USB keyboard covers all of it:
|
||||
1. BIOS: set Primary Display = IGFX. The Auto setting hid the iGPU, so the
|
||||
NanoKVM is blind.
|
||||
2. BIOS/Secure Boot: Prime's call is PENDING. (1) SB off (recommended; matches
|
||||
esh-pve), (2) keep SB and enroll a DKMS MOK via MokManager, or (3) a VFIO VM
|
||||
instead of the LXC.
|
||||
3. vPro: cable I226-LM `enp88s0`, provision MEBx, confirm AMT KVM shows the console.
|
||||
4. Only then move the NanoKVM to pfi-gx10.
|
||||
Detail and order: `servers/nh3-pve/README.md`.
|
||||
|
||||
**Post-boot checklist:**
|
||||
1. **Host on the network:** `ping 10.100.250.60`, then `ip -br link` — names
|
||||
unchanged, `enp2s0f0np0` UP inside vmbr0. If there's no network, open the
|
||||
NanoKVM console at **https://10.100.250.171**.
|
||||
2. **GPU seen:** `lspci | grep -i nvidia` → expect `01:00.0`, and the X710 moved to
|
||||
`03:00` while keeping its pinned name. The BIOS primary display must stay on the
|
||||
iGPU, or the NanoKVM goes black.
|
||||
3. **Guests:**
|
||||
- running: 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3, CT 103
|
||||
nh3-wg, CT 106 nh3-headscale, CT 107 nh3-scale;
|
||||
- staying off (expected): 104 nh3-laser, 108 opnsense-lab.
|
||||
4. **NFS:** touch the automounts — nh3-docker `/mnt/compose` `/mnt/backup`, nh3-dev
|
||||
`/mnt/backup` `/mnt/smithy`.
|
||||
5. **Post office** reachable (`postbox status`, `mode: push`). Kuma "althing post
|
||||
office" green.
|
||||
6. **nh3-dev services:** svos :8770 + hermes-gateway (Miranda's channel), Booth
|
||||
:8090, ttyd/zellij seats, `dev-backup.timer`. Beszel shows nh3-dev /
|
||||
nh3-docker / pbs-nh3 up.
|
||||
7. **Claude sessions:** relaunching them is Prime's call.
|
||||
8. **pfi-gx10:** if Prime ran the AC-pull test on this visit, verify it came back
|
||||
by itself.
|
||||
9. **Next work:** the GPU's purpose is TBD with Prime; likely the second embed/rerank
|
||||
backend. Reuse `playbooks/pve-nvidia-host.yaml` / `gpu-lxc.yaml` (make
|
||||
them host-generic). nh3-pve runs kernel 6.8.12-11 with **no matching
|
||||
proxmox-headers installed**, on PVE 8.4.1 (Debian 12 template only). If the LM
|
||||
port was cabled, MEBx provisioning can be done through the NanoKVM.
|
||||
1. Network: `ping 10.100.250.60`; `ip -br link`. ⚠ The uplink is now
|
||||
`enp2s0f1np1` (the cable moved cages on 09-25 at 1116), not `enp2s0f0np0`. Both
|
||||
are in vmbr0.
|
||||
2. GPU: `lspci` shows `01:00.0` RTX 2000E. If IGFX was set, `00:02.0` is back with
|
||||
`boot_vga=1`. `mokutil --sb-state` shows whether Secure Boot is off.
|
||||
3. Guests running: 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3,
|
||||
CT 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale. Staying off: 104, 108.
|
||||
4. NFS automounts: nh3-docker `/mnt/compose` `/mnt/backup`; nh3-dev `/mnt/backup`
|
||||
`/mnt/smithy`; pbs-nh3 datastore.
|
||||
5. Post office `postbox status` shows push/reachable; Kuma is green.
|
||||
6. nh3-dev services: svos :8770, hermes-gateway (Miranda's channel), Booth :8090,
|
||||
talk, ttyd RO/RW, zellij-web, `dev-backup.timer`.
|
||||
7. Claude sessions: relaunching them is Prime's call.
|
||||
8. pfi-gx10: the AC-restore patch is validated, so a power cut brings it back on
|
||||
its own.
|
||||
9. Resume nh3-ml1, only once Secure Boot is resolved.
|
||||
`scripts/elway root@nh3-pve.nh3.internal --playbook playbooks/pve-nvidia-host.yaml`.
|
||||
Its pre-flight refuses while SB is on without an enrolled key. Then run
|
||||
`playbooks/gpu-lxc.yaml` with the nh3-ml1 `--var` line from its header, deploy
|
||||
`embed-rerank` with `.env` `HOST_NAME=nh3-ml1` `HOST_IP=10.100.50.80`, run a
|
||||
parity check against esh-ml1 (noise floor plus controls), then wire monitoring,
|
||||
DNS and `servers/nh3-ml1/README.md`. Gateway routing is Prime's call; the
|
||||
recommendation is load-share after parity.
|
||||
|
||||
**Left on nh3-pve by the failed 09-25 driver install** (harmless, and needed later):
|
||||
headers 6.8.12-11 plus the series meta, dkms, build-essential; nouveau
|
||||
blacklisted and unloaded; the `.run` staged in `/root/nvidia`.
|
||||
|
||||
### ⭐ esh-ml1 = the fleet's utility GPU host (2026-09-25): embed/rerank (TEI) + reward (vLLM)
|
||||
|
||||
@@ -220,15 +198,17 @@ in `servers/pfi-gx10/README.md`. → `persistent-memory.d/2026-09-24-gx10-ac-res
|
||||
|
||||
### Live threads
|
||||
|
||||
- **Claude sessions on nh3-dev** died in the NH3 outage (only infra-ops,
|
||||
infra-hermes, jekyll survived). Relaunch is Prime's call; not confirmed done.
|
||||
- **Claude sessions on nh3-dev:** Prime relaunched 3 after the 1116 boot. They die again
|
||||
with the ~1430 shutdown, and relaunching them is his call.
|
||||
- **Homepage `hermes-gateway seat` rename resolved (2026-09-25):** the operator's 2026-09-24 handle split is real (`postbox handles` records it; infra-hermes confirmed the seat daemon carries `hermes-gateway`). The rename was already live and is committed as `422a27c`.
|
||||
- **High Seat backup proof DONE 2026-09-25 0111:** infra-hermes reported and I verified that restic snapshot `cd113bef` (nh3-dev, 0100) contains `~/.high-seat/reports/grok.963ddd3f428b7fc7.jsonl`. Told him to retire the daily 0110 check unless svos-dev asked for an ongoing one.
|
||||
- ESH has a single outside route (esh-scale on esh-pve) — noted, untracked.
|
||||
- origin/main is current through `2118449`; only the snapshot commits after it are unpushed (2026-09-25). Push is Prime's call.
|
||||
- origin/main is current through `58390a4`. Everything after it is unpushed (`2bf418c`…`7ddd116` plus this snapshot, 2026-09-25). Pushing is Prime's call.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-25]` **nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls.** Tracked: `7ddd116`, Current state. → `persistent-memory.d/2026-09-25-nh3-ml1-standup.md`
|
||||
- `[2026-09-25]` **nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333).** Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked: `servers/nh3-pve/README.md` (`2b49be4`, `cdd7605`).
|
||||
- `[2026-09-25]` **MS-01 foot-gun: a GPU in the PCIe slot renames every NIC** (the slot's root port takes bus 01, so the X710 goes `enp2s0f0np0`→`enp3s0f0np0`). On nh3-pve that means booting with no uplink. Pinned by MAC with `.link` files before the install, and the AMT port was pulled from vmbr0. → `servers/nh3-pve/README.md`, `playbooks/nh3-pve-pin-nic-names.yaml` (`2118449`)
|
||||
- `[2026-09-25]` **Direct ESH→esh-ml1 consumer path — PARKED** (no ESH-side callers in 7 days; consumers go through the gateway). Tracked at `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`.
|
||||
- `[2026-09-25]` Created empty private repo `corviduo/norn` (git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; the `norn-dev` handle is the operator's to declare.
|
||||
@@ -236,7 +216,7 @@ in `servers/pfi-gx10/README.md`. → `persistent-memory.d/2026-09-24-gx10-ac-res
|
||||
- `[2026-09-25]` **TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend for `qwen3-embedding` + `reranker`; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1.** → `persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md`
|
||||
- `[2026-09-24]` **esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM `order: 2` failover; dead DB alias `reranker-a3-bge-v2-m3` (pre-relocation IP) repaired.** → `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`
|
||||
- `[2026-09-24]` **esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank** (implementation deferred to next session, tracked here + `servers/esh-pve/README.md`). → `persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md`
|
||||
- `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
|
||||
- `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). — `[2026-09-25]` ✅ **VALIDATED**: Prime pulled and replugged AC, and it came up by itself at 1108:54 (`2b49be4`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
|
||||
- `[2026-09-24]` **NH3 power outage recovered** — pbs-nh3 had no `onboot` (set), NFS boot race fixed with automount (`1cbde50`), every other Claude session on nh3-dev died. → `persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md`
|
||||
- `[2026-09-24]` **Miranda standing order is a repo CLAUDE.md operating parameter** (`4b29492`, aligned to the global send protocol in `bcf3342`): high-urgency matters go to her, fixed or not; URGENT only when it cannot wait (she phones Prime). Channel verified end to end (thread `01M3A0RP4Q8T0KNGH8TMFSNDA6`); it depends on svos + hermes-gateway. Prime's callsign **PRiMe / papa romeo mike** is a name, not an authenticator (`617b759`, `62817a2`).
|
||||
- `[2026-09-24]` **task-board mothballed** (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (`e6da607`). Its hooks had sent no traffic in 30 days.
|
||||
@@ -429,6 +409,8 @@ _104 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-09-25]` **The esh-pve NVIDIA DKMS recipe on nh3-pve** — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flights `mokutil --sb-state` (`7ddd116`).
|
||||
- `[2026-09-25]` **elway against an unpinned host-key name** — every `when:` hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 in `when:` should fail.
|
||||
- `[2026-09-24]` **Testing "Restore AC Power Loss" with an OS shutdown** — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button.
|
||||
- `[2026-09-23]` **`booth link --help`** — there is no help flag; it posts `--help` to the operator's link board as a link. Read `booth` with no args for usage.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user