From 5f8a1311a5dc8556169b3953d3710104907b82ad Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Sat, 3 Oct 2026 15:23:19 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=202026-10-03=20~?= =?UTF-8?q?1440=20(nh3-pve=20clean=20on=209,=20RustDesk=201.1.16,=20esh-na?= =?UTF-8?q?s-pve=20PVE=209=20assessed)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../esh-pve-cluster-pve9-upgrade-plan.md | 10 ++++++ persistent-memory.md | 33 +++++-------------- servers/ana-docker/README.md | 2 +- 3 files changed, 20 insertions(+), 25 deletions(-) diff --git a/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md b/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md index 08ed15a..9dd359b 100644 --- a/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md +++ b/docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md @@ -39,6 +39,16 @@ Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`. 3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays). +## esh-nas-pve facts re-read 2026-10-03 (Prime asked whether infra-ops can upgrade it unattended) + +- **QNAP hardware** (`QNAP Systems, Inc.`, Comet Lake HECI present, but no AMT/IPMI we can use), so there is **no remote console**. +- PVE 8.4.20, but still running kernel **6.8.12-13**, up since **2026-08-18**. Its first reboot in ~7 weeks is itself a risk: + do a **reboot test on PVE 8 first**, so a reboot problem is not mistaken for an upgrade problem. +- **No NIC pins** (`/etc/systemd/network/` is empty); the uplink is `enp5s0f0`. Pin before the upgrade. +- No DKMS on this node (the NVIDIA headers lesson applies to `pve`, not here). +- Infra-ops can do the prep remotely. The upgrade reboot should happen only with someone able to reach ESH, or with + Prime's explicit acceptance that a failed boot means an outage until someone is on site. + ## Pre-flight (no guest downtime) 1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the diff --git a/persistent-memory.md b/persistent-memory.md index 9d3d227..6198c26 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-10-03 ~1335 PT (nh3-pve PVE 8→9 upgrade by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_ +_Last updated: 2026-10-03 ~1440 PT (nh3-pve on PVE 9 and clean; RustDesk 1.1.16; esh-nas-pve PVE 9 assessed, awaiting Prime; earlier: PVE 8→9 by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -115,30 +115,11 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections carry their own dates._ +_As of 2026-10-03 ~1440 PT. The newest subsection is first; older subsections carry their own dates._ -### 2026-10-03: live now (as of ~1335 PT) +### 2026-10-03: live now (as of ~1440 PT) -- **[1405 UPDATE] nh3-pve is ON PVE 9.2.21 / kernel 7.0.14-20 (booted 1402).** Post-check: every guest is back (the stopped 104/108 are as before), ZFS healthy, DNS, post office, albok, svos/hermes, mesh and AMT all OK. ⚠ **nh3-ml1 (CT 109) is DOWN:** no `proxmox-headers-7.0` was installed, so the nvidia 580.178.04 DKMS module was never built for 7.0, `nvidia-persistenced` failed, and CT 109 has no /dev/nvidia*. **FIXED 1409 (Prime go, `playbooks/nh3-pve-pve9-postfix.yaml`):** installed `proxmox-default-headers` + headers-7.0.14-20, nvidia DKMS built for 7.0 (580.178.04 compiles fine), nh3-ml1 back (TEI :8001/:8013 healthy, embed dim 1024); intel-microcode installed (non-free-firmware added; takes effect next boot); `pve-enterprise.sources` off; pve8to9 now 0 FAIL / 1 WARN (dkms, benign). VM 102 snapshot `pre-rootgrow-20261002` DELETED (Prime go). No `pre-pve9` snapshot had been taken. **POST-CHECK DONE: nh3-pve clean.** ⚠ Lesson: on a PVE host with DKMS, check that `proxmox-default-headers` is installed BEFORE a major upgrade, or the new kernel boots with no module. -- (Pre-upgrade note, kept:) **nh3-pve PVE 8.4.1 → 9 upgrade: Prime is running it at the console now.** Steps given to him: - - latest 8.4, then `pve8to9 --full`; - - `apt remove systemd-boot` (safe: it boots GRUB via proxmox-boot-tool); - - `zfs snapshot -r rpool/ROOT@pre-pve9`; - - bookworm → trixie in `/etc/apt/sources.list`, then `apt dist-upgrade`; - - **before the reboot, `dkms status` must show nvidia 580.178.04 built for the NEW kernel** (nh3-ml1's GPU); - - reboot. - NIC names are already pinned (`/etc/systemd/network/10-pin-*.link`). **nh3-dev (VM 102, this session's host) goes down with it.** -- **Rollback for nh3-pve:** the ZFS snapshot, from a rescue boot. Never run `zpool upgrade rpool`. -- **POST-CHECK owed once nh3-pve and nh3-dev are back:** - - every guest is running: VMs 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3 (104 nh3-laser and 108 opnsense-lab were already stopped); CTs 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale, 109 nh3-ml1; - - NH3 DNS (AdGuard on nh3-docker 10.100.50.40); - - the althing post office (:8390) and albok-service (:8392); - - the mesh (headscale, nh3-scale); - - Miranda's channel (svos :8770 + hermes-gateway on nh3-dev); - - nh3-ml1's GPU (`nvidia-smi` in CT 109); - - `nh3-pve-amt` connected in MeshCentral; - - Beszel; - - `pveversion` shows 9.x. +- **nh3-pve: DONE. ON PVE 9.2.21 / kernel 7.0.14-20 (booted 1402).** Post-check: every guest is back (the stopped 104/108 are as before), ZFS healthy, DNS, post office, albok, svos/hermes, mesh and AMT all OK. ⚠ **nh3-ml1 (CT 109) is DOWN:** no `proxmox-headers-7.0` was installed, so the nvidia 580.178.04 DKMS module was never built for 7.0, `nvidia-persistenced` failed, and CT 109 has no /dev/nvidia*. **FIXED 1409 (Prime go, `playbooks/nh3-pve-pve9-postfix.yaml`):** installed `proxmox-default-headers` + headers-7.0.14-20, nvidia DKMS built for 7.0 (580.178.04 compiles fine), nh3-ml1 back (TEI :8001/:8013 healthy, embed dim 1024); intel-microcode installed (non-free-firmware added; takes effect next boot); `pve-enterprise.sources` off; pve8to9 now 0 FAIL / 1 WARN (dkms, benign). VM 102 snapshot `pre-rootgrow-20261002` DELETED (Prime go). No `pre-pve9` snapshot had been taken. **POST-CHECK DONE: nh3-pve clean.** ⚠ Lesson: on a PVE host with DKMS, check that `proxmox-default-headers` is installed BEFORE a major upgrade, or the new kernel boots with no module. - **nh3-pve-2** (MS-03, NH3) is live at 10.100.250.62 on nh3-mgmt: - X710 `nic3` on USW port 25; - AMT `nic0` with `auto` + `/etc/sysctl.d/90-amt-port.conf`; @@ -147,11 +128,15 @@ _As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections ca Open and NOT approved yet (Prime's call): disable the ceph enterprise apt source, apt upgrade, and one reboot to prove `.62` and nic0 survive. → `servers/nh3-pve-2/README.md` - **esh-pve-2** (MS-03, ESH) is unplugged by Prime on purpose. vmstore + AMT phone-home are done. The permanent address is deferred by Prime ("address can wait"; `servers/esh-pve-2/README.md`). Run `playbooks/pve-nag-patch.yaml` on it when it is back. - **cira-tunnel-watchdog** on pfi-tacticalrmm has been live since 1232 and drops :4433 tunnels silent >180 s. Watch `journalctl -t cira-tunnel-watchdog`. +- **RustDesk server (ana-docker) updated 1.1.14 → 1.1.16 at 1424** (same key; ports 21115-21119 reachable via rustdesk.phasefinal.com). Private key vaulted first as `ana-docker/rustdesk/id_ed25519`. "Change ID" is refused by the OSS server BY DESIGN (Pro feature; `rendezvous_server.rs` answers TCP RegisterPk with NOT_SUPPORT). Workaround, read at client source and untested: stop RustDesk, delete `enc_id` from RustDesk.toml, set a plain `id`, restart. → `servers/ana-docker/README.md` +- **Booth vs 10.100.10.166 (Prime, ~1420): that client cannot reach the Booth.** The Booth was up and answered from 4 hosts at 1417; .166 is on nh3-dev's own subnet. Prime interrupted my check, so it is OPEN and resumes only on his word. +- **esh-nas-pve → PVE 9: Prime asked (1440) whether I can do it myself. Assessed read-only and ANSWERED; awaiting his call.** Facts: QNAP hardware, no AMT/IPMI (no remote console); PVE 8.4.20 but running kernel 6.8.12-13; up since 2026-08-18; ZFS root `nvme` booted by GRUB; systemd-boot meta present; NO NIC pins (uplink `enp5s0f0`); no DKMS. My answer: I can do all the prep remotely (vzdump, NIC pin, remove systemd-boot, root snapshot, pve8to9, a reboot test on PVE 8 first). The upgrade reboot should happen only with someone able to reach ESH, or with Prime's explicit acceptance of a no-console failure. VMs 100/101 on `pve` must be shut down first (hard NFS mounts). → `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md` - **Done today, nothing open:** - tank verified healthy 0629, disk kept, Miranda told; - albok-service 0.1.2 + the Nemi hourly timer; - the post-deletion Worldtree gate (infra-hermes runs it as a regression check); - - the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve. + - the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve; + - nh3-pve PVE 9 + its post-fixes; nh3-dev's VM 102 snapshot deleted. - **ESH PVE 8→9 plan:** written, NOT executed (needs Prime's green light). `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. The tank blocker is cleared. ### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01) diff --git a/servers/ana-docker/README.md b/servers/ana-docker/README.md index c455431..53b9424 100644 --- a/servers/ana-docker/README.md +++ b/servers/ana-docker/README.md @@ -43,7 +43,7 @@ General-purpose Docker host for the Anaheim colo. Runs everything at `10.250.0.0 | openwebui | 3100 | Chat UI frontend | | sillytavern | 8100 | Chat UI | | mailrise | 8025 | SMTP-to-notification gateway | -| rustdesk (hbbs + hbbr) | host-net 21115-21119 | Remote desktop relay — `rustdesk.phasefinal.com`, server **1.1.16** (updated 2026-10-03 from 1.1.14; `:latest`, image 2026-07-20), public key `vVIChMRsCxCqxjg7h1rDrrxqh5zGpS6ghs+814xV6pI=` (from `id_ed25519.pub` in volume `rustdesk_rustdesk_data`). Client install steps: Booth `rustdesk-client-setup` (kept). Private key vaulted `ana-docker/rustdesk/id_ed25519` (2026-10-03) | +| rustdesk (hbbs + hbbr) | host-net 21115-21119 | Remote desktop relay — `rustdesk.phasefinal.com`, server **1.1.16** (updated 2026-10-03 from 1.1.14; `:latest`, image 2026-07-20), public key `vVIChMRsCxCqxjg7h1rDrrxqh5zGpS6ghs+814xV6pI=` (from `id_ed25519.pub` in volume `rustdesk_rustdesk_data`). Client install steps: Booth `rustdesk-client-setup` (kept). Private key vaulted `ana-docker/rustdesk/id_ed25519` (2026-10-03). ⚠ The client's "Change ID" is refused by this OSS server by design (a Pro feature). Workaround from the client source (`hbb_common` `Config::load`), untested: stop RustDesk, delete `enc_id` from `RustDesk.toml`, set `id = ''`, restart; the client re-encrypts it into `enc_id` | | dockge | 5001 | Docker stack management UI | | beszel | 8090 | Fleet metrics hub (+ local agent); agents on the other hosts report here | | dozzle (hub as `dozzle-hub`) | 8088 | Fleet log viewer; agents on the other hosts report here |