memory: snapshot — 2026-10-03 ~1440 (nh3-pve clean on 9, RustDesk 1.1.16, esh-nas-pve PVE 9 assessed)

This commit is contained in:
vh
2026-10-03 15:23:19 -07:00
parent d4675f2cb5
commit 5f8a1311a5
3 changed files with 20 additions and 25 deletions
@@ -39,6 +39,16 @@ Quorum: 2 nodes × 1 vote, no QDevice, no `two_node`.
3. **`pve8to9` FAIL on both nodes:** the `systemd-boot` meta-package is installed. Both nodes boot
via GRUB, so remove it (`apt remove systemd-boot`; `systemd-boot-efi` stays).
## esh-nas-pve facts re-read 2026-10-03 (Prime asked whether infra-ops can upgrade it unattended)
- **QNAP hardware** (`QNAP Systems, Inc.`, Comet Lake HECI present, but no AMT/IPMI we can use), so there is **no remote console**.
- PVE 8.4.20, but still running kernel **6.8.12-13**, up since **2026-08-18**. Its first reboot in ~7 weeks is itself a risk:
do a **reboot test on PVE 8 first**, so a reboot problem is not mistaken for an upgrade problem.
- **No NIC pins** (`/etc/systemd/network/` is empty); the uplink is `enp5s0f0`. Pin before the upgrade.
- No DKMS on this node (the NVIDIA headers lesson applies to `pve`, not here).
- Infra-ops can do the prep remotely. The upgrade reboot should happen only with someone able to reach ESH, or with
Prime's explicit acceptance that a failed boot means an outage until someone is on site.
## Pre-flight (no guest downtime)
1. **Backups:** one-off `vzdump` of **all ten** guests to pbs-ana in snapshot mode, right before the
+9 -24
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-10-03 ~1335 PT (nh3-pve PVE 8→9 upgrade by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_
_Last updated: 2026-10-03 ~1440 PT (nh3-pve on PVE 9 and clean; RustDesk 1.1.16; esh-nas-pve PVE 9 assessed, awaiting Prime; earlier: PVE 8→9 by Prime at the console; nh3-pve-2 live on .62 with vmstore; CIRA watchdog live; PVE nag patched; tank healthy; albok 0.1.2 + Nemi timer. Prior: 2026-10-01 Parakeet/U11a.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,30 +115,11 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections carry their own dates._
_As of 2026-10-03 ~1440 PT. The newest subsection is first; older subsections carry their own dates._
### 2026-10-03: live now (as of ~1335 PT)
### 2026-10-03: live now (as of ~1440 PT)
- **[1405 UPDATE] nh3-pve is ON PVE 9.2.21 / kernel 7.0.14-20 (booted 1402).** Post-check: every guest is back (the stopped 104/108 are as before), ZFS healthy, DNS, post office, albok, svos/hermes, mesh and AMT all OK. ⚠ **nh3-ml1 (CT 109) is DOWN:** no `proxmox-headers-7.0` was installed, so the nvidia 580.178.04 DKMS module was never built for 7.0, `nvidia-persistenced` failed, and CT 109 has no /dev/nvidia*. **FIXED 1409 (Prime go, `playbooks/nh3-pve-pve9-postfix.yaml`):** installed `proxmox-default-headers` + headers-7.0.14-20, nvidia DKMS built for 7.0 (580.178.04 compiles fine), nh3-ml1 back (TEI :8001/:8013 healthy, embed dim 1024); intel-microcode installed (non-free-firmware added; takes effect next boot); `pve-enterprise.sources` off; pve8to9 now 0 FAIL / 1 WARN (dkms, benign). VM 102 snapshot `pre-rootgrow-20261002` DELETED (Prime go). No `pre-pve9` snapshot had been taken. **POST-CHECK DONE: nh3-pve clean.** ⚠ Lesson: on a PVE host with DKMS, check that `proxmox-default-headers` is installed BEFORE a major upgrade, or the new kernel boots with no module.
- (Pre-upgrade note, kept:) **nh3-pve PVE 8.4.1 → 9 upgrade: Prime is running it at the console now.** Steps given to him:
- latest 8.4, then `pve8to9 --full`;
- `apt remove systemd-boot` (safe: it boots GRUB via proxmox-boot-tool);
- `zfs snapshot -r rpool/ROOT@pre-pve9`;
- bookworm → trixie in `/etc/apt/sources.list`, then `apt dist-upgrade`;
- **before the reboot, `dkms status` must show nvidia 580.178.04 built for the NEW kernel** (nh3-ml1's GPU);
- reboot.
NIC names are already pinned (`/etc/systemd/network/10-pin-*.link`). **nh3-dev (VM 102, this session's host) goes down with it.**
- **Rollback for nh3-pve:** the ZFS snapshot, from a rescue boot. Never run `zpool upgrade rpool`.
- **POST-CHECK owed once nh3-pve and nh3-dev are back:**
- every guest is running: VMs 100 nh3-docker1, 101 nh3-extdev, 102 nh3-dev, 105 pbs-nh3 (104 nh3-laser and 108 opnsense-lab were already stopped); CTs 103 nh3-wg, 106 nh3-headscale, 107 nh3-scale, 109 nh3-ml1;
- NH3 DNS (AdGuard on nh3-docker 10.100.50.40);
- the althing post office (:8390) and albok-service (:8392);
- the mesh (headscale, nh3-scale);
- Miranda's channel (svos :8770 + hermes-gateway on nh3-dev);
- nh3-ml1's GPU (`nvidia-smi` in CT 109);
- `nh3-pve-amt` connected in MeshCentral;
- Beszel;
- `pveversion` shows 9.x.
- **nh3-pve: DONE. ON PVE 9.2.21 / kernel 7.0.14-20 (booted 1402).** Post-check: every guest is back (the stopped 104/108 are as before), ZFS healthy, DNS, post office, albok, svos/hermes, mesh and AMT all OK. ⚠ **nh3-ml1 (CT 109) is DOWN:** no `proxmox-headers-7.0` was installed, so the nvidia 580.178.04 DKMS module was never built for 7.0, `nvidia-persistenced` failed, and CT 109 has no /dev/nvidia*. **FIXED 1409 (Prime go, `playbooks/nh3-pve-pve9-postfix.yaml`):** installed `proxmox-default-headers` + headers-7.0.14-20, nvidia DKMS built for 7.0 (580.178.04 compiles fine), nh3-ml1 back (TEI :8001/:8013 healthy, embed dim 1024); intel-microcode installed (non-free-firmware added; takes effect next boot); `pve-enterprise.sources` off; pve8to9 now 0 FAIL / 1 WARN (dkms, benign). VM 102 snapshot `pre-rootgrow-20261002` DELETED (Prime go). No `pre-pve9` snapshot had been taken. **POST-CHECK DONE: nh3-pve clean.** ⚠ Lesson: on a PVE host with DKMS, check that `proxmox-default-headers` is installed BEFORE a major upgrade, or the new kernel boots with no module.
- **nh3-pve-2** (MS-03, NH3) is live at 10.100.250.62 on nh3-mgmt:
- X710 `nic3` on USW port 25;
- AMT `nic0` with `auto` + `/etc/sysctl.d/90-amt-port.conf`;
@@ -147,11 +128,15 @@ _As of 2026-10-03 ~1335 PT. The newest subsection is first; older subsections ca
Open and NOT approved yet (Prime's call): disable the ceph enterprise apt source, apt upgrade, and one reboot to prove `.62` and nic0 survive. → `servers/nh3-pve-2/README.md`
- **esh-pve-2** (MS-03, ESH) is unplugged by Prime on purpose. vmstore + AMT phone-home are done. The permanent address is deferred by Prime ("address can wait"; `servers/esh-pve-2/README.md`). Run `playbooks/pve-nag-patch.yaml` on it when it is back.
- **cira-tunnel-watchdog** on pfi-tacticalrmm has been live since 1232 and drops :4433 tunnels silent >180 s. Watch `journalctl -t cira-tunnel-watchdog`.
- **RustDesk server (ana-docker) updated 1.1.14 → 1.1.16 at 1424** (same key; ports 21115-21119 reachable via rustdesk.phasefinal.com). Private key vaulted first as `ana-docker/rustdesk/id_ed25519`. "Change ID" is refused by the OSS server BY DESIGN (Pro feature; `rendezvous_server.rs` answers TCP RegisterPk with NOT_SUPPORT). Workaround, read at client source and untested: stop RustDesk, delete `enc_id` from RustDesk.toml, set a plain `id`, restart. → `servers/ana-docker/README.md`
- **Booth vs 10.100.10.166 (Prime, ~1420): that client cannot reach the Booth.** The Booth was up and answered from 4 hosts at 1417; .166 is on nh3-dev's own subnet. Prime interrupted my check, so it is OPEN and resumes only on his word.
- **esh-nas-pve → PVE 9: Prime asked (1440) whether I can do it myself. Assessed read-only and ANSWERED; awaiting his call.** Facts: QNAP hardware, no AMT/IPMI (no remote console); PVE 8.4.20 but running kernel 6.8.12-13; up since 2026-08-18; ZFS root `nvme` booted by GRUB; systemd-boot meta present; NO NIC pins (uplink `enp5s0f0`); no DKMS. My answer: I can do all the prep remotely (vzdump, NIC pin, remove systemd-boot, root snapshot, pve8to9, a reboot test on PVE 8 first). The upgrade reboot should happen only with someone able to reach ESH, or with Prime's explicit acceptance of a no-console failure. VMs 100/101 on `pve` must be shut down first (hard NFS mounts). → `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`
- **Done today, nothing open:**
- tank verified healthy 0629, disk kept, Miranda told;
- albok-service 0.1.2 + the Nemi hourly timer;
- the post-deletion Worldtree gate (infra-hermes runs it as a regression check);
- the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve.
- the PVE subscription popup patched on nh3-pve, nh3-pve-2, pfi-pve and esh-pve;
- nh3-pve PVE 9 + its post-fixes; nh3-dev's VM 102 snapshot deleted.
- **ESH PVE 8→9 plan:** written, NOT executed (needs Prime's green light). `docs/runbooks/esh-pve-cluster-pve9-upgrade-plan.md`. The tank blocker is cleared.
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
+1 -1
View File
@@ -43,7 +43,7 @@ General-purpose Docker host for the Anaheim colo. Runs everything at `10.250.0.0
| openwebui | 3100 | Chat UI frontend |
| sillytavern | 8100 | Chat UI |
| mailrise | 8025 | SMTP-to-notification gateway |
| rustdesk (hbbs + hbbr) | host-net 21115-21119 | Remote desktop relay — `rustdesk.phasefinal.com`, server **1.1.16** (updated 2026-10-03 from 1.1.14; `:latest`, image 2026-07-20), public key `vVIChMRsCxCqxjg7h1rDrrxqh5zGpS6ghs+814xV6pI=` (from `id_ed25519.pub` in volume `rustdesk_rustdesk_data`). Client install steps: Booth `rustdesk-client-setup` (kept). Private key vaulted `ana-docker/rustdesk/id_ed25519` (2026-10-03) |
| rustdesk (hbbs + hbbr) | host-net 21115-21119 | Remote desktop relay — `rustdesk.phasefinal.com`, server **1.1.16** (updated 2026-10-03 from 1.1.14; `:latest`, image 2026-07-20), public key `vVIChMRsCxCqxjg7h1rDrrxqh5zGpS6ghs+814xV6pI=` (from `id_ed25519.pub` in volume `rustdesk_rustdesk_data`). Client install steps: Booth `rustdesk-client-setup` (kept). Private key vaulted `ana-docker/rustdesk/id_ed25519` (2026-10-03). ⚠ The client's "Change ID" is refused by this OSS server by design (a Pro feature). Workaround from the client source (`hbb_common` `Config::load`), untested: stop RustDesk, delete `enc_id` from `RustDesk.toml`, set `id = '<new>'`, restart; the client re-encrypts it into `enc_id` |
| dockge | 5001 | Docker stack management UI |
| beszel | 8090 | Fleet metrics hub (+ local agent); agents on the other hosts report here |
| dozzle (hub as `dozzle-hub`) | 8088 | Fleet log viewer; agents on the other hosts report here |