diff --git a/servers/esh-pve-nas/README.md b/servers/esh-pve-nas/README.md index c3b0e93..94e85d9 100644 --- a/servers/esh-pve-nas/README.md +++ b/servers/esh-pve-nas/README.md @@ -51,6 +51,15 @@ Media services on this box sit at `10.0.50.56` (Plex) and `10.0.50.57` (Jellyfin - **Pools:** `nvme` (2× 931 GB NVMe mirror — 32 G used, 867 G free, holds every guest rootfs), `ssd` (4× 894 GB Intel SATA, 2 mirrors — 1.42 T free), `tank` (12× 14.6 TB raidz2 ×2 — 40 T of 175 T). Plus NFS `/mnt/pve/tank-vmbu` for VM backups. +- ⚠⚠ **THIS HOST IS HALF OF A 2-NODE PROXMOX CLUSTER** — `esh-pve-cluster`, nodes + `pve` (esh-pve, 10.0.250.35, nodeid 1) and `esh-nas-pve` (this host, nodeid 2). + `Expected votes: 2`, `Quorum: 2`, no qdevice. **Taking either node down drops the + survivor below quorum and makes its `/etc/pve` read-only** — guests keep running, + but no config change, no VM start/stop, no storage edit works until the partner + returns. This was undocumented and nearly bit the 2026-08-18 migration, which + rebooted this node without accounting for it. Before any planned reboot of either + node, either accept the read-only window on the survivor or set + `pvecm expected 1` on it for the duration. Discovered 2026-08-18. - ⚠ **CT 103 `esh-nas` (10.0.50.50) runs on THIS host and serves `hard` NFS.** Never reboot this host casually — quiesce the clients first. Measured from the server on 2026-08-18 (`ss -tn '( sport = :2049 )'` inside CT 103), there are diff --git a/servers/esh-pve/README.md b/servers/esh-pve/README.md index 5828feb..fdb1fdc 100644 --- a/servers/esh-pve/README.md +++ b/servers/esh-pve/README.md @@ -33,3 +33,25 @@ Same Proxmox-inspect caveat as `pfi-pve`. ## Placement rule Hypervisor for ESH home-lab VMs. Not part of the PFI colo topology. + +## ⚠ Cluster membership and the dark-tile gotcha + +This node (`pve`) is half of the 2-node **`esh-pve-cluster`** with `esh-nas-pve` +(esh-pve-nas, 10.0.50.55). `Expected votes: 2`, `Quorum: 2`, no qdevice — so +**rebooting either node makes the survivor's `/etc/pve` read-only** until the +partner is back. Guests keep running; config changes do not. Plan reboots of +either node accordingly (`pvecm expected 1` on the survivor, or accept the +read-only window). + +⚠ **A dark/greyed node tile usually means `pvestatd`, not a dead node.** +`pvestatd` SEGV'd here on **2026-05-28 and stayed dead for 82 days** — the node +was quorate and healthy the entire time, with `pve-cluster`, `corosync`, +`pvedaemon`, `pveproxy` and both HA services active and all three guests running. +It is only the *reporting* daemon, so its death is invisible except that the UI +has nothing to render. It has SEGV'd four times (2025-08-18, 2025-09-04, +2025-09-15, 2026-05-28) — treat a recurrence as expected, not novel. + + systemctl reset-failed pvestatd && systemctl restart pvestatd + +Restarted 2026-08-18. Worth a watchdog: nothing alerts when it dies, and the only +symptom is a cosmetic one nobody looks at for months.