diff --git a/persistent-memory.md b/persistent-memory.md index 2d7378d..4ca4fa7 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -128,13 +128,12 @@ went on live, no esh-pve reboot. Full record → Recent decisions. - **Not wired yet:** esh-ml1 in Homepage `docker.yaml` (needs dockerd tcp/2375) and Beszel. The ~18.7 s host-down failover penalty is untuned. -### ⚠ homepage wedged on esh-docker-vm — needs a VM reboot (Prime's call) +### ✅ homepage wedge cleared — esh-docker-vm rebooted 2026-09-25 0613 (Prime) -3rd unkillable wedge (06-03, 09-18, now 2026-09-24 ~2224 PT): node thread D in -`vm_mmap_pgoff`, then `exit_mmap` after a kill; `docker restart` failed. Rest -of the VM healthy. The reboot blips ESH DNS. Signature + safe diagnosis -commands: `servers/esh-docker-vm/README.md`. Option to move Homepage to -ana-docker surfaced to Prime. +Prime ordered the reboot and ruled **do NOT move Homepage** (stays on +esh-docker-vm). Down 0613:44–0616:58 PT; homepage needed a manual +`docker start` (a failed restart had marked it stopped). Signature + recovery: +`servers/esh-docker-vm/README.md`. ### pfi-gx10 is OFF — AC-pull test 2026-09-25 (Prime) diff --git a/servers/esh-docker-vm/README.md b/servers/esh-docker-vm/README.md index 3ff2b64..459ae5f 100644 --- a/servers/esh-docker-vm/README.md +++ b/servers/esh-docker-vm/README.md @@ -97,7 +97,14 @@ state too. Read `/proc//task/*/stat` and `sudo cat .../stack` instead (`ssh infra-ops@10.0.50.45`, NOPASSWD). **Only a VM reboot clears it**, and this VM is ESH's DNS resolver, so the -reboot is a short ESH-wide DNS outage — schedule it. Root cause not +reboot is a short ESH-wide DNS outage — schedule it. + +⚠ **After the reboot, homepage does NOT come back by itself** if anyone tried +`docker restart`/`stop` on it first: docker records it as stopped on purpose, +so `restart: unless-stopped` skips it. Run `docker start homepage`. Measured +2026-09-25: `qm reboot 100` from esh-pve took 3m18s (the graceful shutdown +stalls on the wedged process, then proceeds); the VM was unreachable +0613:44–0616:58 PT; every other container came back unaided. Root cause not established; a kernel-side mmap_lock problem is the leading guess, not a finding. Moving Homepage to ana-docker (as was done for Uptime Kuma on 2026-09-21 for the same box's history) would take the dashboard out of this