# `[2026-08-19]` esh-pve hard-froze for 4.5h — and took the whole house's DNS with it Reported by the operator as "routing or DNS issues on the PVC wifi." It was neither: the internet was healthy the entire time (gateway reporting 3 ms and 209/26 Mbps; 1.1.1.1 and 8.8.8.8 answering at ~3 ms from inside ESH with zero loss). **The house had no name resolution because one VM was down.** ## The SPOF: one resolver, cross-VLAN, no fallback `esh-userland` (VLAN 10, `10.0.10.0/24` — the `PVC` SSID *and* the wired userland LAN) handed out **exactly one DNS server, `10.0.50.45`** — AdGuard, on `esh-docker-vm`, on the **server** VLAN. No secondary. That VM dies, every client on the VLAN loses DNS, and it presents as "the wifi is broken." It was the only network in the house exposed this way. `Default`, `esh-mgmt`, `esh-server` and `esh-cameras` run DNS on auto (the gateway hands itself out); `esh-iot` and `ESH-WG` point at 1.1.1.1 + 8.8.8.8. **Fixed** (operator-approved): `esh-userland` now hands out `10.0.50.45` primary, **`10.0.10.1` (the gateway) secondary** — the UDM's own resolver, verified answering. Applied via the Classic API, `PUT /proxy/network/api/s/default/rest/networkconf/687985eae5d15b673cef1a73` with the full object (GET → modify one field → PUT), `rc: ok`. **This was also the first confirmed WRITE on the ESH UDM key** — previously only the NH3 key was write-tested. See [[reference_unifi_udm_integration_api_keys]]. ⚠️ **A secondary is not clean failover.** macOS/iOS query resolvers in parallel, so once AdGuard is back a real share of lookups go to the gateway and **skip ad-blocking**. This converts a total outage into degraded-but-working. The actual fix for blocking integrity is a second AdGuard instance NOT on esh-pve. ## Root cause: hard freeze, no diagnostics, two suspects `esh-pve` (Minisforum MS-01, i9-13900H, `productname: YajuuSenpai`) froze at **03:34:39**. The journal stops mid-operation — **no panic, no OOM, no MCE, no thermal event**. Powered on with its 10G link up, but not answering ARP. Two changes landed the day before, and they are not exclusive: 1. **New kernel.** A large `apt` batch on **2026-08-18 07:00:21** installed `proxmox-kernel-6.8.12-42-pve`; clean reboot at 07:08:44. Before that the box had **4.5 months of uptime** (Mar 30 → Aug 18) on `6.8.12-16`. First boot on the new kernel lasted **20 hours**. 2. **GPU passthrough.** The last kernel messages of the dead boot are `vfio-pci 0000:01:00.0/.1: enabling device` at **02:55:17** — VM 102 `esh-vm-workstation` starting with `hostpci0: 0000:01:00,pcie=1,x-vga=1`, **39 minutes before the freeze**. A vfio/i915 regression in the newer kernel would produce exactly this signature. `6.8.12-16` is still installed and is the held-in-reserve rollback. **VM 102 is now pinned off** (`qm set 102 --onboot 0`, stopped) per the operator — it is on-demand and there has been no demand. That removes the suspect without a kernel rollback. ## Why nobody could recover it remotely — and the fix Nothing on the box could reboot it: - **`softdog` was the loaded watchdog.** A *software* watchdog cannot rescue a hard kernel freeze: the frozen kernel is the thing that would have to fire its timer. This is the trap — the machine *looked* watchdog-protected. - **Proxmox's `watchdog-mux` held `/dev/watchdog` but never armed it.** It only pets the device while an HA client is connected, and this cluster has no HA resources. - **vPro/AMT was unusable.** The MS-01 reaches the network only via **SFP+** (Intel X710, port 27 on the Garage switch) and presents exactly one MAC. **AMT cannot ride a discrete/SFP+ NIC** — it needs the chipset-integrated Intel PHY, i.e. one of the two i226 RJ45 ports, and both are unplugged. Cabling one and provisioning AMT in MEBx remains the open item for *control*; the watchdog below is the fix for *recovery*. **Fixed:** `playbooks/esh-pve-hardware-watchdog.yaml` — systemd now owns the PCH hardware watchdog (`iTCO_wdt`, `RuntimeWatchdogSec=60`), `softdog` is blacklisted and unloaded, `watchdog-mux` is masked. Verified live: `watchdog0: identity=iTCO_wdt state=active timeout=60s`, held by PID 1, journal `Using hardware watchdog 'iTCO_wdt', version 6`. Playbook re-run proves idempotency (6 skipped / 6 verify OK). Firmware does **not** block the TCO timer here — checked for the `unable to reset NO_REBOOT flag` line before committing to the approach; the board reports `Found a Intel PCH TCO device (Version=6, TCOBASE=0x0400)`. ⚠️ **Masking `watchdog-mux` trades away HA fencing.** If Proxmox HA is ever configured on esh-pve this must be reverted. Not a near-term concern: `esh-pve-cluster` is **two nodes with no qdevice**, so a single node loss already costs quorum and the survivor would fence itself — HA here would reduce availability, not raise it. ⚠️ **The watchdog is configured and armed, but has NOT been proven to fire.** Proving it means deliberately wedging the host. Untested-but-armed is still strictly better than softdog; treat a real firing as unconfirmed until tested. ## Diagnostic corrections worth keeping - **"No route to host" was the dead host, not a routing gap.** Two claims made mid-incident were wrong: that the mgmt VLAN (`10.0.250.0/24`) is not routed over the NH3↔ESH tunnel, and that a firewall isolates it from the server VLAN. Both were artifacts of esh-pve being dead. With it up, `root@esh-pve` SSHes fine from nh3-dev at 7.5 ms, and `10.0.250.1` answers from `esh-pve-nas` in 0.078 ms. **Control-test against a *different* host on the target subnet before concluding "the subnet is unreachable."** - **UDM `uptime` on a client record is association time, not host uptime.** It read 2.2 days while the host had been up 20 hours. Use `journalctl --list-boots` on the host for real boot history. - **`rest/user` `last_seen` is not maintained** (it read ~203 days for hosts that are demonstrably online). `stat/sta` is the live view.