esh-pve hard-froze at 03:34 on 2026-08-19 and stayed frozen ~4.5 hours until a manual power cycle. No panic, no OOM, no MCE — the journal stops mid-operation. The whole ESH site lost DNS with it, because esh-userland (VLAN 10, the PVC SSID and wired userland LAN) was handed exactly one resolver: 10.0.50.45, AdGuard on esh-docker-vm, on a different VLAN, with no secondary. Internet and routing were healthy throughout. Two fixes. 1. DNS: 10.0.10.1 (the gateway, verified resolving) added as secondary on esh-userland via the UDM Classic API. Note this is degradation cover, not clean failover — clients that query resolvers in parallel will bypass AdGuard for a share of lookups. 2. Watchdog: softdog -> iTCO_wdt under systemd (RuntimeWatchdogSec=60), watchdog-mux masked. The box looked watchdog-protected and was not: a software watchdog cannot fire when the kernel it lives in is wedged, and watchdog-mux only pets the device while an HA client is connected, which never happens on a cluster with no HA resources. Firmware does not block the TCO timer here, checked before committing to it. Also pins VM 102 off (onboot: 0). It starts with full GPU passthrough and vfio-pci enabling that device is the last thing the kernel logged, 39 minutes before the freeze. The other suspect is the kernel itself: the host ran 4.5 months on 6.8.12-16, took 6.8.12-42 in an apt batch on 08-18, and died 20 hours into the first boot on it. 6.8.12-16 is still installed as the rollback. The playbook is idempotent — a second run skips all six steps and passes all six verifies. The watchdog is confirmed armed (identity=iTCO_wdt, state=active, held by PID 1) but has NOT been observed firing; proving that needs a deliberate wedge. Memory also corrects two wrong mid-incident calls: the mgmt VLAN is routed over the site tunnel and is not firewalled off — both symptoms were the dead host generating ICMP unreachables.
5.9 KiB
[2026-08-19] esh-pve hard-froze for 4.5h — and took the whole house's DNS with it
Reported by the operator as "routing or DNS issues on the PVC wifi." It was neither: the internet was healthy the entire time (gateway reporting 3 ms and 209/26 Mbps; 1.1.1.1 and 8.8.8.8 answering at ~3 ms from inside ESH with zero loss). The house had no name resolution because one VM was down.
The SPOF: one resolver, cross-VLAN, no fallback
esh-userland (VLAN 10, 10.0.10.0/24 — the PVC SSID and the wired
userland LAN) handed out exactly one DNS server, 10.0.50.45 — AdGuard, on
esh-docker-vm, on the server VLAN. No secondary. That VM dies, every
client on the VLAN loses DNS, and it presents as "the wifi is broken."
It was the only network in the house exposed this way. Default, esh-mgmt,
esh-server and esh-cameras run DNS on auto (the gateway hands itself out);
esh-iot and ESH-WG point at 1.1.1.1 + 8.8.8.8.
Fixed (operator-approved): esh-userland now hands out 10.0.50.45
primary, 10.0.10.1 (the gateway) secondary — the UDM's own resolver,
verified answering. Applied via the Classic API,
PUT /proxy/network/api/s/default/rest/networkconf/687985eae5d15b673cef1a73
with the full object (GET → modify one field → PUT), rc: ok. This was also
the first confirmed WRITE on the ESH UDM key — previously only the NH3 key
was write-tested. See reference_unifi_udm_integration_api_keys.
⚠️ A secondary is not clean failover. macOS/iOS query resolvers in parallel, so once AdGuard is back a real share of lookups go to the gateway and skip ad-blocking. This converts a total outage into degraded-but-working. The actual fix for blocking integrity is a second AdGuard instance NOT on esh-pve.
Root cause: hard freeze, no diagnostics, two suspects
esh-pve (Minisforum MS-01, i9-13900H, productname: YajuuSenpai) froze at
03:34:39. The journal stops mid-operation — no panic, no OOM, no MCE, no
thermal event. Powered on with its 10G link up, but not answering ARP.
Two changes landed the day before, and they are not exclusive:
- New kernel. A large
aptbatch on 2026-08-18 07:00:21 installedproxmox-kernel-6.8.12-42-pve; clean reboot at 07:08:44. Before that the box had 4.5 months of uptime (Mar 30 → Aug 18) on6.8.12-16. First boot on the new kernel lasted 20 hours. - GPU passthrough. The last kernel messages of the dead boot are
vfio-pci 0000:01:00.0/.1: enabling deviceat 02:55:17 — VM 102esh-vm-workstationstarting withhostpci0: 0000:01:00,pcie=1,x-vga=1, 39 minutes before the freeze.
A vfio/i915 regression in the newer kernel would produce exactly this
signature. 6.8.12-16 is still installed and is the held-in-reserve rollback.
VM 102 is now pinned off (qm set 102 --onboot 0, stopped) per the
operator — it is on-demand and there has been no demand. That removes the
suspect without a kernel rollback.
Why nobody could recover it remotely — and the fix
Nothing on the box could reboot it:
softdogwas the loaded watchdog. A software watchdog cannot rescue a hard kernel freeze: the frozen kernel is the thing that would have to fire its timer. This is the trap — the machine looked watchdog-protected.- Proxmox's
watchdog-muxheld/dev/watchdogbut never armed it. It only pets the device while an HA client is connected, and this cluster has no HA resources. - vPro/AMT was unusable. The MS-01 reaches the network only via SFP+ (Intel X710, port 27 on the Garage switch) and presents exactly one MAC. AMT cannot ride a discrete/SFP+ NIC — it needs the chipset-integrated Intel PHY, i.e. one of the two i226 RJ45 ports, and both are unplugged. Cabling one and provisioning AMT in MEBx remains the open item for control; the watchdog below is the fix for recovery.
Fixed: playbooks/esh-pve-hardware-watchdog.yaml — systemd now owns the
PCH hardware watchdog (iTCO_wdt, RuntimeWatchdogSec=60), softdog is
blacklisted and unloaded, watchdog-mux is masked. Verified live:
watchdog0: identity=iTCO_wdt state=active timeout=60s, held by PID 1,
journal Using hardware watchdog 'iTCO_wdt', version 6. Playbook re-run proves
idempotency (6 skipped / 6 verify OK).
Firmware does not block the TCO timer here — checked for the
unable to reset NO_REBOOT flag line before committing to the approach; the
board reports Found a Intel PCH TCO device (Version=6, TCOBASE=0x0400).
⚠️ Masking watchdog-mux trades away HA fencing. If Proxmox HA is ever
configured on esh-pve this must be reverted. Not a near-term concern:
esh-pve-cluster is two nodes with no qdevice, so a single node loss
already costs quorum and the survivor would fence itself — HA here would reduce
availability, not raise it.
⚠️ The watchdog is configured and armed, but has NOT been proven to fire. Proving it means deliberately wedging the host. Untested-but-armed is still strictly better than softdog; treat a real firing as unconfirmed until tested.
Diagnostic corrections worth keeping
- "No route to host" was the dead host, not a routing gap. Two claims made
mid-incident were wrong: that the mgmt VLAN (
10.0.250.0/24) is not routed over the NH3↔ESH tunnel, and that a firewall isolates it from the server VLAN. Both were artifacts of esh-pve being dead. With it up,root@esh-pveSSHes fine from nh3-dev at 7.5 ms, and10.0.250.1answers fromesh-pve-nasin 0.078 ms. Control-test against a different host on the target subnet before concluding "the subnet is unreachable." - UDM
uptimeon a client record is association time, not host uptime. It read 2.2 days while the host had been up 20 hours. Usejournalctl --list-bootson the host for real boot history. rest/userlast_seenis not maintained (it read ~203 days for hosts that are demonstrably online).stat/stais the live view.