Files
esh-pfi-infrastructure/persistent-memory.d/2026-08-19-esh-pve-freeze-dns-spof.md
vh b92097688c fix(esh-pve): hardware watchdog, and close the single-resolver DNS SPOF
esh-pve hard-froze at 03:34 on 2026-08-19 and stayed frozen ~4.5 hours
until a manual power cycle. No panic, no OOM, no MCE — the journal stops
mid-operation. The whole ESH site lost DNS with it, because esh-userland
(VLAN 10, the PVC SSID and wired userland LAN) was handed exactly one
resolver: 10.0.50.45, AdGuard on esh-docker-vm, on a different VLAN, with
no secondary. Internet and routing were healthy throughout.

Two fixes.

1. DNS: 10.0.10.1 (the gateway, verified resolving) added as secondary on
   esh-userland via the UDM Classic API. Note this is degradation cover,
   not clean failover — clients that query resolvers in parallel will
   bypass AdGuard for a share of lookups.

2. Watchdog: softdog -> iTCO_wdt under systemd (RuntimeWatchdogSec=60),
   watchdog-mux masked. The box looked watchdog-protected and was not: a
   software watchdog cannot fire when the kernel it lives in is wedged,
   and watchdog-mux only pets the device while an HA client is connected,
   which never happens on a cluster with no HA resources. Firmware does
   not block the TCO timer here, checked before committing to it.

Also pins VM 102 off (onboot: 0). It starts with full GPU passthrough and
vfio-pci enabling that device is the last thing the kernel logged, 39
minutes before the freeze. The other suspect is the kernel itself: the
host ran 4.5 months on 6.8.12-16, took 6.8.12-42 in an apt batch on
08-18, and died 20 hours into the first boot on it. 6.8.12-16 is still
installed as the rollback.

The playbook is idempotent — a second run skips all six steps and passes
all six verifies. The watchdog is confirmed armed (identity=iTCO_wdt,
state=active, held by PID 1) but has NOT been observed firing; proving
that needs a deliberate wedge.

Memory also corrects two wrong mid-incident calls: the mgmt VLAN is
routed over the site tunnel and is not firewalled off — both symptoms
were the dead host generating ICMP unreachables.
2026-08-19 09:05:52 -07:00

5.9 KiB

[2026-08-19] esh-pve hard-froze for 4.5h — and took the whole house's DNS with it

Reported by the operator as "routing or DNS issues on the PVC wifi." It was neither: the internet was healthy the entire time (gateway reporting 3 ms and 209/26 Mbps; 1.1.1.1 and 8.8.8.8 answering at ~3 ms from inside ESH with zero loss). The house had no name resolution because one VM was down.

The SPOF: one resolver, cross-VLAN, no fallback

esh-userland (VLAN 10, 10.0.10.0/24 — the PVC SSID and the wired userland LAN) handed out exactly one DNS server, 10.0.50.45 — AdGuard, on esh-docker-vm, on the server VLAN. No secondary. That VM dies, every client on the VLAN loses DNS, and it presents as "the wifi is broken."

It was the only network in the house exposed this way. Default, esh-mgmt, esh-server and esh-cameras run DNS on auto (the gateway hands itself out); esh-iot and ESH-WG point at 1.1.1.1 + 8.8.8.8.

Fixed (operator-approved): esh-userland now hands out 10.0.50.45 primary, 10.0.10.1 (the gateway) secondary — the UDM's own resolver, verified answering. Applied via the Classic API, PUT /proxy/network/api/s/default/rest/networkconf/687985eae5d15b673cef1a73 with the full object (GET → modify one field → PUT), rc: ok. This was also the first confirmed WRITE on the ESH UDM key — previously only the NH3 key was write-tested. See reference_unifi_udm_integration_api_keys.

⚠️ A secondary is not clean failover. macOS/iOS query resolvers in parallel, so once AdGuard is back a real share of lookups go to the gateway and skip ad-blocking. This converts a total outage into degraded-but-working. The actual fix for blocking integrity is a second AdGuard instance NOT on esh-pve.

Root cause: hard freeze, no diagnostics, two suspects

esh-pve (Minisforum MS-01, i9-13900H, productname: YajuuSenpai) froze at 03:34:39. The journal stops mid-operation — no panic, no OOM, no MCE, no thermal event. Powered on with its 10G link up, but not answering ARP.

Two changes landed the day before, and they are not exclusive:

  1. New kernel. A large apt batch on 2026-08-18 07:00:21 installed proxmox-kernel-6.8.12-42-pve; clean reboot at 07:08:44. Before that the box had 4.5 months of uptime (Mar 30 → Aug 18) on 6.8.12-16. First boot on the new kernel lasted 20 hours.
  2. GPU passthrough. The last kernel messages of the dead boot are vfio-pci 0000:01:00.0/.1: enabling device at 02:55:17 — VM 102 esh-vm-workstation starting with hostpci0: 0000:01:00,pcie=1,x-vga=1, 39 minutes before the freeze.

A vfio/i915 regression in the newer kernel would produce exactly this signature. 6.8.12-16 is still installed and is the held-in-reserve rollback.

VM 102 is now pinned off (qm set 102 --onboot 0, stopped) per the operator — it is on-demand and there has been no demand. That removes the suspect without a kernel rollback.

Why nobody could recover it remotely — and the fix

Nothing on the box could reboot it:

  • softdog was the loaded watchdog. A software watchdog cannot rescue a hard kernel freeze: the frozen kernel is the thing that would have to fire its timer. This is the trap — the machine looked watchdog-protected.
  • Proxmox's watchdog-mux held /dev/watchdog but never armed it. It only pets the device while an HA client is connected, and this cluster has no HA resources.
  • vPro/AMT was unusable. The MS-01 reaches the network only via SFP+ (Intel X710, port 27 on the Garage switch) and presents exactly one MAC. AMT cannot ride a discrete/SFP+ NIC — it needs the chipset-integrated Intel PHY, i.e. one of the two i226 RJ45 ports, and both are unplugged. Cabling one and provisioning AMT in MEBx remains the open item for control; the watchdog below is the fix for recovery.

Fixed: playbooks/esh-pve-hardware-watchdog.yaml — systemd now owns the PCH hardware watchdog (iTCO_wdt, RuntimeWatchdogSec=60), softdog is blacklisted and unloaded, watchdog-mux is masked. Verified live: watchdog0: identity=iTCO_wdt state=active timeout=60s, held by PID 1, journal Using hardware watchdog 'iTCO_wdt', version 6. Playbook re-run proves idempotency (6 skipped / 6 verify OK).

Firmware does not block the TCO timer here — checked for the unable to reset NO_REBOOT flag line before committing to the approach; the board reports Found a Intel PCH TCO device (Version=6, TCOBASE=0x0400).

⚠️ Masking watchdog-mux trades away HA fencing. If Proxmox HA is ever configured on esh-pve this must be reverted. Not a near-term concern: esh-pve-cluster is two nodes with no qdevice, so a single node loss already costs quorum and the survivor would fence itself — HA here would reduce availability, not raise it.

⚠️ The watchdog is configured and armed, but has NOT been proven to fire. Proving it means deliberately wedging the host. Untested-but-armed is still strictly better than softdog; treat a real firing as unconfirmed until tested.

Diagnostic corrections worth keeping

  • "No route to host" was the dead host, not a routing gap. Two claims made mid-incident were wrong: that the mgmt VLAN (10.0.250.0/24) is not routed over the NH3↔ESH tunnel, and that a firewall isolates it from the server VLAN. Both were artifacts of esh-pve being dead. With it up, root@esh-pve SSHes fine from nh3-dev at 7.5 ms, and 10.0.250.1 answers from esh-pve-nas in 0.078 ms. Control-test against a different host on the target subnet before concluding "the subnet is unreachable."
  • UDM uptime on a client record is association time, not host uptime. It read 2.2 days while the host had been up 20 hours. Use journalctl --list-boots on the host for real boot history.
  • rest/user last_seen is not maintained (it read ~203 days for hosts that are demonstrably online). stat/sta is the live view.