diff --git a/services/esh-vm-docker-watchdog/README.md b/services/esh-vm-docker-watchdog/README.md new file mode 100644 index 0000000..3d0dd30 --- /dev/null +++ b/services/esh-vm-docker-watchdog/README.md @@ -0,0 +1,94 @@ +# esh-vm-docker-watchdog — bound the NFS D-state wedge + +Runs on **esh-pve (10.0.250.35)**, watching **esh-vm-docker (VMID 100 / 10.0.50.45)**. +A watchdog living inside the thing it watches is no watchdog — this one sits on +the hypervisor, which is also the only place that can actually recover the VM. + +## The failure it exists for + +esh-vm-docker mounts NFS from the NAS at 10.0.50.50 with **`hard`** semantics. When +the NAS stalls, I/O blocks in uninterruptible sleep — the D-state — and nothing +inside the guest clears it. Reconfirmed 2026-07-15: `docker stop/rm -f`, +`ctr -n moby task delete`, `systemctl restart docker`, and even `docker exec` +all fail. Only a host reset recovers. Most recent occurrence: 2026-08-16. + +The pre-existing `x-systemd.before=docker.service` fstab fix addresses the **boot +race** — a different bug. It does nothing for a **runtime** stall, which is the +one that keeps biting. + +This watchdog does **not prevent** the wedge. It bounds the outage from "until +someone notices" to roughly ten minutes. + +## Why the probe is an HTTP check and not ping/SSH + +The wedge signature is specifically **"guest OS alive, services dead."** The root +filesystem is local disk, so sshd keeps accepting connections and ICMP keeps +answering straight through a total service outage. **A ping or TCP check reports +HEALTHY while the box is unusable** — which is exactly how the outage goes +unnoticed in the first place. + +So the trigger is an HTTP probe of traefik (the ingress every service sits +behind). The qemu-guest-agent ping is recorded only to *classify* the failure in +the log — agent alive + HTTP dead is the textbook wedge — and never vetoes a +reset. + +## Behaviour + +| knob | default | why | +|---|---|---| +| probe | `http://10.0.50.45:80/` | any HTTP response counts as alive; a 404 from traefik still proves it is serving | +| `PROBE_TIMEOUT` | 10s | bounded so the probe itself can never hang the watchdog | +| interval | 2 min | systemd timer | +| `FAIL_THRESHOLD` | 5 | ~10 min sustained outage before acting — a traefik redeploy cannot trigger a power cycle | +| `COOLDOWN_SEC` | 1800 | a genuinely broken VM cannot be reset-looped | + +Safety rails: resets only when `qm status` reports `running` (a deliberately +stopped VM is left alone), and honours a disable flag. + +```bash +# suspend during planned maintenance +touch /etc/esh-vm-docker-watchdog.disabled +rm /etc/esh-vm-docker-watchdog.disabled # re-arm +``` + +Log: `/var/log/esh-vm-docker-watchdog.log`. State: +`/var/lib/esh-vm-docker-watchdog/`. + +## Install + +```bash +scp esh-vm-docker-watchdog.sh root@10.0.250.35:/usr/local/sbin/ +scp esh-vm-docker-watchdog.{service,timer} root@10.0.250.35:/etc/systemd/system/ +ssh root@10.0.250.35 'chmod 0755 /usr/local/sbin/esh-vm-docker-watchdog.sh && + systemctl daemon-reload && systemctl enable --now esh-vm-docker-watchdog.timer' +``` + +## Verify without power-cycling anything + +The failure path can be exercised safely by pointing the probe at a closed port +and raising the threshold out of reach: + +```bash +PROBE_URL=http://10.0.50.45:9/ FAIL_THRESHOLD=99 PROBE_TIMEOUT=3 \ + /usr/local/sbin/esh-vm-docker-watchdog.sh +tail -3 /var/log/esh-vm-docker-watchdog.log +``` + +Validated this way on install (2026-08-16): healthy → silent no-op; disable flag +→ `SKIP`; simulated outage → counts and classifies as the D-state signature; +recovery → counter cleared. VM never reset during testing. + +## What it does NOT cover + +`/mnt/books` is still a **`hard`** mount and is still able to wedge the box — +deliberately. Calibre's library holds a SQLite `metadata.db`, and `soft` / +`softerr` semantics risk corrupting it on a stall mid-write. Operator ruling +2026-08-16: keep `hard`, accept the wedge, bound it with this watchdog. Revisit +only alongside moving the calibre library off NFS. + +`/mnt/backup` (restic, daily 01:00) is also still `hard`, but is touched only by +a scheduled job rather than a long-running container, so a stall there wedges +restic and not dockerd. + +Related: `playbooks/fix-esh-nfs-boot-ordering.yaml` (the boot-race fix), +park item `harden-esh-docker-vm-against-the-recurring-nfs` (id 28). diff --git a/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.service b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.service new file mode 100644 index 0000000..5d720aa --- /dev/null +++ b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.service @@ -0,0 +1,11 @@ +[Unit] +Description=Watchdog for esh-vm-docker (VMID 100) — reset on NFS D-state wedge +Documentation=https://gitea.phasefinal.com/vh/esh-pfi-infrastructure +After=pve-guests.service +Wants=network-online.target + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/esh-vm-docker-watchdog.sh +# Probe + qm calls are all bounded; this is a backstop against a stuck run. +TimeoutStartSec=120 diff --git a/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.sh b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.sh new file mode 100644 index 0000000..f99b53b --- /dev/null +++ b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.sh @@ -0,0 +1,107 @@ +#!/bin/bash +# esh-vm-docker-watchdog — bound the NFS D-state wedge on esh-vm-docker (VMID 100). +# +# WHY THIS EXISTS +# esh-vm-docker mounts /mnt/books from the NAS at 10.0.50.50 with `hard` NFS +# semantics (required: calibre's library holds a SQLite metadata.db, and soft +# semantics risk corrupting it). When the NAS stalls, I/O blocks in +# uninterruptible sleep — the classic D-state — and NOTHING clears it from +# inside the guest. Confirmed 2026-07-15: `docker stop/rm -f`, `ctr task +# delete`, `systemctl restart docker`, and even `docker exec` all fail. Only a +# host reset recovers. This watchdog does not PREVENT the wedge; it bounds the +# outage from "until someone notices" to a few minutes. +# +# THE PROBE, AND WHY IT IS SHAPED THIS WAY +# The wedge signature is specifically "guest OS alive, services dead" — sshd +# keeps answering because / is local disk, so an SSH or ping check reports +# HEALTHY through a total service outage. That is why the trigger is an HTTP +# probe of traefik (the ingress every service sits behind) and NOT a TCP/ping +# check. The guest-agent ping is recorded only to CLASSIFY the failure in the +# log, never to veto a reset. +# +# SAFETY RAILS +# - Resets only when `qm status` reports the VM `running`. A deliberately +# stopped VM is left alone. +# - Requires FAIL_THRESHOLD consecutive failures, so a traefik redeploy or a +# brief blip cannot trigger a power cycle. +# - COOLDOWN_SEC between resets, so a genuinely broken VM cannot be +# reset-looped forever. +# - Touch the disable flag to suspend it during planned maintenance: +# touch /etc/esh-vm-docker-watchdog.disabled +# +# Install: see README.md in this directory. Runs on esh-pve (10.0.250.35), NOT +# on the guest — a watchdog living inside the thing it watches is no watchdog. + +set -uo pipefail + +VMID="${VMID:-100}" +TARGET="${TARGET:-10.0.50.45}" +PROBE_URL="${PROBE_URL:-http://${TARGET}:80/}" +PROBE_TIMEOUT="${PROBE_TIMEOUT:-10}" +FAIL_THRESHOLD="${FAIL_THRESHOLD:-5}" +COOLDOWN_SEC="${COOLDOWN_SEC:-1800}" +STATE_DIR="${STATE_DIR:-/var/lib/esh-vm-docker-watchdog}" +LOG="${LOG:-/var/log/esh-vm-docker-watchdog.log}" +DISABLE_FLAG="${DISABLE_FLAG:-/etc/esh-vm-docker-watchdog.disabled}" + +mkdir -p "$STATE_DIR" +FAILFILE="$STATE_DIR/consecutive_failures" +LASTRESET="$STATE_DIR/last_reset_epoch" +[ -f "$FAILFILE" ] || echo 0 > "$FAILFILE" +[ -f "$LASTRESET" ] || echo 0 > "$LASTRESET" + +log() { printf '%s %s\n' "$(date -Is)" "$*" >> "$LOG"; } + +if [ -f "$DISABLE_FLAG" ]; then + echo 0 > "$FAILFILE" + log "SKIP disabled by $DISABLE_FLAG" + exit 0 +fi + +# Only act on a VM that is SUPPOSED to be up. +status=$(qm status "$VMID" 2>/dev/null | awk '{print $2}') +if [ "$status" != "running" ]; then + echo 0 > "$FAILFILE" + log "SKIP vm $VMID status=${status:-unknown} (not running)" + exit 0 +fi + +# Any HTTP response at all means the ingress is serving. --max-time bounds it so +# the probe itself can never hang the watchdog. +if curl -sS -o /dev/null --max-time "$PROBE_TIMEOUT" "$PROBE_URL" 2>/dev/null; then + prev=$(cat "$FAILFILE") + echo 0 > "$FAILFILE" + [ "$prev" -gt 0 ] && log "OK probe recovered after $prev consecutive failure(s)" + exit 0 +fi + +fails=$(( $(cat "$FAILFILE") + 1 )) +echo "$fails" > "$FAILFILE" + +# Classification only — an agent that still answers while HTTP is dead is the +# textbook D-state wedge, and is the case we most want in the log. +if qm agent "$VMID" ping >/dev/null 2>&1; then + kind="guest-agent ALIVE (services wedged — D-state signature)" +else + kind="guest-agent DEAD (guest hung or down)" +fi +log "FAIL $fails/$FAIL_THRESHOLD probe=$PROBE_URL $kind" + +[ "$fails" -ge "$FAIL_THRESHOLD" ] || exit 0 + +now=$(date +%s) +since=$(( now - $(cat "$LASTRESET") )) +if [ "$since" -lt "$COOLDOWN_SEC" ]; then + log "HOLD threshold reached but only ${since}s since last reset (cooldown ${COOLDOWN_SEC}s) — NOT resetting" + exit 0 +fi + +log "RESET issuing 'qm reset $VMID' after $fails consecutive failures — $kind" +if qm reset "$VMID" >/dev/null 2>&1; then + echo "$now" > "$LASTRESET" + echo 0 > "$FAILFILE" + log "RESET ok" +else + log "RESET FAILED — qm reset returned non-zero; manual intervention needed" + exit 1 +fi diff --git a/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.timer b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.timer new file mode 100644 index 0000000..df06cb6 --- /dev/null +++ b/services/esh-vm-docker-watchdog/esh-vm-docker-watchdog.timer @@ -0,0 +1,14 @@ +[Unit] +Description=Run the esh-vm-docker wedge watchdog every 2 minutes +Documentation=https://gitea.phasefinal.com/vh/esh-pfi-infrastructure + +[Timer] +# 5 consecutive failures x 2 min => ~10 minutes of sustained outage before a +# reset. Long enough that a traefik redeploy cannot trigger a power cycle. +OnBootSec=5min +OnUnitActiveSec=2min +AccuracySec=15s +Unit=esh-vm-docker-watchdog.service + +[Install] +WantedBy=timers.target