fix(headscale-ddns): say why it failed, retry the WAN lookup, and track it at all
The failed-START notifier built earlier today had its first REAL firing at
15:28: headscale-ddns.service exited 1 after succeeding all afternoon. The
detection worked. The alarm was also useless, and that is the finding.
Both failure paths exited 1 IN SILENCE, so the message said "exit status 1" and
nothing else. An alarm you cannot act on costs the same triage as no alarm at
all -- the notifier did its job and the subject script had no diagnostics for it
to carry.
Cause was transient and harmless: icanhazip.com did not answer inside its 10s
cap, so the IP came back empty and the regex guard refused it. No DNS impact --
the record already held the right address, verified against 1.1.1.1 before
touching anything, and the next timer run succeeded. Arithmetic confirms it:
~17s vault read + 10s curl timeout = 27s against the 28s the failing run took.
Fixed, both verified by making them fail:
- every exit path names its cause; a missing vault key names the key, an
EMPTY token is distinguished from a failed read, and a dead WAN lookup adds
"DNS left unchanged" because that is the fact the reader needs
- the WAN lookup retries 3x with ANNOUNCED attempts -- one third-party blip
should not page a human, and a silent retry would hide a degrading
dependency
⚠ ALSO: this script was not tracked anywhere. A fix to the thing every mesh
client resolves through lived on exactly one disk. Script, unit and timer are
in the repo now.
Measured and recorded: the vault read is 17 of the script's 18 seconds, every
10 minutes. Not a fault, but it bounds any retry budget and it is fleet-wide --
svos-dev's alarm unit carries the same 17-second note.
This commit is contained in:
@@ -0,0 +1,49 @@
|
||||
# headscale-ddns — keep `headscale.phasefinal.com` pointed at the NH3 WAN v4
|
||||
|
||||
A user timer on nh3-dev, every 10 minutes: read the WAN v4, PATCH the Cloudflare
|
||||
A record if it moved. It is how every mesh client finds the control plane, so a
|
||||
silent failure here eventually costs the mesh.
|
||||
|
||||
Installed at `~/.local/bin/headscale-ddns.sh` with user units
|
||||
`headscale-ddns.{service,timer}`. **Tracked here since 2026-09-22** — it was
|
||||
running untracked before that, so a fix to it lived on exactly one disk.
|
||||
|
||||
## The 2026-09-22 failure, and what it exposed
|
||||
|
||||
The unit failed at 15:28 (exit 1) after succeeding all afternoon. Cause was
|
||||
transient — `icanhazip.com` did not answer inside its 10s cap, so `$IP` came back
|
||||
empty and the regex guard refused it. No DNS impact: the record already held the
|
||||
right address and the next timer run succeeded.
|
||||
|
||||
**What actually mattered was that the alarm carried no cause.** Both failure
|
||||
paths were `|| exit 1` in silence, so the failed-START notifier fired correctly
|
||||
and said only "exit status 1". An alarm you cannot act on costs the same triage
|
||||
as no alarm.
|
||||
|
||||
Two fixes, both verified by making them fail:
|
||||
|
||||
- **Every exit path now says why** — a missing vault key names the key; an empty
|
||||
token is distinguished from a failed read; a dead WAN lookup says so and adds
|
||||
*"DNS left unchanged"*, which is the fact the reader needs.
|
||||
- **The WAN lookup retries 3×** with announced attempts. One third-party blip
|
||||
should not page a human, and a *silent* retry would hide a degrading
|
||||
dependency.
|
||||
|
||||
## ⚠ The vault read is 17 of the script's 18 seconds
|
||||
|
||||
Measured 2026-09-22: `secret get` takes **~17s**, everything else ~1s. It runs
|
||||
every 10 minutes. That is not a fault — but it bounds any retry budget here, and
|
||||
it is a fleet-wide cost worth knowing: svos-dev's own alarm unit carries the same
|
||||
17-second note. Any script in a timer that reads the vault pays it.
|
||||
|
||||
## Triage
|
||||
|
||||
```sh
|
||||
systemctl --user status headscale-ddns.service
|
||||
journalctl --user -u headscale-ddns.service -n 50 --no-pager
|
||||
dig +short headscale.phasefinal.com @1.1.1.1 # the thing that actually matters
|
||||
/home/lkraven/.local/bin/headscale-ddns.sh # safe to run by hand; idempotent
|
||||
```
|
||||
|
||||
A failed unit clears itself on the next timer run; it does not need
|
||||
`reset-failed` unless you want the state gone immediately.
|
||||
@@ -0,0 +1,15 @@
|
||||
[Unit]
|
||||
Description=Update headscale.phasefinal.com A record to NH3 WAN v4
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/home/lkraven/.local/bin/headscale-ddns.sh
|
||||
|
||||
# /home/lkraven/.config/systemd/user/headscale-ddns.service.d/10-onfailure-althing.conf
|
||||
# Installed as <unit>.service.d/10-onfailure-althing.conf on every fleet user
|
||||
# unit. A drop-in rather than an edit to the unit file so it is visible in
|
||||
# `systemctl --user cat`, reversible by deleting one file, and survives the
|
||||
# unit being reinstalled by its own deploy.
|
||||
#
|
||||
# %n is the failing unit's full name; the template receives it as %I.
|
||||
[Unit]
|
||||
OnFailure=althing-notify-failure@%n.service
|
||||
Executable
+28
@@ -0,0 +1,28 @@
|
||||
#!/bin/bash
|
||||
# Keep headscale.phasefinal.com's A record on NH3's current WAN v4. Token from the vault at run time.
|
||||
set -u
|
||||
SECRET=/home/lkraven/development/eshpfi-management/services/secrets-broker/secret
|
||||
# ⚠ SAY WHY. Both failure paths below used to `|| exit 1` in silence, and on
|
||||
# 2026-09-22 15:28 this unit failed for real: the failed-START alarm fired
|
||||
# correctly and carried NO CAUSE, because the script had printed nothing. An
|
||||
# alarm you cannot act on costs the same triage as no alarm at all.
|
||||
T=$($SECRET get nh3-dev/.config/cloudflare/infra-ops-dns-token 2>/dev/null) || {
|
||||
echo "FATAL: vault read failed for nh3-dev/.config/cloudflare/infra-ops-dns-token" >&2; exit 1; }
|
||||
[ -n "$T" ] || { echo "FATAL: vault returned an EMPTY token (read succeeded, value blank)" >&2; exit 1; }
|
||||
# The WAN lookup leans on a third party, so one blip should not page a human.
|
||||
# Measured 2026-09-22: the whole script takes ~18s of which ~17s is the vault
|
||||
# read, so three tries at a 10s cap is bounded and still well inside the 10-min
|
||||
# timer. Retries are announced -- a silent retry hides a degrading dependency.
|
||||
IP=""
|
||||
for try in 1 2 3; do
|
||||
IP=$(curl -s -m 10 -4 https://icanhazip.com | tr -d '[:space:]')
|
||||
[[ "$IP" =~ ^[0-9.]+$ ]] && break
|
||||
echo "WARN: WAN lookup attempt $try/3 returned [$IP]" >&2
|
||||
sleep 2
|
||||
done
|
||||
[[ "$IP" =~ ^[0-9.]+$ ]] || {
|
||||
echo "FATAL: could not determine WAN v4 after 3 attempts (icanhazip.com unreachable or returning junk); DNS left unchanged" >&2; exit 1; }
|
||||
ZID=$(curl -s -m 15 -H "Authorization: Bearer $T" "https://api.cloudflare.com/client/v4/zones?name=phasefinal.com" | python3 -c 'import sys,json; print(json.load(sys.stdin)["result"][0]["id"])')
|
||||
read -r RID CUR < <(curl -s -m 15 -H "Authorization: Bearer $T" "https://api.cloudflare.com/client/v4/zones/$ZID/dns_records?type=A&name=headscale.phasefinal.com" | python3 -c 'import sys,json; r=json.load(sys.stdin)["result"][0]; print(r["id"], r["content"])')
|
||||
[ "$CUR" = "$IP" ] && { echo "unchanged $IP"; exit 0; }
|
||||
curl -s -m 15 -X PATCH -H "Authorization: Bearer $T" -H "Content-Type: application/json" "https://api.cloudflare.com/client/v4/zones/$ZID/dns_records/$RID" -d "{\"content\":\"$IP\"}" | python3 -c 'import sys,json; d=json.load(sys.stdin); print("updated", d["success"], d["result"]["content"])'
|
||||
@@ -0,0 +1,7 @@
|
||||
[Unit]
|
||||
Description=headscale DDNS every 10 minutes
|
||||
[Timer]
|
||||
OnBootSec=2min
|
||||
OnUnitActiveSec=10min
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
Reference in New Issue
Block a user