Files
esh-pfi-infrastructure/services/headscale-ddns/README.md
T
vh 30517fd603 fix(headscale-ddns): say why it failed, retry the WAN lookup, and track it at all
The failed-START notifier built earlier today had its first REAL firing at
15:28: headscale-ddns.service exited 1 after succeeding all afternoon. The
detection worked. The alarm was also useless, and that is the finding.

Both failure paths exited 1 IN SILENCE, so the message said "exit status 1" and
nothing else. An alarm you cannot act on costs the same triage as no alarm at
all -- the notifier did its job and the subject script had no diagnostics for it
to carry.

Cause was transient and harmless: icanhazip.com did not answer inside its 10s
cap, so the IP came back empty and the regex guard refused it. No DNS impact --
the record already held the right address, verified against 1.1.1.1 before
touching anything, and the next timer run succeeded. Arithmetic confirms it:
~17s vault read + 10s curl timeout = 27s against the 28s the failing run took.

Fixed, both verified by making them fail:
  - every exit path names its cause; a missing vault key names the key, an
    EMPTY token is distinguished from a failed read, and a dead WAN lookup adds
    "DNS left unchanged" because that is the fact the reader needs
  - the WAN lookup retries 3x with ANNOUNCED attempts -- one third-party blip
    should not page a human, and a silent retry would hide a degrading
    dependency

⚠ ALSO: this script was not tracked anywhere. A fix to the thing every mesh
client resolves through lived on exactly one disk. Script, unit and timer are
in the repo now.

Measured and recorded: the vault read is 17 of the script's 18 seconds, every
10 minutes. Not a fault, but it bounds any retry budget and it is fleet-wide --
svos-dev's alarm unit carries the same 17-second note.
2026-09-22 15:32:26 -07:00

50 lines
2.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# headscale-ddns — keep `headscale.phasefinal.com` pointed at the NH3 WAN v4
A user timer on nh3-dev, every 10 minutes: read the WAN v4, PATCH the Cloudflare
A record if it moved. It is how every mesh client finds the control plane, so a
silent failure here eventually costs the mesh.
Installed at `~/.local/bin/headscale-ddns.sh` with user units
`headscale-ddns.{service,timer}`. **Tracked here since 2026-09-22** — it was
running untracked before that, so a fix to it lived on exactly one disk.
## The 2026-09-22 failure, and what it exposed
The unit failed at 15:28 (exit 1) after succeeding all afternoon. Cause was
transient — `icanhazip.com` did not answer inside its 10s cap, so `$IP` came back
empty and the regex guard refused it. No DNS impact: the record already held the
right address and the next timer run succeeded.
**What actually mattered was that the alarm carried no cause.** Both failure
paths were `|| exit 1` in silence, so the failed-START notifier fired correctly
and said only "exit status 1". An alarm you cannot act on costs the same triage
as no alarm.
Two fixes, both verified by making them fail:
- **Every exit path now says why** — a missing vault key names the key; an empty
token is distinguished from a failed read; a dead WAN lookup says so and adds
*"DNS left unchanged"*, which is the fact the reader needs.
- **The WAN lookup retries 3×** with announced attempts. One third-party blip
should not page a human, and a *silent* retry would hide a degrading
dependency.
## ⚠ The vault read is 17 of the script's 18 seconds
Measured 2026-09-22: `secret get` takes **~17s**, everything else ~1s. It runs
every 10 minutes. That is not a fault — but it bounds any retry budget here, and
it is a fleet-wide cost worth knowing: svos-dev's own alarm unit carries the same
17-second note. Any script in a timer that reads the vault pays it.
## Triage
```sh
systemctl --user status headscale-ddns.service
journalctl --user -u headscale-ddns.service -n 50 --no-pager
dig +short headscale.phasefinal.com @1.1.1.1 # the thing that actually matters
/home/lkraven/.local/bin/headscale-ddns.sh # safe to run by hand; idempotent
```
A failed unit clears itself on the next timer run; it does not need
`reset-failed` unless you want the state gone immediately.