The failed-START notifier built earlier today had its first REAL firing at
15:28: headscale-ddns.service exited 1 after succeeding all afternoon. The
detection worked. The alarm was also useless, and that is the finding.
Both failure paths exited 1 IN SILENCE, so the message said "exit status 1" and
nothing else. An alarm you cannot act on costs the same triage as no alarm at
all -- the notifier did its job and the subject script had no diagnostics for it
to carry.
Cause was transient and harmless: icanhazip.com did not answer inside its 10s
cap, so the IP came back empty and the regex guard refused it. No DNS impact --
the record already held the right address, verified against 1.1.1.1 before
touching anything, and the next timer run succeeded. Arithmetic confirms it:
~17s vault read + 10s curl timeout = 27s against the 28s the failing run took.
Fixed, both verified by making them fail:
- every exit path names its cause; a missing vault key names the key, an
EMPTY token is distinguished from a failed read, and a dead WAN lookup adds
"DNS left unchanged" because that is the fact the reader needs
- the WAN lookup retries 3x with ANNOUNCED attempts -- one third-party blip
should not page a human, and a silent retry would hide a degrading
dependency
⚠ ALSO: this script was not tracked anywhere. A fix to the thing every mesh
client resolves through lived on exactly one disk. Script, unit and timer are
in the repo now.
Measured and recorded: the vault read is 17 of the script's 18 seconds, every
10 minutes. Not a fault, but it bounds any retry budget and it is fleet-wide --
svos-dev's alarm unit carries the same 17-second note.
50 lines
2.3 KiB
Markdown
50 lines
2.3 KiB
Markdown
# headscale-ddns — keep `headscale.phasefinal.com` pointed at the NH3 WAN v4
|
||
|
||
A user timer on nh3-dev, every 10 minutes: read the WAN v4, PATCH the Cloudflare
|
||
A record if it moved. It is how every mesh client finds the control plane, so a
|
||
silent failure here eventually costs the mesh.
|
||
|
||
Installed at `~/.local/bin/headscale-ddns.sh` with user units
|
||
`headscale-ddns.{service,timer}`. **Tracked here since 2026-09-22** — it was
|
||
running untracked before that, so a fix to it lived on exactly one disk.
|
||
|
||
## The 2026-09-22 failure, and what it exposed
|
||
|
||
The unit failed at 15:28 (exit 1) after succeeding all afternoon. Cause was
|
||
transient — `icanhazip.com` did not answer inside its 10s cap, so `$IP` came back
|
||
empty and the regex guard refused it. No DNS impact: the record already held the
|
||
right address and the next timer run succeeded.
|
||
|
||
**What actually mattered was that the alarm carried no cause.** Both failure
|
||
paths were `|| exit 1` in silence, so the failed-START notifier fired correctly
|
||
and said only "exit status 1". An alarm you cannot act on costs the same triage
|
||
as no alarm.
|
||
|
||
Two fixes, both verified by making them fail:
|
||
|
||
- **Every exit path now says why** — a missing vault key names the key; an empty
|
||
token is distinguished from a failed read; a dead WAN lookup says so and adds
|
||
*"DNS left unchanged"*, which is the fact the reader needs.
|
||
- **The WAN lookup retries 3×** with announced attempts. One third-party blip
|
||
should not page a human, and a *silent* retry would hide a degrading
|
||
dependency.
|
||
|
||
## ⚠ The vault read is 17 of the script's 18 seconds
|
||
|
||
Measured 2026-09-22: `secret get` takes **~17s**, everything else ~1s. It runs
|
||
every 10 minutes. That is not a fault — but it bounds any retry budget here, and
|
||
it is a fleet-wide cost worth knowing: svos-dev's own alarm unit carries the same
|
||
17-second note. Any script in a timer that reads the vault pays it.
|
||
|
||
## Triage
|
||
|
||
```sh
|
||
systemctl --user status headscale-ddns.service
|
||
journalctl --user -u headscale-ddns.service -n 50 --no-pager
|
||
dig +short headscale.phasefinal.com @1.1.1.1 # the thing that actually matters
|
||
/home/lkraven/.local/bin/headscale-ddns.sh # safe to run by hand; idempotent
|
||
```
|
||
|
||
A failed unit clears itself on the next timer run; it does not need
|
||
`reset-failed` unless you want the state gone immediately.
|