Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md
T
vh 148a5a34da memory: snapshot — three silent fleet faults found and fixed in one afternoon
An infrastructure day with no training work, and the through-line is that
every fault was invisible to monitoring. NH3↔Anaheim had been crossing a
throttled DERP relay rather than a direct path for long enough to carry 78 GB;
`.internal` DNS was failing roughly one lookup in ten from two independent
causes; SearXNG had exactly one working general web engine. Nothing alarmed on
any of it. All three surfaced because tts-dev had a 1545 ms voice-loop budget
and chose to measure rather than adapt around the problem.

Also landed: althing v3.6.3, which makes hyphenated search work for the first
time on a fleet whose hostnames are nearly all hyphenated; FleetTools, a
capability index autoloaded by Claude, Codex and Grok from one symlinked file;
Miranda's Hermes plugin moved from a copy to a repo symlink; Worldtree's
env.sh secrets vaulted.

Six detail files. The in-flight section is rewritten and shrinks 142 lines to
64 — it opens on the one thing this session did NOT verify, lv-mccarthy's run
outcome, which was left untouched and must not be assumed good.

Two foot-guns recorded, both mine: the ESH egress experiment reverted on a
diagnosis the rollback itself falsified, and `!ENV` in searxng settings, which
has no constructor in that build and crash-looped the container ten times.

No archival this run. 165 of 169 dated entries are under the 14-day guard and
the remaining four all carry open deferred pointers, so the index stays over
the soft cap at 480 lines — an over-cap file that keeps live decisions beats a
scannable one that lost one.
2026-09-18 20:09:01 -07:00

95 lines
4.2 KiB
Markdown

# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
Two independent faults, fixed in order. Together they were costing roughly one
`.internal` lookup in ten either a hard failure or a five-second stall, on every
DHCP client at every site.
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
resolver missed a packet, the resolver waited out its timeout, fell through to
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
definitive "no such host" rather than a retry. Binary failure: instant, or five
seconds then `gaierror`.
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
`resolv.conf`.
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
editing it survives exactly until the next lease renewal — tts-dev's original
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
and then silently reverted. Change it at the source:
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
`dns-server1` / `dns-server2`.
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
**Operator's ruling was broader than the proposal**: each site's backup resolver
should be *another site's* resolver. ESH already worked this way; the other two
were set, and the ring closes.
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
server (`sfsrv`) and the fortilink switch-management server were deliberately left
alone — tenant property.
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
ratelimit: 20
ratelimit_subnet_len_ipv4: 24
ratelimit_whitelist: []
**20 queries per second shared across an entire /24 — per subnet, not per client.**
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
and anything over the line is **silently dropped**, costing the client its full 5 s
resolver timeout.
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
| | before | after |
|---|---|---|
| median | 5044 ms | 72 ms |
| timed out (>4 s) | **40 of 60** | **0 of 60** |
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
firewalls; that setting exists to blunt internet-side DNS amplification, which is
not this.
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
live member. That the ring made this rollout safe is the point of having built it
first, an hour earlier.
## End-to-end, the original symptom
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
| | original | after ring only | after ratelimit 0 |
|---|---|---|---|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
| by NAME failed | 2/40 | 0/40 | **0/40** |
| by IP p90 (control) | 25 ms | — | 11.7 ms |
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
median**. Resolution has stopped being a cost rather than become a smaller one.
## What this says about monitoring
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
of adapting around the problem. Both faults were invisible to Beszel and Uptime
Kuma — every host was up, every service answered.
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
afternoon. The two changes compose.