Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md
T
vh 148a5a34da memory: snapshot — three silent fleet faults found and fixed in one afternoon
An infrastructure day with no training work, and the through-line is that
every fault was invisible to monitoring. NH3↔Anaheim had been crossing a
throttled DERP relay rather than a direct path for long enough to carry 78 GB;
`.internal` DNS was failing roughly one lookup in ten from two independent
causes; SearXNG had exactly one working general web engine. Nothing alarmed on
any of it. All three surfaced because tts-dev had a 1545 ms voice-loop budget
and chose to measure rather than adapt around the problem.

Also landed: althing v3.6.3, which makes hyphenated search work for the first
time on a fleet whose hostnames are nearly all hyphenated; FleetTools, a
capability index autoloaded by Claude, Codex and Grok from one symlinked file;
Miranda's Hermes plugin moved from a copy to a repo symlink; Worldtree's
env.sh secrets vaulted.

Six detail files. The in-flight section is rewritten and shrinks 142 lines to
64 — it opens on the one thing this session did NOT verify, lv-mccarthy's run
outcome, which was left untouched and must not be assumed good.

Two foot-guns recorded, both mine: the ESH egress experiment reverted on a
diagnosis the rollback itself falsified, and `!ENV` in searxng settings, which
has no constructor in that build and crash-looped the container ten times.

No archival this run. 165 of 169 dated entries are under the 14-day guard and
the remaining four all carry open deferred pointers, so the index stays over
the soft cap at 480 lines — an over-cap file that keeps live decisions beats a
scannable one that lost one.
2026-09-18 20:09:01 -07:00

4.2 KiB

[2026-09-18] .internal DNS was failing ~10% of lookups — two causes, both fleet-wide

Two independent faults, fixed in order. Together they were costing roughly one .internal lookup in ten either a hard failure or a five-second stall, on every DHCP client at every site.

Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone

DHCP handed out 10.100.50.40, 1.1.1.1 on both NH3 VLANs. When the internal resolver missed a packet, the resolver waited out its timeout, fell through to Cloudflare, and got an authoritative NXDOMAIN — so a transient miss became a definitive "no such host" rather than a retry. Binary failure: instant, or five seconds then gaierror.

Anaheim was worse. Its FortiGate handed clients itself as resolver and forwarded to 1.1.1.1/1.0.0.1, so no ANA host could resolve .internal at all — ana-docker, which HOSTS the Anaheim AdGuard, had 1.1.1.1 in its own resolv.conf.

⚠ /etc/resolv.conf on these hosts is DHCP-managed (dhclient on ens18). Hand editing it survives exactly until the next lease renewal — tts-dev's original framing was "fix nh3-dev's resolv.conf", which would have worked, been verified, and then silently reverted. Change it at the source:

  • UniFi: dhcpd_dns_1 / dhcpd_dns_2 per network object.
  • FortiGate: config system dhcp server, set dns-service specify + dns-server1 / dns-server2.
  • Then sudo dhclient -1 -v ens18 to pick it up without releasing the lease.

Operator's ruling was broader than the proposal: each site's backup resolver should be another site's resolver. ESH already worked this way; the other two were set, and the ring closes.

ESH  primary 10.0.50.45    backup 10.100.50.40  (NH3)   [pre-existing]
NH3  primary 10.100.50.40  backup 10.250.50.70  (ANA)   [changed]
ANA  primary 10.250.50.70  backup 10.0.50.45    (ESH)   [changed]

All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP server (sfsrv) and the fortilink switch-management server were deliberately left alone — tenant property.

Result: hard failures 3/40 → 0/40. The 5 s stalls remained.

Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three

ratelimit: 20
ratelimit_subnet_len_ipv4: 24
ratelimit_whitelist: []

20 queries per second shared across an entire /24 — per subnet, not per client. Every host on a VLAN draws from one bucket, so one busy container starves the rest, and anything over the line is silently dropped, costing the client its full 5 s resolver timeout.

Demonstrated rather than inferred — 60 concurrent queries at one resolver:

before after
median 5044 ms 72 ms
timed out (>4 s) 40 of 60 0 of 60

Set to ratelimit: 0 on all three. These are LAN-only resolvers behind three firewalls; that setting exists to blunt internet-side DNS amplification, which is not this.

⚠ Restarting AdGuard is required for a config change and briefly drops that site's primary resolver — do the three ONE AT A TIME so the cross-site ring always has a live member. That the ring made this rollout safe is the point of having built it first, an hour earlier.

End-to-end, the original symptom

40 TCP connects to a service by name vs by IP; the by-IP arm is the control.

original after ring only after ratelimit 0
by NAME p90 120 ms 5017 ms 13.5 ms
by NAME >1 s 3/40 5/40 0/40
by NAME failed 2/40 0/40 0/40
by IP p90 (control) 25 ms — 11.7 ms

tts-dev confirmed from talk's own request path: name and IP 1.5 ms apart at the median. Resolution has stopped being a cost rather than become a smaller one.

What this says about monitoring

Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor today was a voice loop with a 1545 ms budget whose owner chose to measure instead of adapting around the problem. Both faults were invisible to Beszel and Uptime Kuma — every host was up, every service answered.

Related: 2026-09-18-nh3-ana-derp-relay — the ANA AdGuard is only a sensible cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same afternoon. The two changes compose.