An infrastructure day with no training work, and the through-line is that every fault was invisible to monitoring. NH3↔Anaheim had been crossing a throttled DERP relay rather than a direct path for long enough to carry 78 GB; `.internal` DNS was failing roughly one lookup in ten from two independent causes; SearXNG had exactly one working general web engine. Nothing alarmed on any of it. All three surfaced because tts-dev had a 1545 ms voice-loop budget and chose to measure rather than adapt around the problem. Also landed: althing v3.6.3, which makes hyphenated search work for the first time on a fleet whose hostnames are nearly all hyphenated; FleetTools, a capability index autoloaded by Claude, Codex and Grok from one symlinked file; Miranda's Hermes plugin moved from a copy to a repo symlink; Worldtree's env.sh secrets vaulted. Six detail files. The in-flight section is rewritten and shrinks 142 lines to 64 — it opens on the one thing this session did NOT verify, lv-mccarthy's run outcome, which was left untouched and must not be assumed good. Two foot-guns recorded, both mine: the ESH egress experiment reverted on a diagnosis the rollback itself falsified, and `!ENV` in searxng settings, which has no constructor in that build and crash-looped the container ten times. No archival this run. 165 of 169 dated entries are under the 14-day guard and the remaining four all carry open deferred pointers, so the index stays over the soft cap at 480 lines — an over-cap file that keeps live decisions beats a scannable one that lost one.
95 lines
4.2 KiB
Markdown
95 lines
4.2 KiB
Markdown
# [2026-09-18] `.internal` DNS was failing ~10% of lookups — two causes, both fleet-wide
|
|
|
|
Two independent faults, fixed in order. Together they were costing roughly one
|
|
`.internal` lookup in ten either a hard failure or a five-second stall, on every
|
|
DHCP client at every site.
|
|
|
|
## Fault 1 — a PUBLIC resolver was the fallback for a PRIVATE zone
|
|
|
|
DHCP handed out `10.100.50.40, 1.1.1.1` on both NH3 VLANs. When the internal
|
|
resolver missed a packet, the resolver waited out its timeout, fell through to
|
|
Cloudflare, and got an **authoritative NXDOMAIN** — so a transient miss became a
|
|
definitive "no such host" rather than a retry. Binary failure: instant, or five
|
|
seconds then `gaierror`.
|
|
|
|
**Anaheim was worse.** Its FortiGate handed clients *itself* as resolver and
|
|
forwarded to `1.1.1.1`/`1.0.0.1`, so **no ANA host could resolve `.internal` at
|
|
all** — `ana-docker`, which HOSTS the Anaheim AdGuard, had `1.1.1.1` in its own
|
|
`resolv.conf`.
|
|
|
|
⚠ **`/etc/resolv.conf` on these hosts is DHCP-managed** (dhclient on ens18). Hand
|
|
editing it survives exactly until the next lease renewal — tts-dev's original
|
|
framing was "fix nh3-dev's resolv.conf", which would have worked, been verified,
|
|
and then silently reverted. Change it at the source:
|
|
|
|
- **UniFi**: `dhcpd_dns_1` / `dhcpd_dns_2` per network object.
|
|
- **FortiGate**: `config system dhcp server`, `set dns-service specify` +
|
|
`dns-server1` / `dns-server2`.
|
|
- Then `sudo dhclient -1 -v ens18` to pick it up **without releasing** the lease.
|
|
|
|
**Operator's ruling was broader than the proposal**: each site's backup resolver
|
|
should be *another site's* resolver. ESH already worked this way; the other two
|
|
were set, and the ring closes.
|
|
|
|
ESH primary 10.0.50.45 backup 10.100.50.40 (NH3) [pre-existing]
|
|
NH3 primary 10.100.50.40 backup 10.250.50.70 (ANA) [changed]
|
|
ANA primary 10.250.50.70 backup 10.0.50.45 (ESH) [changed]
|
|
|
|
All six members verified to answer the private zone. ⚠ The SureFire tenant DHCP
|
|
server (`sfsrv`) and the fortilink switch-management server were deliberately left
|
|
alone — tenant property.
|
|
|
|
**Result: hard failures 3/40 → 0/40. The 5 s stalls remained.**
|
|
|
|
## Fault 2 — THE ROOT CAUSE: AdGuard rate-limiting, identical on all three
|
|
|
|
ratelimit: 20
|
|
ratelimit_subnet_len_ipv4: 24
|
|
ratelimit_whitelist: []
|
|
|
|
**20 queries per second shared across an entire /24 — per subnet, not per client.**
|
|
Every host on a VLAN draws from one bucket, so one busy container starves the rest,
|
|
and anything over the line is **silently dropped**, costing the client its full 5 s
|
|
resolver timeout.
|
|
|
|
Demonstrated rather than inferred — 60 concurrent queries at one resolver:
|
|
|
|
| | before | after |
|
|
|---|---|---|
|
|
| median | 5044 ms | 72 ms |
|
|
| timed out (>4 s) | **40 of 60** | **0 of 60** |
|
|
|
|
Set to `ratelimit: 0` on all three. These are LAN-only resolvers behind three
|
|
firewalls; that setting exists to blunt internet-side DNS amplification, which is
|
|
not this.
|
|
|
|
⚠ **Restarting AdGuard is required for a config change and briefly drops that site's
|
|
primary resolver — do the three ONE AT A TIME** so the cross-site ring always has a
|
|
live member. That the ring made this rollout safe is the point of having built it
|
|
first, an hour earlier.
|
|
|
|
## End-to-end, the original symptom
|
|
|
|
40 TCP connects to a service by name vs by IP; the by-IP arm is the control.
|
|
|
|
| | original | after ring only | after ratelimit 0 |
|
|
|---|---|---|---|
|
|
| by NAME p90 | 120 ms | 5017 ms | **13.5 ms** |
|
|
| by NAME >1 s | 3/40 | 5/40 | **0/40** |
|
|
| by NAME failed | 2/40 | 0/40 | **0/40** |
|
|
| by IP p90 (control) | 25 ms | — | 11.7 ms |
|
|
|
|
tts-dev confirmed from talk's own request path: name and IP **1.5 ms apart at the
|
|
median**. Resolution has stopped being a cost rather than become a smaller one.
|
|
|
|
## What this says about monitoring
|
|
|
|
Nothing on this fleet watches DNS success rate. The fleet's actual DNS monitor
|
|
today was a voice loop with a 1545 ms budget whose owner chose to measure instead
|
|
of adapting around the problem. Both faults were invisible to Beszel and Uptime
|
|
Kuma — every host was up, every service answered.
|
|
|
|
Related: [[2026-09-18-nh3-ana-derp-relay]] — the ANA AdGuard is only a sensible
|
|
cross-site backup because that fix took NH3→ANA from ~1.2 s to ~15 ms the same
|
|
afternoon. The two changes compose.
|