feat(dns): fleet .internal naming — git-sourced, agent-managed, three resolvers
Names for fleet hosts so addresses stop needing to be memorised. Built because IPv6 makes that hopeless — and, more to the point, because v6 addresses are derived rather than assigned, so they cannot reliably be written down once and trusted either. dns/internal.yaml source of truth: 38 hosts + 4 service aliases scripts/dns-sync.py reconciles AdGuard resolvers against it stacks/adguard-ana/ the colo's resolver, which did not exist Naming is <host>.<site>.internal with sites ana/esh/nh3 (operator's call). .internal is ICANN-reserved for this; .local is reserved for mDNS, which is why searxng.pfi.local was a collision that merely happened to work. Same posture as deploy-stack.sh: file is intent, resolvers are derived state, you see a diff before anything changes. Every name is published to every resolver, so the site label says where a host IS, not who knows about it. Two properties that matter: - Authority is scoped to the ZONE, not the resolver. ESH carries hand-made esteban.net rewrites predating this; they are read, ignored and preserved. Resolver-wide authority would have silently deleted them. - Within .internal it IS authoritative, so UI-added names get removed. That is the point — one place to look. Colo gap closed: ana-docker had no resolver at all (hosts went straight to 1.1.1.1). Its AdGuard runs API on 8053 because 8080/3000 were taken, so the port is carried per-site in the yaml rather than assumed by the script. It ships with no blocklists — a false positive on a server network breaks service-to-service calls for no upside. Auth is a dedicated infra-ops AdGuard user, not the operator's account, password vaulted at nh3-dev/adguard-infra-ops-password. Pre-change configs backed up on each resolver. Both resolvers stayed answering across the restart. searxng.pfi.local -> searxng.ana.internal, with the old Host() kept alongside so nothing breaks mid-migration. matrix.pfi.local deliberately NOT migrated: a Matrix server_name is baked into every user id, room id and signing key, so renaming it rebuilds the homeserver's identity rather than changing a DNS name. The v6 column is empty and correct — no fleet host has a global v6 address yet. The file documents why addresses must be pinned statically before they go in, since a record that silently stops matching is worse than no record.
This commit is contained in:
+125
@@ -0,0 +1,125 @@
|
||||
# Fleet internal DNS — `*.internal`
|
||||
|
||||
Names for fleet hosts so nobody has to remember addresses. Built 2026-08-19
|
||||
because IPv6 makes memorising them hopeless — and, more to the point, because
|
||||
v6 addresses are *derived* rather than assigned, so they cannot be reliably
|
||||
memorised **or** written down once and trusted.
|
||||
|
||||
```
|
||||
dns/internal.yaml the source of truth — hosts, sites, aliases
|
||||
scripts/dns-sync.py reconciles the resolvers against it
|
||||
```
|
||||
|
||||
## Adding a name
|
||||
|
||||
Edit `dns/internal.yaml`, then:
|
||||
|
||||
```bash
|
||||
scripts/dns-sync.py --dry-run # see the diff
|
||||
scripts/dns-sync.py # apply, with a prompt
|
||||
```
|
||||
|
||||
That is the whole workflow. It is deliberately the same shape as
|
||||
`deploy-stack.sh`: a file in git is the intent, the running system is derived
|
||||
state, and you see a diff before anything changes.
|
||||
|
||||
## Naming
|
||||
|
||||
`<host>.<site>.internal`, sites **`ana`** (Anaheim colo), **`esh`** (home lab),
|
||||
**`nh3`** (office).
|
||||
|
||||
`.internal` is ICANN-reserved for private use, which is why it is used here
|
||||
rather than `.local` (reserved for mDNS — the old `searxng.pfi.local` was a
|
||||
standards collision that happened to work) or an invented TLD that could later
|
||||
collide with a real one.
|
||||
|
||||
**Every name is published to every resolver.** The site label says where a host
|
||||
*is*, not which resolver knows about it — `ana-docker.ana.internal` resolves
|
||||
from ESH and NH3 too.
|
||||
|
||||
Irvine is not a fourth zone: `irv-ml1` is reachable only through NH3's
|
||||
WireGuard tunnel and numbered out of NH3's `10.100.79.0/24`, so it lives under
|
||||
`nh3`. Worth revisiting if Irvine ever becomes a site in its own right.
|
||||
|
||||
## The resolvers
|
||||
|
||||
| site | resolver | API port |
|
||||
|---|---|---|
|
||||
| ana | ana-docker `10.250.50.70` | **8053** |
|
||||
| esh | esh-docker-vm `10.0.50.45` | 8080 |
|
||||
| nh3 | nh3-docker `10.100.50.40` | 8080 |
|
||||
|
||||
ana is the odd one out — `:8080` and `:3000` were already taken on that host —
|
||||
so the port is carried per-site in `internal.yaml` rather than assumed by the
|
||||
script.
|
||||
|
||||
The colo resolver (`stacks/adguard-ana/`) was stood up as part of this work;
|
||||
before it, colo hosts resolved straight against `1.1.1.1` and the site had no
|
||||
way to answer for internal names. ESH and NH3 run older, unmanaged compose
|
||||
files, left alone on purpose — adopting three live resolvers into this repo
|
||||
while also introducing a new naming system is two risky changes at once.
|
||||
|
||||
## Two properties worth not breaking
|
||||
|
||||
**Authority is scoped to the zone, not the resolver.** Only rewrites ending in
|
||||
`.internal` are managed. The ESH resolver carries hand-made `esteban.net`
|
||||
entries that predate this system; the sync reads them, ignores them, and leaves
|
||||
them alone. If this ever grows to manage another zone, that scoping is the
|
||||
thing to be careful with — resolver-wide authority would silently delete
|
||||
somebody else's work.
|
||||
|
||||
**Within the zone it is authoritative.** Names added by hand in the AdGuard UI
|
||||
*will* be deleted by the next sync. That is the point: one place to look.
|
||||
|
||||
## Credential
|
||||
|
||||
`scripts/dns-sync.py` authenticates as a dedicated **`infra-ops`** AdGuard user,
|
||||
not as the operator's account, and pulls the password from the vault:
|
||||
|
||||
```bash
|
||||
secret get nh3-dev/adguard-infra-ops-password
|
||||
```
|
||||
|
||||
⚠️ The vault appends a trailing newline on read. The script strips it, because
|
||||
a password carrying a stray `\n` fails auth in a way that looks exactly like a
|
||||
wrong password.
|
||||
|
||||
The existing `lkraven` AdGuard user was left untouched. Config backups from
|
||||
before the user was added are on each resolver as
|
||||
`AdGuardHome.yaml.bak-preinfraops-*`.
|
||||
|
||||
## ⚠️ IPv6 — the reason this exists, and still the unfinished half
|
||||
|
||||
The `v6:` column is empty and that is correct as of 2026-08-19: **no fleet host
|
||||
has a global IPv6 address yet.** ESH's `/56` is live only on `esh-cameras`,
|
||||
NH3's LANs are back to `ipv6_interface_type: none`, the colo has no v6 at all.
|
||||
|
||||
When v6 arrives, **do not paste in whatever `ip -6 addr` shows.** SLAAC gives
|
||||
hosts either EUI-64 addresses (MAC-coupled) or privacy-extension ones (which
|
||||
rotate), and UniFi has no v6 equivalent of a DHCP reservation. An address only
|
||||
belongs in this file once it has been pinned **statically on the host itself**.
|
||||
A record that silently stops matching reality is worse than no record — the
|
||||
name keeps resolving and starts lying.
|
||||
|
||||
The suggested convention when that happens: give each server a static address
|
||||
out of its site's `/64` whose low-order bits echo the v4 host octet
|
||||
(`esh-docker-vm` at `…::45`), so the addresses are both declarable and
|
||||
semi-memorable.
|
||||
|
||||
## Not migrated: `matrix.pfi.local`
|
||||
|
||||
`searxng.pfi.local` moved to `searxng.ana.internal` (both names still route,
|
||||
so nothing breaks mid-migration; drop the fallback `Host()` in
|
||||
`stacks/searxng/compose.yaml` once the Traefik log shows the old one unused).
|
||||
|
||||
**`matrix.pfi.local` was deliberately left alone.** A Matrix `server_name` is
|
||||
baked into every user ID, room ID and signing key, and federation identity is
|
||||
derived from it — renaming it is not a DNS change, it is rebuilding the
|
||||
homeserver's identity and invalidating its history. It stays on `.local`.
|
||||
|
||||
## Still open
|
||||
|
||||
Colo hosts still point at `1.1.1.1`, so they do not yet *use* the new resolver
|
||||
— they only get answers if something asks it directly. Repointing a whole
|
||||
site's DNS is a bigger change than standing the service up, so it is a separate
|
||||
operator-approved step.
|
||||
Reference in New Issue
Block a user