Files
esh-pfi-infrastructure/dns
vh efddb4e511 feat(scriberr): stand up transcription on ana-ml2, pinned to GPU1
Scriberr transcribes audio and video locally with WhisperX and
speaker diarization, and it lands on ana-ml2 rather than ana-docker
because the work is GPU-shaped: ana-docker offers eight cores already
shared with fifty containers and thirty-seven gigabytes of disk,
against ninety-six cores, terabytes on /tank and idle capacity on
GPU1. The reservation names device 1 explicitly, since GPU0 is fully
committed to the gen seat, and the container is confirmed to see that
card alone.

The image is built from source, which is not a preference. These are
Blackwell cards at sm_120; the published CUDA image covers Pascal
through Ada only, and the blackwell image the upstream README
documents has never been published at all. The path upstream actually
ships for sm_120 is Dockerfile.cuda.12.9, carrying CUDA 12.9 and cu128
torch, so that is what gets built. The compose header says so, because
the obvious cleanup is to swap in the published image and that would
silently drop the deployment to CPU.

Two configuration details are load-bearing and documented where
someone would go to change them. The application runs as uid 10001
rather than the usual 1000: that Dockerfile moves its user aside for
Ubuntu 24.04's own uid-1000 account and chowns /app accordingly, while
the entrypoint's remapping covers only the data directories, so at
1000 the process cannot open its database and restarts forever behind
a SQLite error that reads as though the machine were out of memory.
Secure cookies stay off while the service is reached over plain HTTP,
or sessions are dropped by the browser and login appears to loop for
no visible reason.

Storage is bind-mounted onto /tank because model weights run to
several gigabytes and the root pool on that host is nearly full.

Also adds the scriberr service alias to internal DNS, following the
existing alias convention so consumers name the service rather than
the box.
2026-08-23 19:28:33 -07:00
..

Fleet internal DNS — *.internal

Names for fleet hosts so nobody has to remember addresses. Built 2026-08-19 because IPv6 makes memorising them hopeless — and, more to the point, because v6 addresses are derived rather than assigned, so they cannot be reliably memorised or written down once and trusted.

dns/internal.yaml     the source of truth — hosts, sites, aliases
scripts/dns-sync.py   reconciles the resolvers against it

Adding a name

Edit dns/internal.yaml, then:

scripts/dns-sync.py --dry-run     # see the diff
scripts/dns-sync.py               # apply, with a prompt

That is the whole workflow. It is deliberately the same shape as deploy-stack.sh: a file in git is the intent, the running system is derived state, and you see a diff before anything changes.

Naming

<host>.<site>.internal, sites ana (Anaheim colo), esh (home lab), nh3 (office).

.internal is ICANN-reserved for private use, which is why it is used here rather than .local (reserved for mDNS — the old searxng.pfi.local was a standards collision that happened to work) or an invented TLD that could later collide with a real one.

Every name is published to every resolver. The site label says where a host is, not which resolver knows about it — ana-docker.ana.internal resolves from ESH and NH3 too.

Irvine is not a fourth zone: irv-ml1 is reachable only through NH3's WireGuard tunnel and numbered out of NH3's 10.100.79.0/24, so it lives under nh3. Worth revisiting if Irvine ever becomes a site in its own right.

The resolvers

site resolver API port
ana ana-docker 10.250.50.70 8053
esh esh-docker-vm 10.0.50.45 8080
nh3 nh3-docker 10.100.50.40 8080

ana is the odd one out — :8080 and :3000 were already taken on that host — so the port is carried per-site in internal.yaml rather than assumed by the script.

The colo resolver (stacks/adguard-ana/) was stood up as part of this work; before it, colo hosts resolved straight against 1.1.1.1 and the site had no way to answer for internal names. ESH and NH3 run older, unmanaged compose files, left alone on purpose — adopting three live resolvers into this repo while also introducing a new naming system is two risky changes at once.

Two properties worth not breaking

Authority is scoped to the zone, not the resolver. Only rewrites ending in .internal are managed. The ESH resolver carries hand-made esteban.net entries that predate this system; the sync reads them, ignores them, and leaves them alone. If this ever grows to manage another zone, that scoping is the thing to be careful with — resolver-wide authority would silently delete somebody else's work.

Within the zone it is authoritative. Names added by hand in the AdGuard UI will be deleted by the next sync. That is the point: one place to look.

Credential

scripts/dns-sync.py authenticates as a dedicated infra-ops AdGuard user, not as the operator's account, and pulls the password from the vault:

secret get nh3-dev/adguard-infra-ops-password

⚠️ The vault appends a trailing newline on read. The script strips it, because a password carrying a stray \n fails auth in a way that looks exactly like a wrong password.

The existing lkraven AdGuard user was left untouched. Config backups from before the user was added are on each resolver as AdGuardHome.yaml.bak-preinfraops-*.

⚠️ IPv6 — the reason this exists, and still the unfinished half

The v6: column is empty and that is correct as of 2026-08-19: no fleet host has a global IPv6 address yet. ESH's /56 is live only on esh-cameras, NH3's LANs are back to ipv6_interface_type: none, the colo has no v6 at all.

When v6 arrives, do not paste in whatever ip -6 addr shows. SLAAC gives hosts either EUI-64 addresses (MAC-coupled) or privacy-extension ones (which rotate), and UniFi has no v6 equivalent of a DHCP reservation. An address only belongs in this file once it has been pinned statically on the host itself. A record that silently stops matching reality is worse than no record — the name keeps resolving and starts lying.

The suggested convention when that happens: give each server a static address out of its site's /64 whose low-order bits echo the v4 host octet (esh-docker-vm at …::45), so the addresses are both declarable and semi-memorable.

Not migrated: matrix.pfi.local

searxng.pfi.local moved to searxng.ana.internal (both names still route, so nothing breaks mid-migration; drop the fallback Host() in stacks/searxng/compose.yaml once the Traefik log shows the old one unused).

matrix.pfi.local was deliberately left alone. A Matrix server_name is baked into every user ID, room ID and signing key, and federation identity is derived from it — renaming it is not a DNS change, it is rebuilding the homeserver's identity and invalidating its history. It stays on .local.

Still open

Colo hosts still point at 1.1.1.1, so they do not yet use the new resolver — they only get answers if something asks it directly. Repointing a whole site's DNS is a bigger change than standing the service up, so it is a separate operator-approved step.