Lessons from DNS caching and NXDOMAIN
“It’s always DNS” is a joke because it’s so often true. One of the more confusing versions of the problem is when you’ve fixed a DNS record, you can see the correct answer from one machine, and another machine still insists the name doesn’t exist. The usual culprit is negative caching: resolvers remembering that a name returned NXDOMAIN.
What NXDOMAIN means
When a resolver asks for a name and the authoritative server says “this name does not exist,” the response code is NXDOMAIN. That’s different from a name that exists but has no record of the type you asked for. That case is NOERROR with an empty answer, sometimes called NODATA. Both are cached, and both cause the same kind of confusion.
Negative answers get cached too
Most people know positive answers are cached for the record’s TTL. Fewer realize that resolvers also cache “doesn’t exist.” Per RFC 2308, the length of that negative cache comes from the zone’s SOA record. It’s the lower of the SOA record’s own TTL and its minimum field, the last number in the SOA.
So the classic sequence goes like this:
- You test
newapp.example.combefore creating the record. It’s NXDOMAIN. - Your resolver caches that negative answer for, say, an hour.
- You create the record a minute later.
- For the next hour, that resolver keeps telling everyone the name doesn’t exist.
The record was fine the whole time. You just asked too early.
Find out how long you’ll wait
Look at the zone’s SOA:
dig +noall +answer example.com SOA
The final number in the answer is the negative caching value. When you query a name that doesn’t exist, the SOA comes back in the authority section, and its TTL counts down in cached responses:
dig +noall +authority newapp.example.com A
Run that against your local resolver a couple of times and you can watch the remaining time tick down.
Compare resolvers to locate the stale cache
The fastest way to troubleshoot is to ask several resolvers the same question:
# the authoritative server: the truth
dig newapp.example.com A @<authoritative-nameserver> +norecurse
# a public resolver
dig newapp.example.com A @1.1.1.1
# whatever your machine uses by default
dig newapp.example.com A
If the authoritative server has the record and your local resolver doesn’t, you’re looking at a cache problem, not a record problem.
Caches are layered
Flushing “the DNS cache” rarely clears just one thing. A single lookup can pass through several caches:
- The application. Browsers keep their own host cache. Chrome’s can be cleared from
chrome://net-internals/#dns. - The operating system.
resolvectl flush-cacheson systemd-resolved,ipconfig /flushdnson Windows, orsudo dscacheutil -flushcache; sudo killall -HUP mDNSResponderon macOS. - The local resolver. A home router, Pi-hole, Unbound or firewall resolver has its own cache, often with its own negative-cache settings.
- The upstream resolver. An ISP or public resolver you can’t flush. Some public resolvers offer a purge tool for specific names.
Flushing your laptop does nothing if the router is the one holding the stale NXDOMAIN.
Habits that avoid the problem
- Create the record before you test it. Don’t poke at a name “to see if it’s there yet.”
- Test against the authoritative server first. It bypasses every cache and tells you whether the change is published.
- Keep negative TTLs reasonable. For zones that change often, a value of a few minutes to an hour is a sensible balance.
- Lower TTLs before planned changes. Drop the TTL a day before a migration, make the change, and raise it again afterwards.
- Cap negative caching on your own resolver. Unbound, for example, has
cache-max-negative-ttl, which limits how long any NXDOMAIN is held no matter what the zone says.
The takeaway
When a DNS fix “didn’t work,” check whether it actually didn’t, or whether something is remembering an old answer. Ask the authoritative server, compare a couple of resolvers, and work out which layer is caching. Most of the time the record is fine and the clock just hasn’t run out yet.
Questions or corrections? Email me.
