The most frustrating DNS problem isn't a total outage — it's the one that heals itself before you can screenshot it. Sites fail for three seconds, reload, and work fine. Your monitoring shows green, but users keep complaining. Intermittent DNS failures have a specific fingerprint: they surface under load, after DHCP renewals, during ISP maintenance windows, or when two resolvers disagree on an answer. This guide walks through a disciplined methodology to identify exactly which layer is failing and how to fix it permanently.

Define the Failure Pattern First

Before touching anything, write down the exact failure mode. "Intermittent" can mean several distinct things, each with different root causes:

  • Random individual lookups fail — SERVFAIL or NXDOMAIN for a domain that resolves fine on immediate retry.
  • All DNS fails for 30–120 seconds, then recovers — typically a DHCP lease renewal clobbering the DNS server assignment.
  • Works on some devices, not others — different resolvers configured, or IPv4/IPv6 preference mismatch between devices.
  • Fails under load only — rate-limiting on the upstream resolver, or a congested DNS relay on the router.
  • Fails at specific times of day — ISP maintenance windows, TTL expiry during peak hours, or a cron job hammering the local resolver.

That pattern is your primary diagnostic signal. Don't skip this step.

Root Causes, Ranked by Frequency

1. Flaky ISP Resolver

The most common cause by far. Your ISP's recursive resolver is a shared, overloaded service that suffers packet loss, rate-limits, or silent drops under congestion. The symptom is a random SERVFAIL that disappears on retry because the second query hits a different server in their anycast pool. It appears random because it is — at the resolver tier.

2. DHCP Lease Renewal Clobbering DNS

On many consumer routers, when a DHCP lease renews, the client temporarily loses its DNS server assignment for 1–5 seconds. Applications that open new TCP connections during this window see DNS failures. This explains the "brief outage every few hours" pattern that correlates with lease duration (commonly 2h or 4h on ISP-supplied modems).

3. Negative Caching of NXDOMAIN

If an authoritative server briefly returns NXDOMAIN — due to a zone propagation race, a misconfigured API-driven DNS update, or a momentary timeout — your resolver caches that negative answer for the SOA record's negative TTL, sometimes up to 3600 seconds. Every device using that resolver gets the wrong answer for an hour even after the underlying problem is fixed.

4. Router DNS Relay and Proxy Bugs

Consumer routers act as DNS proxies by default — they hand out 192.168.x.1 as the DNS server and forward upstream. This relay is buggy in most firmware: it may truncate large UDP responses, fail to handle EDNS0 properly, drop responses for queries that take longer than ~800ms, or silently stop forwarding during firmware updates or memory pressure.

5. DNS Round-Robin Returning Dead IPs

If the domain returns multiple A records and one of those IPs is down or blocking connections, roughly half of all sessions fail. Each individual DNS lookup succeeds — the failure is at the TCP layer, not DNS. This is one of the most common intermittent DNS misdiagnoses.

6. IPv6 Resolver Unreachability

Most home networks now have native IPv6. Operating systems prefer AAAA records and IPv6 resolver addresses. If your ISP's IPv6 resolver is reachable but the upstream path to the authoritative server has IPv6 routing issues, lookups fail intermittently only for domains with AAAA records, and only on devices that prefer IPv6 — which in 2026 is most modern hardware.

7. DoH and DoT Fallback Loops

Browsers and operating systems — iOS 14+, Windows 11, Android 9+ — now use DNS-over-HTTPS or DNS-over-TLS with fallback to plain UDP. If the DoH endpoint is slow or intercepted by a corporate proxy, you get 3–8 second delays before fallback resolves the query. This looks like intermittent slowness rather than hard failure, and it bypasses OS-level DNS cache flushes entirely.

8. DNSSEC Validation Failures

If the domain has DNSSEC enabled and the resolver validates signatures, a misconfigured DS record, an expired RRSIG, or an incomplete key rollover causes the resolver to return SERVFAIL — even though the domain resolves fine from a non-validating resolver. This creates a split-brain effect: users on 8.8.8.8 or 1.1.1.1 fail while users on non-validating ISP resolvers succeed.

Step-by-Step Diagnosis

Step 1: Capture a Failure in Real Time

You cannot diagnose intermittent faults by checking when things are working. Set up continuous monitoring before you do anything else:

watch -n 2 "dig @8.8.8.8 example.com A +short +time=2 +tries=1 || echo FAILED $(date +%T)"

On Windows PowerShell:

while ($true) { $r = Resolve-DnsName example.com -ErrorAction SilentlyContinue if (-not $r) { Write-Host "$(Get-Date -Format HH:mm:ss) FAILED" } else { Write-Host "$(Get-Date -Format HH:mm:ss) OK" } Start-Sleep 2 }

Leave this running in the background while you use the system normally. The timestamp on each failure correlates with DHCP lease events, cron jobs, or load patterns.

Step 2: Isolate the Layer

Query three different resolver tiers and compare:

# Your local router (from DHCP) dig @$(awk '/nameserver/{print $2; exit}' /etc/resolv.conf) example.com A +short # Google Public DNS (removes ISP from the path) dig @8.8.8.8 example.com A +short +time=3 # Authoritative server directly (bypasses all caching) dig example.com NS +short dig @ns1.example-ns.com example.com A +time=3

If the local resolver fails but 8.8.8.8 succeeds: fault is your LAN or router. If 8.8.8.8 fails but the authoritative works: ISP or upstream resolver. If the authoritative itself fails: DNS configuration on the domain owner's side.

Step 3: Check for Negative Caching

# Check SOA for negative TTL (last numeric field) dig @8.8.8.8 example.com SOA +short # ns1.example.com. admin.example.com. 2026010101 3600 900 604800 300 # ^^^ # negative TTL = 300 seconds # Query for NXDOMAIN status dig @8.8.8.8 example.com A

If you're seeing cached NXDOMAIN and the negative TTL is high, flush the resolver cache (see OS-specific commands below) and retest immediately.

Step 4: Test for DNSSEC Breakage

# Without CD flag — validation enforced dig @8.8.8.8 example.com A +dnssec # With CD flag — validation disabled dig @8.8.8.8 example.com A +cd # If +cd works but the normal query SERVFAILs: DNSSEC break confirmed # Check which resolver validates vs does not: dig @9.9.9.9 example.com A # Quad9: validates DNSSEC dig @9.9.9.10 example.com A # Quad9: no DNSSEC enforcement
💡 Getting inconsistent results across devices? Different resolvers may have stale or broken answers cached. Use the DNS Propagation Checker to query 20+ global resolver nodes simultaneously — it shows you exactly which ones have NXDOMAIN, SERVFAIL, or stale records cached right now.

Step 5: Bypass the Router Relay

Temporarily point a device directly at a public resolver to test whether the router is the fault source:

  • Windows: Control Panel → Network and Sharing Center → change adapter settings → right-click adapter → Properties → Internet Protocol Version 4 → Use the following DNS server addresses: 8.8.8.8 / 8.8.4.4
  • Mac: System Settings → Network → your interface → Details → DNS tab → click + → add 8.8.8.8
  • Linux: Add nameserver 8.8.8.8 as the first line of /etc/resolv.conf, or configure via NetworkManager with nmcli con mod "connection-name" ipv4.dns "8.8.8.8 1.1.1.1"
  • iOS: Settings → Wi-Fi → tap your network → Configure DNS → Manual → remove existing entries → add 8.8.8.8
  • Android: Settings → Network & Internet → Wi-Fi → long-press network → Modify network → Advanced options → IP settings: Static → DNS 1: 8.8.8.8

If intermittent failures stop, your router's DNS relay is the culprit.

OS-Specific Resolver Flush Commands

Windows

# Flush DNS resolver cache (run as Administrator) ipconfig /flushdns # Check current resolver configuration ipconfig /all | findstr /i "DNS Servers" # Query with explicit server (verify bypass works) nslookup example.com 8.8.8.8 # Windows 11 with DoH — check DoH server configuration Get-DnsClientDohServerAddress

macOS

# Flush DNS cache (macOS 12 Monterey and later) sudo dscacheutil -flushcache; sudo killall -HUP mDNSResponder # Inspect what resolver macOS is actually using scutil --dns | head -40 # Real-time DNS resolution trace dns-sd -Q example.com A

Linux with systemd-resolved

# Check status and cache statistics resolvectl status resolvectl statistics # Flush all DNS caches resolvectl flush-caches # Query with verbose output including DNSSEC status resolvectl query example.com --legend=yes # For non-systemd systems using nscd: sudo nscd -i hosts
💡 Need to inspect a specific record type — SOA, TXT, AAAA, MX — without installing command-line tools? The DNS Lookup tool lets you query any record type against any resolver or authoritative nameserver from your browser, useful for comparing answers across multiple resolvers in seconds.

iOS and Android

Neither platform exposes a direct DNS cache flush to users. On iOS, toggling Airplane Mode on and off forces a full network reinit including resolver reset — the fastest option. On Android 9+ with Private DNS (DoT) configured, go to Settings → Network → Advanced → Private DNS and toggle between Off and your server to force a resolver reinit. For Android 12+, simply toggling Wi-Fi off and on resets the DNS cache.

Fixing Each Root Cause

Switching Resolvers at the Router

Set reliable public resolvers at the router level so every LAN device benefits without per-device changes. Admin paths by router brand:

  • TP-Link (standard firmware): tplinkwifi.net or 192.168.0.1 → Advanced → Network → DHCP Server → Primary DNS / Secondary DNS
  • ASUS: asusrouter.com or 192.168.1.1 → WAN → Internet Connection → DNS Server → set to Manual → 8.8.8.8 / 8.8.4.4
  • Netgear Orbi: orbilogin.com → Advanced → Setup → Internet Setup → Domain Name Server (DNS) Address
  • Linksys Velop/MR series: linksyssmartwifi.com → Connectivity → Local Network → DHCP Server → DNS
  • Xiaomi: miwifi.com or 192.168.31.1 → More Settings → LAN Settings → DNS
  • Netgear R-series: routerlogin.net → Internet → Domain Name Server (DNS) Address
  • Most generic routers at 192.168.1.1: WAN Settings or Internet → DNS

On OpenWrt: edit /etc/config/dhcp, add option server '8.8.8.8' and option server '1.1.1.1' under the dnsmasq stanza, then run service dnsmasq restart. On DD-WRT: Services → Services → Additional DNSMasq Options: add server=8.8.8.8 on its own line.

Fixing DHCP Lease DNS Gaps

Increase DHCP lease time to 24–48h on consumer LANs — shorter leases mean more renewal events, more gaps. Set static DHCP reservations for critical machines. On Linux with NetworkManager, prevent lease renewals from clobbering resolv.conf by adding dns=none under [main] in /etc/NetworkManager/NetworkManager.conf and managing /etc/resolv.conf manually.

Fixing Router Relay Bugs

Two approaches: update router firmware (most EDNS0 and large-response handling bugs were fixed in 2023–2024 firmware releases), or disable the DNS relay entirely by setting the DHCP "DNS Server" field to the public resolver IP directly (8.8.8.8) instead of the router's LAN IP. Clients then talk to the upstream resolver directly with no router in the DNS path.

Fixing DNSSEC Failures

If you own the domain: log into your registrar, find the DNSSEC settings, and verify the DS record in the parent zone matches your current DNSKEY. A key rollover that didn't update the DS record at the registrar is the most common cause. The fix is updating or re-entering the DS record through the registrar's UI — not through your DNS host. If you don't own the domain, use a non-enforcing resolver: Quad9's 9.9.9.10 (with DNSSEC disabled) or temporarily set your device DNS to a non-validating ISP resolver while the domain owner fixes the chain.

Fixing DoH and DoT Fallback Delays

On Windows 11: Settings → Network & Internet → your connection → Hardware properties → DNS Server Assignment — verify the DoH toggle state matches your network's capabilities. On networks where DoH is transparently intercepted by a proxy, disable it via Group Policy: Computer Configuration → Administrative Templates → Network → DNS Client → Turn off DNS over HTTPS. On Android 9+: Settings → Network & Internet → Advanced → Private DNS → Off to revert to plain UDP if the DoT server is unreliable.

Confirming the Fix Worked

One successful lookup after a change proves nothing. Run a sustained test:

# 60 queries at 2-second intervals — count failures fail=0; for i in $(seq 1 60); do dig @8.8.8.8 example.com A +short +time=2 +tries=1 > /dev/null 2>&1 || { echo "FAIL $i"; ((fail++)); } sleep 2 done; echo "Failures: $fail / 60"

Zero failures across 60 queries over two minutes is reasonable short-term confirmation. For production systems, configure a monitoring check in Uptime Kuma, Prometheus blackbox exporter, or a simple cron + curl that alerts specifically on DNS resolution failure — not just HTTP status code, which can mask upstream DNS errors.

Preventing Recurrence

  • Configure two upstream resolvers from different providers (8.8.8.8 + 1.1.1.1) so a single provider outage doesn't take everything down.
  • For domains you control, keep TTLs at 300–600 seconds in normal operation. Drop them to 60s 48 hours before any DNS change and restore afterward.
  • Monitor DNSSEC chain health on your own domains — most registrar dashboards show RRSIG expiry. Tools like Zonemaster alert on chain breaks before they cause user-visible failures.
  • If running a local forwarder (dnsmasq, Unbound, Pi-hole), always configure at least two upstream forwarders. For Unbound add two forward-addr: lines. For dnsmasq add two server= lines.
  • Keep router firmware current — EDNS0 handling, TCP fallback for large responses, and IPv6 resolver support have all improved substantially in 2024–2026 firmware generations.

Common Misdiagnoses

"It must be DNS" when it's TCP or routing: If dig succeeds but connections fail, DNS returned the right IP but that IP is unreachable or blocking. Confirm with traceroute or mtr after verifying the DNS answer is correct. Don't restart DNS services when the problem is upstream routing.

Blaming the CDN for normal behavior: Different queries legitimately return different IPs via anycast or round-robin. That's expected. The problem only exists if some of those IPs are failing health checks. If dig example.com A returns three IPs and one of them drops connections, that's a CDN infrastructure issue — contact the CDN provider with the specific IP.

Confusing browser DNS cache with OS cache: Chrome maintains its own DNS cache independent of the OS. Flushing the OS cache leaves Chrome's cache intact. Check chrome://net-internals/#dns and click "Clear host cache" if the issue is browser-specific. Firefox has a similar cache at about:networking#dns.

Treating SERVFAIL and NXDOMAIN as the same: SERVFAIL means the resolver failed to get an answer — network issue, DNSSEC break, or resolver overload. NXDOMAIN means the domain genuinely does not exist at that resolver at that moment. They have different root causes, different caching behavior, and different fixes. Reading the status code in your dig output is the first diagnostic step, not an afterthought.

IPv6, DoH, and DNSSEC in 2026

Three features have moved from edge case to default-on in the past two years and are now active causes of intermittent failures on older hardware and configurations.

IPv6 AAAA preference: Windows 11 23H2+, macOS 14+, iOS 17+, and Android 14 all prefer AAAA records and IPv6 DNS resolver addresses. If your router's IPv6 DNS proxy is buggy — common on TP-Link hardware before firmware 1.4.x and ASUS firmware before 3.0.0.6.x — only IPv6-initiated lookups fail, which is roughly half of all queries on a dual-stack network. The symptom correlates with specific domains (those with AAAA records) rather than being truly random. Fix: update firmware, or set your DHCP-distributed DNS server addresses to IPv4-only addresses of your chosen resolver until you confirm the IPv6 resolver path is clean.

DoH interception and ECH: Cloudflare, Google Public DNS, and Apple's iCloud Private Relay all route DNS over HTTPS on port 443 in current OS defaults. Corporate and ISP proxies that perform TLS inspection without Encrypted Client Hello support cause DoH queries to silently fail. Fallback to plain UDP adds 2–8 seconds per lookup cycle. Per Google's Public DNS documentation, DoH endpoint reachability can be verified by fetching https://dns.google/resolve?name=example.com&type=A directly from the affected network.

DNSSEC at scale: Roughly 38% of active TLDs enforce DNSSEC validation as of 2026. Automated key rollovers — increasingly common as registrars move to programmatic DNSSEC management — can push a broken DS record to the parent zone within minutes. A botched rollover causes SERVFAIL for all validating resolvers worldwide while non-validating resolvers continue answering normally. If a domain suddenly fails for users on 8.8.8.8 and 1.1.1.1 but works on ISP resolvers, run dig @8.8.8.8 example.com A +cd — if it resolves with +cd but fails without it, the DNSSEC chain is broken and the domain owner needs to fix the DS record at their registrar immediately.