Intermittent DNS Cache Issue with nscd on glibc 2.28 After /etc/resolv.conf Updates
debianyu
debianyu@163.com
Thu Dec 18 10:11:46 GMT 2025
We are experiencing an intermittent DNS caching issue in our production environment that appears to be
related to nscd's handling of /etc/resolv.conf updates. After researching similar problems, we found your
patch discussion at https://inbox.sourceware.org/libc-alpha/87ikyhqfwy.fsf@oldenburg.str.redhat.com/
which seems highly relevant to our situation.
Environment Details:
OS: Custom Linux distribution (based on CentOS 8)
glibc version: 2.28
nscd version: 2.28 (bundled with glibc)
Network management: systemd-networkd and NetworkManager coexistence
Problem Description:
We periodically perform environment migrations where DNS infrastructure changes. After updating /etc/resolv.conf with new nameserver entries, we observe that ping and other glibc-based applications intermittently continue using the old, now-unreachable nameservers.
Key Observations:
1 The problem is intermittent - occurring in approximately 1-3% of environment transitions
2 /etc/resolv.conf is definitively updated with new nameservers (verified via timestamps and content checks)
3 nslookup and dig work correctly (they bypass nscd cache)
4 The issue persists across multiple DNS queries, not just the first one after the change
5 Executing systemctl restart nscd immediately resolves the problem
6 nscd configuration has check-files hosts yes set
7 We've verified that nscd is running and appears to be functioning normally aside from this issue
Questions:
1 Is the patch discussed in the referenced thread included in glibc 2.28?
2 If not, are there known workarounds for this issue in glibc 2.28?
3 Does nscd's check-files mechanism fully handle all edge cases of resolv.conf updates?
4 Are there known issues with nscd and mixed network management environments (systemd-networkd + NetworkManager)?
Temporary Workaround:
We currently work around this by:
1 Forcing nscd restart after DNS changes: systemctl restart nscd
However, we seek a more robust solution that doesn't require service restarts during critical migrations.
Request:
Could you provide guidance on whether this is a known issue in glibc 2.28 and if there are any patches or
configuration changes that would resolve it? We'd be happy to provide additional diagnostic information or
test potential fixes.
More information about the Libc-help
mailing list