Hey all.
We've involved JTAC on this issue, but we haven't heard anything from them and we are in a hurry to find out what this issue might be.
The thing is that we have an SRX4100 were we see traffic drops during every commit. Even the easiest commits triggers the issue, ie. a change of description in a interface name. We observe that during the commits, the nsd process is very eager and consumes almost 100% CPU during a short periode of time. This timespan matches the timespan were we see packet drops.
There are several things that we have observed:
- This issue is only seen on our SRX4100s configured with MNHA. We have several SRX4100s set up in chassis cluster without the issue.
- The SRX has almost 1000 policy rules, but we've removed all of them and this does not seem to solve anything. Also this is way below the number of rules that the SRX4100 is capable of.
- We've removed several DNS-name address book entries that didn't resolve to anything, but this did not help.
- We've tried to remove all SRG1+ except one, to see if that helped. But no fix.
- This is not load related because this SRX has been taken out of the production network and does not currently handle any traffic.
We have seen this PR related to nsd, but it is fixed in 23.4R2-S4 and this is the same version as we are currently running. Also our problem does not match the triggers in the PR.
I understand that this is hard for any of you to find out what might be the cause, but I am interested to hear if anyone has seen similar issues and most of all if anyone has any pointers to how we can get this temporarily fixed until a fixed software version arrives.
------------------------------
Best regards
Vidar Stokke
------------------------------