No subject
Noah Goldstein
goldstein.w.n@gmail.com
Mon Jul 10 15:58:09 GMT 2023
On Mon, Jul 10, 2023 at 12:23 AM Sajan Karumanchi
<sajan.karumanchi@gmail.com> wrote:
>
> Noah,
> I verified your patches on the master branch that impacts the non-threshold
> parameter on x86 CPUs. This patch modifies the non-temporal threshold value
> from 24MB(3/4th of L3$) to 8MB(1/4th of L3$) on ZEN4.
> From the Glibc benchmarks, we saw a significant performance drop ranging
> from 15% to 99% for size ranges of 8MB to 16MB.
> I also ran the new tool developed by you on all Zen architectures and the
> results conclude that 3/4th L3 size holds good on AMD CPUs.
> Hence the current patch degrades the performance of AMD CPUs.
> We strongly recommend marking this change to Intel CPUs only.
>
So it shouldn't actually go down. I think what is missing is:
```
get_common_cache_info (&shared, &shared_per_thread, &threads, core);
```
In the AMD case shared == shared_per_thread which shouldn't really
be the case.
The intended new calculation is: Total_L3_Size / Scale
as opposed to: (L3_Size / NThread) / Scale"
Before just going with default for AMD, maybe try out the following patch?
```
---
sysdeps/x86/dl-cacheinfo.h | 1 +
1 file changed, 1 insertion(+)
diff --git a/sysdeps/x86/dl-cacheinfo.h b/sysdeps/x86/dl-cacheinfo.h
index c98fa57a7b..c1866ca898 100644
--- a/sysdeps/x86/dl-cacheinfo.h
+++ b/sysdeps/x86/dl-cacheinfo.h
@@ -717,6 +717,7 @@ dl_init_cacheinfo (struct cpu_features *cpu_features)
level3_cache_assoc = handle_amd (_SC_LEVEL3_CACHE_ASSOC);
level3_cache_linesize = handle_amd (_SC_LEVEL3_CACHE_LINESIZE);
+ get_common_cache_info (&shared, &shared_per_thread, &threads, core);
if (shared <= 0)
/* No shared L3 cache. All we have is the L2 cache. */
shared = core;
--
2.34.1
```
> Thanks,
> Sajan K.
>
More information about the Libc-alpha
mailing list