This is the mail archive of the
libc-alpha@sourceware.org
mailing list for the glibc project.
Re: Commit 27d3ce1467990f89126e228559dec8f84b96c60e?
- From: "H.J. Lu" <hjl dot tools at gmail dot com>
- To: "Carlos O'Donell" <carlos at redhat dot com>
- Cc: libc-alpha <libc-alpha at sourceware dot org>, Florian Weimer <fweimer at redhat dot com>
- Date: Tue, 3 Dec 2019 09:42:55 -0800
- Subject: Re: Commit 27d3ce1467990f89126e228559dec8f84b96c60e?
- References: <8e209ae2-969e-d4d5-bc00-0111c85198a6@redhat.com>
On Tue, Nov 26, 2019 at 7:01 AM Carlos O'Donell <carlos@redhat.com> wrote:
>
> HJ,
>
> In commit 27d3ce1467990f89126e228559dec8f84b96c60e we stop
> setting bit_arch_Fast_Copy_Backward for Intel Core processors
> as an optimization to improve performance.
>
> It turns out that this change also improves performance for
> Haswell servers. Was it the intent of this change to *also*
> improve performance for Haswell? The comments don't indicate
> this and I was worried that it might be an unintentional change
> in this case. The particular CPU was a E5-2650 v3.
>
> If we step back and look at the overall sequence of changes and
> performance it looks like this:
>
> The performance regression is between this change:
> c3d8dc45c9df199b8334599a6cbd98c9950dba62 - Triggers default: handling + TSX handling.
> - Causes a 21% lmbench regression for an E5-2650 v3.
>
> and this change (the one we are discussing):
> 27d3ce1467990f89126e228559dec8f84b96c60e - Removes bit_arch_Fast_Copy_Backward.
> - Restores the performance loss.
>
> My worry is that the two are unrelated, and that we've only
> just made back performance at the expense of the other change
> and we could be doing better.
>
> As our Intel expert what do you think is going on here?
My change should be a NOP on Haswell since Fast_Copy_Backward is
used only in x86_64/multiarch/ifunc-memmove.h:
if (CPU_FEATURES_ARCH_P (cpu_features, AVX_Fast_Unaligned_Load))
{
if (CPU_FEATURES_CPU_P (cpu_features, ERMS))
return OPTIMIZE (avx_unaligned_erms);
return OPTIMIZE (avx_unaligned);
}
if (!CPU_FEATURES_CPU_P (cpu_features, SSSE3)
|| CPU_FEATURES_ARCH_P (cpu_features, Fast_Unaligned_Copy))
{
if (CPU_FEATURES_CPU_P (cpu_features, ERMS))
return OPTIMIZE (sse2_unaligned_erms);
return OPTIMIZE (sse2_unaligned);
}
if (CPU_FEATURES_ARCH_P (cpu_features, Fast_Copy_Backward))
return OPTIMIZE (ssse3_back);
return OPTIMIZE (ssse3);
and AVX_Fast_Unaligned_Load is set on Haswell.
> --
> Cheers,
> Carlos.
>
> [1] https://ark.intel.com/content/www/us/en/ark/products/81705/intel-xeon-processor-e5-2650-v3-25m-cache-2-30-ghz.html
>
--
H.J.