[PATCH] x86-64: Optimize strcat/strncat, strcpy/strncpy and stpcpy/stpncpy with AVX2
H.J. Lu
hjl.tools@gmail.com
Wed Oct 31 18:36:00 GMT 2018
On Mon, Oct 8, 2018 at 6:59 AM
<leonardo.sandoval.gonzalez@linux.intel.com> wrote:
>
> From: Leonardo Sandoval <leonardo.sandoval.gonzalez@linux.intel.com>
>
> Optimize x86-64 strcat/strncat, strcpy/strncpy and stpcpy/stpncpy with AVX2.
> It uses vector comparison as much as possible. In general, the larger the
> source string, the greater performance gain observed, reaching speedups of
> 1.6x compared to SSE2 unaligned routines. Select AVX2 strcat/strncat,
> strcpy/strncpy and stpcpy/stpncpy on AVX2 machines where vzeroupper is
> preferred and AVX unaligned load is fast.
>
> * sysdeps/x86_64/multiarch/Makefile (sysdep_routines): Add
> strcat-avx2, strncat-avx2, strcpy-avx2, strncpy-avx2,
> stpcpy-avx2 and stpncpy-avx2.
> * sysdeps/x86_64/multiarch/ifunc-impl-list.c:
> (__libc_ifunc_impl_list): Add tests for __strcat_avx2,
> __strncat_avx2, __strcpy_avx2, __strncpy_avx2, __stpcpy_avx2
> and __stpncpy_avx2.
> * sysdeps/x86_64/multiarch/{ifunc-unaligned-ssse3.h =>
> ifunc-strcpy.h}: rename header for a more generic name.
> * sysdeps/x86_64/multiarch/ifunc-strcpy.h:
> (IFUNC_SELECTOR): Return OPTIMIZE (avx2) on AVX 2 machines if
> AVX unaligned load is fast and vzeroupper is preferred.
> * sysdeps/x86_64/multiarch/stpcpy-avx2.S: New file
> * sysdeps/x86_64/multiarch/stpncpy-avx2.S: Likewise
> * sysdeps/x86_64/multiarch/strcat-avx2.S: Likewise
> * sysdeps/x86_64/multiarch/strcpy-avx2.S: Likewise
> * sysdeps/x86_64/multiarch/strncat-avx2.S: Likewise
> * sysdeps/x86_64/multiarch/strncpy-avx2.S: Likewise
LGTM.
Thanks.
H.J.
More information about the Libc-alpha
mailing list