[PATCH] x86-64: Optimize strchr/strchrnul/wcschr with AVX2

H.J. Lu hjl.tools@gmail.com
Fri Jun 2 19:49:00 GMT 2017


On Thu, Jun 1, 2017 at 11:13 AM, H.J. Lu <hongjiu.lu@intel.com> wrote:
> Optimize strchr/strchrnul/wcschr with AVX2 to search 32 bytes with vector
> instructions.  It is as fast as SSE2 versions for size <= 16 bytes and up
> to 1X faster for or size > 16 bytes on Haswell.  Select AVX2 version on
> AVX2 machines where vzeroupper is preferred and AVX unaligned load is fast.
>
> NB: It uses TZCNT instead of BSF since TZCNT produces the same result
> as BSF for non-zero input.  TZCNT is faster than BSF and is executed
> as BSF if machine doesn't support TZCNT.
>
> Any comments?
>
> H.J.
>         * sysdeps/x86_64/multiarch/Makefile (sysdep_routines): Add
>         strchr-avx2, strchrnul-avx2 and wcschr-avx2.
>         * sysdeps/x86_64/multiarch/ifunc-impl-list.c
>         (__libc_ifunc_impl_list): Add tests for __strchr_avx2,
>         __strchrnul_avx2, __strchrnul_sse2, __wcschr_avx2 and
>         __wcschr_sse2.
>         * sysdeps/x86_64/multiarch/strchr-avx2.S: New file.
>         * sysdeps/x86_64/multiarch/strchrnul-avx2.S: Likewise.
>         * sysdeps/x86_64/multiarch/strchrnul.S: Likewise.
>         * sysdeps/x86_64/multiarch/wcschr-avx2.S: Likewise.
>         * sysdeps/x86_64/multiarch/wcschr.S: Likewise.
>         * sysdeps/x86_64/multiarch/strchr.S (strchr): Add support for
>         __strchr_avx2.
>

Updated patch with IFUNC selector in C.

-- 
H.J.
-------------- next part --------------
A non-text attachment was scrubbed...
Name: 0006-x86-64-Optimize-strchr-strchrnul-wcschr-with-AVX2.patch
Type: text/x-patch
Size: 21728 bytes
Desc: not available
URL: <http://sourceware.org/pipermail/libc-alpha/attachments/20170602/fbd7e2b6/attachment.bin>


More information about the Libc-alpha mailing list