[PATCH v5 1/1] x86-64: Add vector acos/acosf implementation to libmvec

H.J. Lu hjl.tools@gmail.com
Sun Dec 19 20:26:31 GMT 2021


On Sun, Dec 19, 2021 at 12:29:07PM -0600, GNU C Library wrote:
> On Sun, Dec 19, 2021 at 11:19 AM Sunil K Pandey via Libc-alpha
> <libc-alpha@sourceware.org> wrote:
> >
> > Implement vectorized acos/acosf containing SSE, AVX, AVX2 and
> > AVX512 versions for libmvec as per vector ABI.  It also contains
> > accuracy and ABI tests for vector acos/acosf with regenerated ulps.
> > ---
> 
> Have a few small comments but generally okay with a patch like this
> one going out in
> 2.35.

...

> 
> > +#define poly_coeff_6                   896
> > +#define poly_coeff_7                   960
> > +#define poly_coeff_8                   1024
> > +#define poly_coeff_9                   1088
> > +#define poly_coeff_10                  1152
> > +#define poly_coeff_11                  1216
> > +#define poly_coeff_12                  1280
> > +#define PiH                            1344
> > +#define Pi2H                           1408
> 
> There is enough memory here it may pay to make the accesses

Did you enough registers?

> sequential in memory.

This is based on Intel compiler generated codes.  We will evaluate
Intel compiler changes.

...

> > +
> > +#include <sysdep.h>
> > +        vfmadd231pd {rn-sae}, %zmm3, %zmm11, %zmm10
> > +        andl      %eax, %ecx
> drop I think
> 
> > +        vmovups   poly_coeff_12+__svml_dacos_data_internal(%rip), %zmm11
> > +        kmovw     %ecx, %k3
> kandw %k4, %k2, %k3

This may not be faster since mask register can only go to port 0.  We
will evaluate register allocation in Intel compiler.


Thanks.

H.J.


More information about the Libc-alpha mailing list