[PATCH v5 1/1] x86-64: Add vector acos/acosf implementation to libmvec
H.J. Lu
hjl.tools@gmail.com
Sun Dec 19 20:26:31 GMT 2021
On Sun, Dec 19, 2021 at 12:29:07PM -0600, GNU C Library wrote:
> On Sun, Dec 19, 2021 at 11:19 AM Sunil K Pandey via Libc-alpha
> <libc-alpha@sourceware.org> wrote:
> >
> > Implement vectorized acos/acosf containing SSE, AVX, AVX2 and
> > AVX512 versions for libmvec as per vector ABI. It also contains
> > accuracy and ABI tests for vector acos/acosf with regenerated ulps.
> > ---
>
> Have a few small comments but generally okay with a patch like this
> one going out in
> 2.35.
...
>
> > +#define poly_coeff_6 896
> > +#define poly_coeff_7 960
> > +#define poly_coeff_8 1024
> > +#define poly_coeff_9 1088
> > +#define poly_coeff_10 1152
> > +#define poly_coeff_11 1216
> > +#define poly_coeff_12 1280
> > +#define PiH 1344
> > +#define Pi2H 1408
>
> There is enough memory here it may pay to make the accesses
Did you enough registers?
> sequential in memory.
This is based on Intel compiler generated codes. We will evaluate
Intel compiler changes.
...
> > +
> > +#include <sysdep.h>
> > + vfmadd231pd {rn-sae}, %zmm3, %zmm11, %zmm10
> > + andl %eax, %ecx
> drop I think
>
> > + vmovups poly_coeff_12+__svml_dacos_data_internal(%rip), %zmm11
> > + kmovw %ecx, %k3
> kandw %k4, %k2, %k3
This may not be faster since mask register can only go to port 0. We
will evaluate register allocation in Intel compiler.
Thanks.
H.J.
More information about the Libc-alpha
mailing list