[PATCH] Improve hypot performance

Wilco Dijkstra Wilco.Dijkstra@arm.com
Wed Dec 1 17:14:27 GMT 2021


Hi Adhemerval/Paul,

So let me benchmark this on Skylake:

        recip-throughput  latency
mainline:    28.7292      52.7198
Adhemerval:  34.8628      66.6
New no fma:  22.224       60.9611
-mavx -mfma: 16.1479      29.7945

The fma version wins without contest on x86. It's ~1.8x faster both in
throughput and latency and more than twice as fast as Adhemerval's
version. The non-fma version is 29.3% faster in throughput but with 15.6%
higher latency. So the speedups closely match what I get on AArch64.

Note both my patch and Adhemerval's remove the wrappers. The results
are the best out of 4 runs using 100 times more iterations than the default
in bench-skeleton.c (the default seems too low).

Cheers,
Wilco


More information about the Libc-alpha mailing list