[PATCH] Improve hypot performance
Wilco Dijkstra
Wilco.Dijkstra@arm.com
Wed Dec 1 17:14:27 GMT 2021
Hi Adhemerval/Paul,
So let me benchmark this on Skylake:
recip-throughput latency
mainline: 28.7292 52.7198
Adhemerval: 34.8628 66.6
New no fma: 22.224 60.9611
-mavx -mfma: 16.1479 29.7945
The fma version wins without contest on x86. It's ~1.8x faster both in
throughput and latency and more than twice as fast as Adhemerval's
version. The non-fma version is 29.3% faster in throughput but with 15.6%
higher latency. So the speedups closely match what I get on AArch64.
Note both my patch and Adhemerval's remove the wrappers. The results
are the best out of 4 runs using 100 times more iterations than the default
in bench-skeleton.c (the default seems too low).
Cheers,
Wilco
More information about the Libc-alpha
mailing list