[PATCH] aarch64: Improve codegen of AdvSIMD logf function family

Wilco Dijkstra Wilco.Dijkstra@arm.com
Tue Dec 17 15:30:36 GMT 2024


Hi Joana,

> Load the polynomial evaluation coefficients into 2 vectors and use lanewise MLAs.
> 8% improvement in throughput microbenchmark on Neoverse V1 for log2 and log;
> 2% improvement in throughput microbenchmarkon Neoverse V1 for log10.

OK - pushed (needed an extra return at the end).

Cheers,
Wilco


More information about the Libc-alpha mailing list