[EXT] [PATCH 6/7] powerpc: Remove modf optimization
Peter Bergner
bergner@tenstorrent.com
Wed May 28 20:31:15 GMT 2025
On 5/28/25 7:59 AM, Adhemerval Zanella wrote:
> The generic implementation is slight more optimized than the powerpc
> one, where it has a more optimized inf/nan check (by not using FP
> unit checks, along with branch prediction hints), and removed one
> branch by issuing trunc instead of a combination of floor/ceil (which
> also generated less code).
>
> On power10 with gcc 14.2.1:
>
> reciprocal-throughput master patch difference
> workload-0_1 1.1351 0.9067 20.12%
> workload-1_maxint 1.4230 0.9040 36.47%
> workload-maxint_maxfloat 1.5038 0.9076 39.65%
> workload-integral 1.1280 0.9111 19.23%
>
> latency master patch difference
> workload-0_1 1.1440 2.7117 -137.03%
> workload-1_maxint 4.0556 2.7070 33.25%
> workload-maxint_maxfloat 3.2122 2.7164 15.43%
> workload-integral 3.2381 2.7281 15.75%
I like the idea of using an optimized generic routine over an arch
specific routine, but I'm confused by your data. For the data above,
it looks like the throughput is less with the patch (ie, smaller number)
and the latency is longer with the patch (ie, bigger number). ...and not
just on Power, but x86 as well. Is that just my confusion? If so, what
are these data values actually specifying?
Peter
More information about the Libc-alpha
mailing list