[PATCH 8/8] math: Use binary search on tgammaf slow path
Adhemerval Zanella Netto
adhemerval.zanella@linaro.org
Mon Oct 13 18:48:16 GMT 2025
On 13/10/25 11:35, Paul Zimmermann wrote:
> Hi Adhemerval,
>
>> From: Adhemerval Zanella <adhemerval.zanella@linaro.org>
>> Cc: Wilco Dijkstra <Wilco.Dijkstra@arm.com>,
>> Paul Zimmermann <Paul.Zimmermann@inria.fr>,
>> DJ Delorie <dj@redhat.com>
>> Date: Fri, 10 Oct 2025 14:49:26 -0300
>>
>> Checked on x86_64-linux-gnu and aarch64-linux-gnu.
>> ---
>> sysdeps/ieee754/flt-32/e_lgammaf_r.c | 56 ++++++++++++++++------------
>> 1 file changed, 32 insertions(+), 24 deletions(-)
>>
>> diff --git a/sysdeps/ieee754/flt-32/e_lgammaf_r.c b/sysdeps/ieee754/flt-32/e_lgammaf_r.c
>> index 75ec25fb9e1..3e3d6c53a10 100644
>> --- a/sysdeps/ieee754/flt-32/e_lgammaf_r.c
>> +++ b/sysdeps/ieee754/flt-32/e_lgammaf_r.c
>> @@ -116,26 +116,7 @@ __ieee754_lgammaf_r (float x, int *signgamp)
>> float f;
>> float df;
>> } tb[] = {
>> - { -0x1.efc2a2p+14, -0x1.222dbcp+18, -0x1p-7 },
>> - { -0x1.627346p+7, -0x1.73235ep+9, -0x1p-16 },
>> - { -0x1.08b14p+4, -0x1.f0cbe6p+4, -0x1p-21 },
>> - { -0x1.69d628p+3, -0x1.0eac2ap+4, -0x1p-21 },
>> - { -0x1.904902p+2, -0x1.65532cp+2, 0x1p-23 },
>> - { -0x1.9272d2p+1, -0x1.170b98p-8, 0x1p-33 },
>> - { -0x1.625edap+1, 0x1.6a6c4ap-5, -0x1p-30 },
>> - { -0x1.5fc2aep+1, 0x1.c0a484p-11, -0x1p-36 },
>> - { -0x1.5fb43ep+1, 0x1.5b697p-17, 0x1p-42 },
>> - { -0x1.5fa20cp+1, -0x1.132f7ap-10, 0x1p-35 },
>> - { -0x1.580c1ep+1, -0x1.5787c6p-4, 0x1p-29 },
>> - { -0x1.3a7fcap+1, -0x1.e4cf24p-24, -0x1p-49 },
>> - { -0x1.c2f04p-30, 0x1.43a6f6p+4, 0x1p-21 },
>> - { -0x1.ade594p-30, 0x1.446ab2p+4, -0x1p-21 },
>> - { -0x1.437e74p-40, 0x1.b7dec2p+4, -0x1p-21 },
>> - { -0x1.d85bfep-43, 0x1.d31592p+4, -0x1p-21 },
>> - { -0x1.f51c8ep-49, 0x1.0a572ap+5, -0x1p-20 },
>> - { -0x1.108a5ap-66, 0x1.6d7b18p+5, -0x1p-20 },
>> - { -0x1.ecf3fep-73, 0x1.8f8e5ap+5, -0x1p-20 },
>> - { -0x1.25cb66p-123, 0x1.547a44p+6, -0x1p-19 },
>> + /* NB: the entries should be sorted by the asuint (x) value. */
>> { 0x1.ecf3fep-73, 0x1.8f8e5ap+5, -0x1p-20 },
>> { 0x1.108a5ap-66, 0x1.6d7b18p+5, -0x1p-20 },
>> { 0x1.a68bbcp-42, 0x1.c9c6e8p+4, 0x1p-21 },
>> @@ -151,6 +132,26 @@ __ieee754_lgammaf_r (float x, int *signgamp)
>> { 0x1.dcbbaap+99, 0x1.fc5772p+105, 0x1p+80 },
>> { 0x1.58ace8p+112, 0x1.9e4f66p+118, -0x1p+93 },
>> { 0x1.87bdfp+115, 0x1.e465aep+121, 0x1p+96 },
>> + { -0x1.25cb66p-123, 0x1.547a44p+6, -0x1p-19 },
>> + { -0x1.ecf3fep-73, 0x1.8f8e5ap+5, -0x1p-20 },
>> + { -0x1.108a5ap-66, 0x1.6d7b18p+5, -0x1p-20 },
>> + { -0x1.f51c8ep-49, 0x1.0a572ap+5, -0x1p-20 },
>> + { -0x1.d85bfep-43, 0x1.d31592p+4, -0x1p-21 },
>> + { -0x1.437e74p-40, 0x1.b7dec2p+4, -0x1p-21 },
>> + { -0x1.ade594p-30, 0x1.446ab2p+4, -0x1p-21 },
>> + { -0x1.c2f04p-30, 0x1.43a6f6p+4, 0x1p-21 },
>> + { -0x1.3a7fcap+1, -0x1.e4cf24p-24, -0x1p-49 },
>> + { -0x1.580c1ep+1, -0x1.5787c6p-4, 0x1p-29 },
>> + { -0x1.5fa20cp+1, -0x1.132f7ap-10, 0x1p-35 },
>> + { -0x1.5fb43ep+1, 0x1.5b697p-17, 0x1p-42 },
>> + { -0x1.5fc2aep+1, 0x1.c0a484p-11, -0x1p-36 },
>> + { -0x1.625edap+1, 0x1.6a6c4ap-5, -0x1p-30 },
>> + { -0x1.9272d2p+1, -0x1.170b98p-8, 0x1p-33 },
>> + { -0x1.904902p+2, -0x1.65532cp+2, 0x1p-23 },
>> + { -0x1.69d628p+3, -0x1.0eac2ap+4, -0x1p-21 },
>> + { -0x1.08b14p+4, -0x1.f0cbe6p+4, -0x1p-21 },
>> + { -0x1.627346p+7, -0x1.73235ep+9, -0x1p-16 },
>> + { -0x1.efc2a2p+14, -0x1.222dbcp+18, -0x1p-7 }
>> };
>
> after checking within CORE-MATH, I found that the following entries are not
> needed in tb[] (tested on x86_64, both with and without -march=x86-64-v2):
> -0x1.08b14p+4, -0x1.904902p+2, -0x1.9272d2p+1, -0x1.625edap+1, -0x1.5fc2aep+1,
> -0x1.5fb43ep+1, -0x1.5fa20cp+1, -0x1.3a7fcap+1.
Ack.
>
> Please could you double-check them?
>
>> float fx = floor (x);
>> @@ -355,10 +356,17 @@ __ieee754_lgammaf_r (float x, int *signgamp)
>> if (__glibc_unlikely (tl <= 31u))
>> {
>> t = asuint (x);
>> - for (unsigned i = 0; i < array_length (tb); i++)
>> - {
>> - if (t == asuint (tb[i].x))
>> - return tb[i].f + tb[i].df;
>> + int a = 0, b = array_length (tb), m = (a + b) / 2;
>> + while (a <= b)
>> + { /* Binary search. */
>> + uint32_t tbi = asuint (tb[m].x);
>> + if (t == tbi)
>> + return tb[m].f + tb[m].df;
>> + else if (t > tbi)
>> + a = m + 1;
>> + else
>> + b = m - 1;
>> + m = (a + b) / 2;
>> }
>> }
>> return r;
>
> this sounds good to me, although I write binary search in a different way:
> if you check t == tbi after the loop, you have only one comparison in the
> loop. The invariant is either t < asuint (tb[0].x), in which case 'a' remains
> 0, or asuint (tb[a].x) <= t < asuint (tb[b].x), in which case you only have to
> compare t and asuint (tb[a].x) at the end, the loop condition is
> while (a + 1 < b), and you set b to m if t < tbi, and a to m otherwise.
>
> Paul
Hum, I am having some trouble to implement this approach. This approach is
similar to the ones from:
src/binary64/acosh/acosh.c: while (a <= b) { // binary search
src/binary64/asinh/asinh.c: while (a <= b) { // binary search
src/binary64/atanh/atanh.c: while (a <= b) { // binary search
src/binary64/lgamma/lgamma.c: while (a <= b) { // binary search
src/binary64/tanpi/tanpi.c: while (a <= b) { // binary search
src/binary64/tgamma/tgamma.c: while (a <= b) { // binary search
src/binary80/exp/expl.c: binary search would be more efficient.
src/binary80/exp2/exp2l.c: while (a <= b) { // binary search
More information about the Libc-alpha
mailing list