[PATCH] math: Fix x86_64 build for -Os (BZ 33367)
H.J. Lu
hjl.tools@gmail.com
Mon Sep 8 21:05:01 GMT 2025
On Mon, Sep 8, 2025 at 2:01 PM H.J. Lu <hjl.tools@gmail.com> wrote:
>
> On Mon, Sep 8, 2025 at 12:55 PM Adhemerval Zanella Netto
> <adhemerval.zanella@linaro.org> wrote:
> >
> >
> >
> > On 08/09/25 16:34, H.J. Lu wrote:
> > > On Mon, Sep 8, 2025 at 12:09 PM Adhemerval Zanella Netto
> > > <adhemerval.zanella@linaro.org> wrote:
> > >>
> > >>
> > >>
> > >> On 08/09/25 15:59, H.J. Lu wrote:
> > >>> On Mon, Sep 8, 2025 at 9:30 AM H.J. Lu <hjl.tools@gmail.com> wrote:
> > >>>>
> > >>>> On Mon, Sep 8, 2025 at 9:04 AM Adhemerval Zanella Netto
> > >>>> <adhemerval.zanella@linaro.org> wrote:
> > >>>>>
> > >>>>>
> > >>>>>
> > >>>>> On 08/09/25 12:56, H.J. Lu wrote:
> > >>>>>> On Mon, Sep 8, 2025 at 8:27 AM Adhemerval Zanella Netto
> > >>>>>> <adhemerval.zanella@linaro.org> wrote:
> > >>>>>>>
> > >>>>>>>
> > >>>>>>>
> > >>>>>>> On 08/09/25 12:22, H.J. Lu wrote:
> > >>>>>>>> On Mon, Sep 8, 2025 at 7:26 AM H.J. Lu <hjl.tools@gmail.com> wrote:
> > >>>>>>>>>
> > >>>>>>>>> On Sat, Sep 6, 2025 at 3:22 PM Adhemerval Zanella
> > >>>>>>>>> <adhemerval.zanella@linaro.org> wrote:
> > >>>>>>>>>>
> > >>>>>>>>>> The compiler might not inline the trunc function call for
> > >>>>>>>>>> USE_TRUNC_BUILTIN, which requires adding the required objects.
> > >>>>>>>>>>
> > >>>>>>>>>> Route trunc calls for the optimized modf variants to thei
> > >>>>>>>>>> ABI-compatible version. This avoids invoking the ifunc variant
> > >>>>>>>>>> for x86_64-v1 and requires fewer objects to be added.
> > >>>>>>>>>>
> > >>>>>>>>>> Checked on x86_64, x86_64-v2, x86_64-v3, and x86_64-v4. Used -O2 and
> > >>>>>>>>>> -Os options. Performed a full make check on x86_64 with both
> > >>>>>>>>>> optimizations.
> > >>>>>>>>>> ---
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/Makefile | 19 ++++++++++++-------
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modf-avx.c | 5 +++++
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modf-sse4_1.c | 5 +++++
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modf.c | 1 -
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modff-avx.c | 5 +++++
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modff-sse4_1.c | 5 +++++
> > >>>>>>>>>> sysdeps/x86_64/fpu/multiarch/s_modff.c | 1 -
> > >>>>>>>>>> 7 files changed, 32 insertions(+), 9 deletions(-)
> > >>>>>>>>>>
> > >>>>>>>>>> diff --git a/sysdeps/x86_64/fpu/multiarch/Makefile b/sysdeps/x86_64/fpu/multiarch/Makefile
> > >>>>>>>>>> index 708b14297f..baefe8dd07 100644
> > >>>>>>>>>> --- a/sysdeps/x86_64/fpu/multiarch/Makefile
> > >>>>>>>>>> +++ b/sysdeps/x86_64/fpu/multiarch/Makefile
> > >>>>>>>>>> @@ -26,6 +26,11 @@ CFLAGS-s_sinf-fma.c = -mfma -mavx2
> > >>>>>>>>>> CFLAGS-s_cosf-fma.c = -mfma -mavx2
> > >>>>>>>>>> CFLAGS-s_sincosf-fma.c = -mfma -mavx2
> > >>>>>>>>>>
> > >>>>>>>>>> +sysdep_calls += \
> > >>>>>>>>>> + s_trunc-sse4_1 \
> > >>>>>>>>>> + s_truncf-sse4_1 \
> > >>>>>>>>>> +# sysdep_calls
> > >>>>>>>>>> +
> > >>>>>>>>>> # Check if ISA level is 2 or above.
> > >>>>>>>>>> ifeq (,$(filter $(have-x86-isa-level),$(x86-isa-level-2-or-above)))
> > >>>>>>>>>> sysdep_calls += \
> > >>>>>>>>>> @@ -36,6 +41,10 @@ endif
> > >>>>>>>>>>
> > >>>>>>>>>> # Check if ISA level is 3 or above.
> > >>>>>>>>>> ifneq (,$(filter $(have-x86-isa-level),$(x86-isa-level-3-or-above)))
> > >>>>>>>>>> +sysdep_calls += \
> > >>>>>>>>>> + s_trunc-avx \
> > >>>>>>>>>> + s_truncf-avx \
> > >>>>>>>>>> +# sysdep_calls
> > >>>>>>>>>> libm-sysdep_routines += \
> > >>>>>>>>>> s_ceil-avx \
> > >>>>>>>>>> s_ceilf-avx \
> > >>>>>>>>>> @@ -47,8 +56,6 @@ libm-sysdep_routines += \
> > >>>>>>>>>> s_rintf-avx \
> > >>>>>>>>>> s_roundeven-avx \
> > >>>>>>>>>> s_roundevenf-avx \
> > >>>>>>>>>> - s_trunc-avx \
> > >>>>>>>>>> - s_truncf-avx \
> > >>>>>>>>>> # libm-sysdep_routines
> > >>>>>>>>>> else
> > >>>>>>>>>> sysdep_calls += \
> > >>>>>>>>>> @@ -114,8 +121,6 @@ libm-sysdep_routines += \
> > >>>>>>>>>> s_tan-avx \
> > >>>>>>>>>> s_tan-fma \
> > >>>>>>>>>> s_tanh-fma \
> > >>>>>>>>>> - s_trunc-sse4_1 \
> > >>>>>>>>>> - s_truncf-sse4_1 \
> > >>>>>>>>>> # libm-sysdep_routines
> > >>>>>>>>>> ifeq ($(have-x86-isa-level),baseline)
> > >>>>>>>>>> libm-sysdep_routines += \
> > >>>>>>>>>> @@ -132,13 +137,13 @@ libm-sysdep_routines += \
> > >>>>>>>>>> s_trunc-c \
> > >>>>>>>>>> s_truncf-c \
> > >>>>>>>>>> # libm-sysdep_routines
> > >>>>>>>>>> -endif
> > >>>>>>>>>> +endif # ($(have-x86-isa-level),baseline)
> > >>>>>>>>>> +endif # (,$(filter $(have-x86-isa-level),$(x86-isa-level-3-or-above)))
> > >>>>>>>>>>
> > >>>>>>>>>> # $(sysdep_calls) functions are built both for libc and libm. While the
> > >>>>>>>>>> # libc objects have the prefix s_, the libm ones are prefixed with m_.
> > >>>>>>>>>> -sysdep_routines += $(sysdep_calls)
> > >>>>>>>>>> +sysdep_routines += $(sysdep_calls)
> > >>>>>>>>>> libm-sysdep_routines += $(sysdep_calls:s_%=m_%)
> > >>>>>>>>>> -endif
> > >>>>>>>>>>
> > >>>>>>>>>> CFLAGS-e_asin-fma4.c = -mfma4
> > >>>>>>>>>> CFLAGS-e_atan2-fma4.c = -mfma4
> > >>>>>>>>>> diff --git a/sysdeps/x86_64/fpu/multiarch/s_modf-avx.c b/sysdeps/x86_64/fpu/multiarch/s_modf-avx.c
> > >>>>>>>>>> index ab4f03db0e..25617c4fdc 100644
> > >>>>>>>>>> --- a/sysdeps/x86_64/fpu/multiarch/s_modf-avx.c
> > >>>>>>>>>> +++ b/sysdeps/x86_64/fpu/multiarch/s_modf-avx.c
> > >>>>>>>>>> @@ -1,3 +1,8 @@
> > >>>>>>>>>> +#include <sysdeps/x86/isa-level.h>
> > >>>>>>>>>> +#if MINIMUM_X86_ISA_LEVEL < SSE4_1_X86_ISA_LEVEL
> > >>>>>>>>>> +asm ("__trunc = __trunc_sse41");
> > >>>>>>>>>
> > >>>>>>>>> It looks odd for __modf_avx to call __trunc_sse41. But this is a
> > >>>>>>>>> separate issue.
> > >>>>>>>>> We do have s_truncf-avx.S, but it isn't used by default. I think we should
> > >>>>>>>>> make it available for s_modf.c and s_modf-avx.c first.
> > >>>>>>>>
> > >>>>>>>> Can we use inline functions?
> > >>>>>>>
> > >>>>>>> It might an option, and I think it simplified the Makefile requirement.
> > >>>>>>> I will update the patch with this.
> > >>>>>>>
> > >>>>>>> Do we have a way to force compiler always expand the builin
> > >>>>>>> regardless of the optimization level?
> > >>>>>>
> > >>>>>> #include <math.h>
> > >>>>>> #include <x86intrin.h>
> > >>>>>>
> > >>>>>> double
> > >>>>>> __trunc (double x)
> > >>>>>> {
> > >>>>>> #ifdef __AVX__
> > >>>>>> asm ("vroundsd $11, %1, %1, %0" : "=v" (x) : "v" (x));
> > >>>>>> #elif defined __SSE4_1__
> > >>>>>> asm ("roundsd $11, %1, %0" : "=x" (x) : "x" (x));
> > >>>>>> #else
> > >>>>>> x = trunc (x);
> > >>>>>> #endif
> > >>>>>> return x;
> > >>>>>> }
> > >>>>>>
> > >>>>> Yeah, that was my idea to change it. But my question was if there is an extra
> > >>>>> compiler option to avoid the libcall generation for -Os if the builtin lowering
> > >>>>> does allow to use an instruction.
> > >>>>
> > >>>> I don't know.
> > >>>>
> > >>>>> As a side note, I think it is a missed optimization that gcc is creating a
> > >>>>> libcall for x86_64-v2 and forward.
> > >>>>
> > >>>> Can you open a GCC bug and CC me?
> > >>>>
> > >>>
> > >>> I opened:
> > >>>
> > >>> https://gcc.gnu.org/bugzilla/show_bug.cgi?id=121861
> > >>>
> > >>
> > >> Thanks, using inline resulted in a simpler patch. I will send a new version.
> > >
> > > Will adding -fno-trapping-math with -mavx/-msse4.1 break these 2 functions?
> > > If not, we can add -fno-trapping-math with -mavx/-msse4.1.
> > >
> >
> > The -fno-trapping-math forces the roundss/vroundss; however I am not if this is the
>
> But roundss/vroundss are exactly what compiler generates for
> trunc/truncf with -O2.
>
> > best approach since it seems to be a unrelated flag.
>
My GCC patch:
https://patchwork.sourceware.org/project/gcc/list/?series=51782
I am testing this:
diff --git a/sysdeps/x86_64/fpu/multiarch/Makefile
b/sysdeps/x86_64/fpu/multiarch/Makefile
index 708b14297f..0b2caac7a0 100644
--- a/sysdeps/x86_64/fpu/multiarch/Makefile
+++ b/sysdeps/x86_64/fpu/multiarch/Makefile
@@ -158,11 +158,14 @@ CFLAGS-s_sin-avx.c = -msse2avx -DSSE2AVX
CFLAGS-s_tan-avx.c = -msse2avx -DSSE2AVX
CFLAGS-s_sincos-avx.c = -msse2avx -DSSE2AVX
-CFLAGS-s_modf-sse4_1.c = -msse4.1 -fno-builtin-modff32x -fno-builtin-modff64
-CFLAGS-s_modff-sse4_1.c = -msse4.1 -fno-builtin-modff32
+CFLAGS-s_modf-sse4_1.c = -msse4.1 -fno-builtin-modff32x \
+ -fno-builtin-modff64 -fno-trapping-math
+CFLAGS-s_modff-sse4_1.c = -msse4.1 -fno-builtin-modff32 \
+ -fno-trapping-math
-CFLAGS-s_modf-avx.c = -mavx -fno-builtin-modff32x -fno-builtin-modff64
-CFLAGS-s_modff-avx.c = -mavx -fno-builtin-modff32
+CFLAGS-s_modf-avx.c = -mavx -fno-builtin-modff32x -fno-builtin-modff64 \
+ -fno-trapping-math
+CFLAGS-s_modff-avx.c = -mavx -fno-builtin-modff32 -fno-trapping-math
endif
ifeq ($(subdir),mathvec)
--
H.J.
More information about the Libc-alpha
mailing list