negative performance speedup with -fno-plt
Adhemerval Zanella Netto
adhemerval.zanella@linaro.org
Mon Sep 16 21:28:38 GMT 2024
Building glibc with -fno-plt most likely will not yield much gain,
just a handful of symbols are called as external ones (you can check
with the localplt.data, since they are arch-specific and has changes
over time). Currently the main ones are malloc functions to allow
malloc interposition
On 16/09/24 15:36, Farid Zakaria via Libc-help wrote:
> Thank you for the advice; I will look into that with perf if it has statistics.
>
> I noticed also that I had only built python + libpython without a PLT
> but glibc itself continued to have a PLT
> (I'm using NixOS to make changes).
>
> Looks like trying to override CFLAGS to build glibc without PLT is
> challenging; it didn't like me just setting CFLAGS :)
>
> I wonder if removing PLT from glibc will exacerbate the problem I'm
> noticing but also have me see more speedups.
> For those curious here is my python benchmark thus far:
> https://pasteboard.co/RdcGYZH4m0kN.png
>
> On Sun, Sep 15, 2024 at 11:42 PM Florian Weimer <fweimer@redhat.com> wrote:
>>
>> * Farid Zakaria via Libc-help:
>>
>>> I'm comparing it against a Python built with BIND_NOW (-z,now) to
>>> account for no-lazy binding in both. I'm a bit stumped at what could
>>> be causing some negative speedups. Anyone got leads that might cause a
>>> difference?
>>
>> Look at mispredicted indirect branches. With the PLT, all the different
>> calls to external functions share one indirect branch, but without it,
>> you get an indirect branch for each call site. It must be predicted
>> indepedently, and your CPU might not be able to track so man indirect
>> branches.
>>
>> Thanks,
>> Florian
>>
More information about the Libc-help
mailing list