[PATCH] malloc: Optimize the madvise behaviour on the main heap

Dev Jain dev.jain@arm.com
Tue Nov 25 17:25:11 GMT 2025


On 25/11/25 8:20 pm, Adhemerval Zanella Netto wrote:
>
> On 22/11/25 07:58, Dev Jain wrote:
>> On 20/11/25 10:51 pm, Adhemerval Zanella Netto wrote:
>>> On 17/11/25 02:40, Dev Jain wrote:
>>>> On 14/11/25 9:57 pm, Adhemerval Zanella Netto wrote:
>>>>> On 09/11/25 04:29, Dev Jain wrote:
>>>>>> Linux handles virtual memory in Virtual Memory Areas (VMAs). The
>>>>>> madvise(MADV_HUGEPAGE) call works on a VMA granularity, which sets the
>>>>>> VM_HUGEPAGE flag on the VMA. Therefore, if we can guarantee that a VMA
>>>>>> has been marked with VM_HUGEPAGE already, then we do not need to call
>>>>>> madvise() on that VMA again.
>>>>>>
>>>>>> For mp_.thp_pagesize != 0, currently we align the new brk to the thp size.
>>>>>> This means that after the first extension, all such brk extensions are
>>>>>> guaranteed to produce an extension size >= thp size: madvise_thp() will
>>>>>> invoke the madvise() syscall only if size >= thp size, and the other
>>>>>> condition is related to the sysctl setting, wherein mp_.thp_mode will be
>>>>>> same throughout the lifetime of the process. Therefore, currently we invoke
>>>>>> the madvise() syscall on the heap on each extension, which is unnecessary.
>>>>>>
>>>>>> First, pass the total heap size, instead of the extension size, to
>>>>>> madvise_thp: Linux does not care about the size passed, in case the
>>>>>> madvise() syscall is invoked with MADV_HUGEPAGE flag, because the flag
>>>>>> will be set on the entire VMA, no matter for what portion of the VMA
>>>>>> the syscall is invoked.
>>>>>>
>>>>>> This enables us to do the following: if the old heap size >= thp size,
>>>>>> we can guarantee that madvise() was invoked on one of the previous
>>>>>> extensions of the heap. So, avoid making the syscall in this case.
>>>>>>
>>>>>> The tricky part is computing the size of the heap, i.e the current program
>>>>>> break minus the initial program break. In case the first ever attempt at
>>>>>> extending the break fails, mp_.sbrk_base will be set to an mmapped address.
>>>>>> Therefore, we need some other way of remembering the initial location
>>>>>> of the program break. We can reuse some code for this: MORECORE (0), when
>>>>>> invoked for the first time ever, will give us the initial program break.
>>>>>> ---
>>>>>> The patch applies on 259adb087dd9. Built on Aarch64, all malloc tests pass.
>>>>> I am not sure if the calculation of
>>>>>
>>>>>>    malloc/malloc.c | 23 +++++++++++++++++++++--
>>>>>>    1 file changed, 21 insertions(+), 2 deletions(-)
>>>>>>
>>>>>> diff --git a/malloc/malloc.c b/malloc/malloc.c
>>>>>> index 0b21bdf1bd..277fb9e9ec 100644
>>>>>> --- a/malloc/malloc.c
>>>>>> +++ b/malloc/malloc.c
>>>>>> @@ -1938,6 +1938,12 @@ struct malloc_par
>>>>>>      /* First address handed out by MORECORE/sbrk.  */
>>>>>>      char *sbrk_base;
>>>>>>    +  /* The initial location of program break. This will most likely be equal
>>>>>> +     to sbrk_base; in case the first ever extension attempt of brk fails,
>>>>>> +     sbrk_base will point to an mmapped address (see sysmalloc_mmap_fallback),
>>>>>> +     in which case these two values will not be equal.  */
>>>>>> +  char *init_sbrk_base;
>>>>>> +
>>>>>>    #if USE_TCACHE
>>>>>>      /* Maximum number of small buckets to use.  */
>>>>>>      size_t tcache_small_bins;
>>>>>> @@ -2667,6 +2673,9 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>          if (__glibc_unlikely (mp_.thp_pagesize != 0))
>>>>>>        {
>>>>>>          uintptr_t lastbrk = (uintptr_t) MORECORE (0);
>>>>>> +      if (mp_.init_sbrk_base == NULL)
>>>>>> +        mp_.init_sbrk_base = (char *) lastbrk;
>>>>>> +
>>>>>>          uintptr_t top = ALIGN_UP (lastbrk + size, mp_.thp_pagesize);
>>>>>>          size = top - lastbrk;
>>>>>>        }
>>>>>> @@ -2682,8 +2691,18 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>          if ((ssize_t) size > 0)
>>>>>>            {
>>>>>>              brk = (char *) (MORECORE ((long) size));
>>>>>> -      if (brk != (char *) (MORECORE_FAILURE))
>>>>>> -        madvise_thp (brk, size);
>>>>>> +      if (brk != (char *) (MORECORE_FAILURE)) {
>>>>>> +        size_t old_size = (size_t) (brk - mp_.init_sbrk_base);
>>>>>> +
>>>>>> +        /*
>>>>>> +          If heap already marked with MADV_HUGEPAGE, skip madvise(). Note
>>>>>> +          that, we don't need to check mp_.init_sbrk_base != NULL; if it
>>>>>> +          is NULL, it implies that mp_.thp_pagesize == 0, in which case
>>>>>> +          madvise_thp() will not invoke madvise().
>>>>>> +         */
>>>>>> +        if (old_size < mp_.thp_pagesize)
>>>>>> +        madvise_thp (brk, old_size + size);
>>>>> I am not sure if this calculation is fully correct, with a simple testcase:
>>>>>
>>>>> $ cat t.c
>>>>> #include <stdlib.h>
>>>>> #include <pthread.h>
>>>>>
>>>>> static void *tf (void* arg)
>>>>> {
>>>>>     for (int i = 0; i < 1024; i++)
>>>>>       malloc (4096);
>>>>>     return NULL;
>>>>> }
>>>>>
>>>>> int main ()
>>>>> {
>>>>>     for (int i = 0; i < 1024; i++)
>>>>>       malloc (4096);
>>>>>
>>>>>     pthread_t t;
>>>>>     pthread_create (&t, NULL, tf, NULL);
>>>>>     pthread_join (t, NULL);
>>>>> }
>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 elf/ld.so --library-path . ./t
>>>>> madvise(0xaaaade400000, 3354624, MADV_HUGEPAGE) = -1 ENOMEM (Cannot allocate memory)
>>>>> madvise(0xf7fa28ed0000, 65536, 0x66 /* MADV_??? */) = 0
>>>>> strace: Process 313335 attached
>>>>> [pid 313335] madvise(0xf7fa24000000, 2097152, MADV_HUGEPAGE) = 0
>>>>> [pid 313335] madvise(0xf7fa28ed0000, 8314880, MADV_DONTNEED) = 0
>>>>> [pid 313335] +++ exited with 0 +++
>>>>> +++ exited with 0 +++
>>>>>
>>>>> Where without this patch:
>>>>>
>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 ./t
>>>>> madvise(0xc3da42fb9000, 2097152, MADV_HUGEPAGE) = 0
>>>>> madvise(0xc3da43200000, 2097152, MADV_HUGEPAGE) = 0
>>>>> strace: Process 313347 attached
>>>>> [pid 313347] madvise(0xfa19510a0000, 8314880, MADV_DONTNEED) = 0
>>>>> [pid 313347] +++ exited with 0 +++
>>>>> +++ exited with 0 +++
>>>> Thanks for testing! Although without the patch, I get this:
>>>>
>>>> madvise(0xfffff8000000, 2097152, MADV_HUGEPAGE) = 0
>>>> madvise(0xfffff8000000, 4194304, MADV_HUGEPAGE) = 0
>>>> madvise(0xfffff8000000, 6291456, MADV_HUGEPAGE) = 0
>>>> madvise(0xfffff7400000, 65536, 0x66 /* MADV_??? */) = -1 EINVAL (Invalid argument)
>>>> strace: Process 3162240 attached
>>>> [pid 3162240] madvise(0xfffff0000000, 2097152, MADV_HUGEPAGE) = 0
>>>> [pid 3162240] madvise(0xfffff7400000, 8314880, MADV_DONTNEED) = 0
>>>> [pid 3162240] +++ exited with 0 +++
>>>> +++ exited with 0 +++
>>>>
>>>> I don't know where the EINVAL is coming from, that shouldn't be happening, because we
>>>> have hardcoded MADV_HUGEPAGE everywhere. I played around and I think that when madvise() is not invoked
>>>> ever, then strace shows it returning EINVAL. I am not familiar with how strace actually
>>>> works and how to interpret the output.
>>> The EINVAL most likely come from the size not being multiple of the large
>>> page size, but I am not sure.  I could not understand why the flags are being
>>> shown as invalid (0x66), since I step on every madvise syscall trap and
>>> I did not see this valug.
>> For MADV_HUGEPAGE, the size does not matter, if you just do a madvise(brk, 1, MADV_HUGEPAGE)
>> it will work.
> The EINVAL from the 0x66 flags is actually the MADV_GUARD_INSTALL,
> and neither the kernel nor strace has support for it.
>
>>>> But, there is an obvious error in this patch. I shouldn't be making the madvise() call
>>>> on brk pointer, but on init_sbrk_base. Calling it on brk will cross the heap VMA, and
>>>> the target range may include a hole. I checked the kernel source and it returns -ENOMEM
>>>> if it encounters a gap.
>>>>
>>>> But I have a question. In madvise_thp, we align the pointer down if it is not aligned to
>>>> pagesize. If before that pointer, there is a gap, then madvise() will return -ENOMEM,
>>>> although, the madvise() operation will continue on the VMA which we have specified.
>>>> So, the operation will do what we want it to do, but return -ENOMEM. Do you think
>>>> this is something which needs to be fixed? The fix is simple: Align the pointer up
>>>> instead of down, and subtract from size instead of add. Then we can check again
>>>> whether size >= thp_pagesize, and continue with the syscall.
>>> Are you referring to the case where the initial heap is not largepage aligned, then
>>> the align down to a possible invalid address? If kernel returns ENOMEM in such case
>>> (I have not checked yet), your suggestion might work although I think it will not
>>> cover the initial page.
>> I am referring to this:
>>
>> /* Linux requires the input address to be page-aligned, and unaligned
>> inputs happens only for initial data segment. */
>> if(__glibc_unlikely(!PTR_IS_ALIGNED(p, GLRO(dl_pagesize))))
>> {
>> void*q = PTR_ALIGN_DOWN(p, GLRO(dl_pagesize));
>> size += PTR_DIFF(p, q);
>> p = q;
>> }
>>
>> In this case we should be aligning the pointer up and subtracting from size.
>> I'll send a patch to fix this - although it does not really matter, but I think
>> we should do this cleanly and not suppress the -ENOMEM even when the MADV_HUGEPAGE
>> operation will work.
> I think the ENOMEM from MADV_HUGEPAGE come from trying to madvise on unmapped regions.
>  From mm/madvise.c:
>
> 1609 /*
> 1610  * Walk the vmas in range [start,end), and call the madvise_vma_behavior
> 1611  * function on each one.  The function will get start and end parameters that
> 1612  * cover the overlap between the current vma and the original range.  Any
> 1613  * unmapped regions in the original range will result in this function returning
> 1614  * -ENOMEM while still calling the madvise_vma_behavior function on all of the
> 1615  * existing vmas in the range.  Must be called with the mmap_lock held for
> 1616  * reading or writing.
> 1617  */
> 1618 static
> 1619 int madvise_walk_vmas(struct madvise_behavior *madv_behavior)
> 1620 {
> [...]
> 1644         for (;;) {
> 1645                 /* Still start < end. */
> 1646                 if (!vma)
> 1647                         return -ENOMEM;
> 1648
> 1649                 /* Here start < (last_end|vma->vm_end). */
> 1650                 if (range->start < vma->vm_start) {
> 1651                         /*
> 1652                          * This indicates a gap between VMAs in the input
> 1653                          * range. This does not cause the operation to abort,
> 1654                          * rather we simply return -ENOMEM to indicate that this
> 1655                          * has happened, but carry on.
> 1656                          */
> 1657                         unmapped_error = -ENOMEM;
> 1658                         range->start = vma->vm_start;
> 1659                         if (range->start >= last_end)
> 1660                                 break;
> 1661                 }
>
> With the patch applied, on first madvise:
>
> (gdb) bt
> #0  0x0000fffff7e9774c in __GI_madvise () at ../sysdeps/unix/syscall-template.S:120
> #1  0x0000fffff7e3b11c in madvise_thp (p=<optimized out>, size=<optimized out>) at malloc.c:2085
> #2  madvise_thp (p=0xaaaaaac00000, size=<optimized out>) at malloc.c:2063
> #3  sysmalloc (nb=nb@entry=4112, av=av@entry=0xfffff7f70a50 <main_arena>) at malloc.c:2704
> #4  0x0000fffff7e3c160 in _int_malloc (av=av@entry=0xfffff7f70a50 <main_arena>, bytes=bytes@entry=4096) at malloc.c:4664
> #5  0x0000fffff7e3c358 in __libc_malloc2 (bytes=4096) at malloc.c:3479
> #6  0x0000fffff7f80948 in ?? ()
> #7  0x0000fffff7f80918 in ?? ()
> Backtrace stopped: not enough registers or memory available to unwind further
> (gdb) i r x0 x1 x2
> x0             0xaaaaaac00000      187649985871872
> x1             0x355000            3493888
> x2             0xe                 14
> (gdb) info proc mappings
> process 3827900
> Mapped address spaces:
>
>            Start Addr           End Addr       Size     Offset  Perms  objfile
>        0xaaaaaaaab000     0xaaaaaae00000   0x355000        0x0  rw-p   [heap]
>        0xfffff7d90000     0xfffff7f52000   0x1c2000        0x0  r-xp   [...]/libc.so
>        0xfffff7f52000     0xfffff7f6d000    0x1b000   0x1c2000  ---p   [...]/libc.so
>        0xfffff7f6d000     0xfffff7f70000     0x3000   0x1cd000  r--p   [...]/libc.so
>        0xfffff7f70000     0xfffff7f72000     0x2000   0x1d0000  rw-p   [...]/libc.so
>        0xfffff7f72000     0xfffff7f7e000     0xc000        0x0  rw-p
>        0xfffff7f80000     0xfffff7f81000     0x1000        0x0  r-xp   [...]/t
>        0xfffff7f81000     0xfffff7f9f000    0x1e000     0x1000  ---p   [...]/t
>        0xfffff7f9f000     0xfffff7fa0000     0x1000     0xf000  r--p   [...]/t
>        0xfffff7fa0000     0xfffff7fa1000     0x1000    0x10000  rw-p   [...]/t
>        0xfffff7fb0000     0xfffff7fd8000    0x28000        0x0  r-xp   [...]/ld.so
>        0xfffff7fee000     0xfffff7ff0000     0x2000    0x2e000  r--p   [...]/ld.so
>        0xfffff7ff0000     0xfffff7ff1000     0x1000    0x30000  rw-p   [...]/ld.so
>        0xfffff7ff1000     0xfffff7ff2000     0x1000        0x0  rw-p
>        0xfffff7ff8000     0xfffff7ffa000     0x2000        0x0  rw-p
>        0xfffff7ffa000     0xfffff7ffe000     0x4000        0x0  r--p   [vvar]
>        0xfffff7ffe000     0xfffff8000000     0x2000        0x0  r-xp   [vdso]
>        0xfffffffde000    0x1000000000000    0x22000        0x0  rw-p   [stack]
>
> And the patch tries to madvise the region of [0xaaaaaac00000,0xaaaaaaf55000], which is
> larger than the allocated heap.  I am not sure if kernel enables khugepage/THP page
> by page, meaning that it will have partial enablement if the memory region is is outside
> of the mapped range.

Missed to answer your doubt: as explained in the patch description, kernel *does not* enable
THP page by page. The granularity is VMA, not page. If the VMA which you are targetting is
the range R1 and you do a madvise on range R2, all you need to ensure is that R1's intersection
with R2 contains at least one page. Of course, to preserve sanity, if you have a VMA in the range
[4, 9], you will never write code doing a madvise call on the range [9, 20]: it does the job, but
looks insane, may return ENOMEM (which we ultimately ignore) if there is a gap after 9th page, and
may end up having the effect of madvise on neighbouring VMAs.

>
>> After fixing that, there is some error in this patch, I got a regression on SPEC with THP=madvise
>> and glibc.malloc.hugetlb=1. I suspect it has something to do with the madvise_thp (snd_brk, correction)
>> down in the code. As in, suppose the first extension of the heap is less than thp_pagesize, correction
>> is also less than thp_pagesize, but the sum of them is >= thp_pagesize. This will result in VM_HUGEPAGE
>> never being set in either of the madvise_thp() calls, but the resultant heap size will be >= thp_pagesize,
>> and my patch will avoid making the syscall, so VM_HUGEPAGE never gets set on the main heap.
> Is it a performance regression or a correctness one? And I am not sure either what might be
> causing this, I will need to dig into to understand it better.
>


More information about the Libc-alpha mailing list