[PATCH] malloc: Optimize the madvise behaviour on the main heap

Dev Jain dev.jain@arm.com
Tue Nov 25 18:26:09 GMT 2025


On 25/11/25 11:06 pm, Adhemerval Zanella Netto wrote:
>
> On 25/11/25 14:25, Dev Jain wrote:
>> On 25/11/25 8:20 pm, Adhemerval Zanella Netto wrote:
>>> On 22/11/25 07:58, Dev Jain wrote:
>>>> On 20/11/25 10:51 pm, Adhemerval Zanella Netto wrote:
>>>>> On 17/11/25 02:40, Dev Jain wrote:
>>>>>> On 14/11/25 9:57 pm, Adhemerval Zanella Netto wrote:
>>>>>>> On 09/11/25 04:29, Dev Jain wrote:
>>>>>>>> Linux handles virtual memory in Virtual Memory Areas (VMAs). The
>>>>>>>> madvise(MADV_HUGEPAGE) call works on a VMA granularity, which sets the
>>>>>>>> VM_HUGEPAGE flag on the VMA. Therefore, if we can guarantee that a VMA
>>>>>>>> has been marked with VM_HUGEPAGE already, then we do not need to call
>>>>>>>> madvise() on that VMA again.
>>>>>>>>
>>>>>>>> For mp_.thp_pagesize != 0, currently we align the new brk to the thp size.
>>>>>>>> This means that after the first extension, all such brk extensions are
>>>>>>>> guaranteed to produce an extension size >= thp size: madvise_thp() will
>>>>>>>> invoke the madvise() syscall only if size >= thp size, and the other
>>>>>>>> condition is related to the sysctl setting, wherein mp_.thp_mode will be
>>>>>>>> same throughout the lifetime of the process. Therefore, currently we invoke
>>>>>>>> the madvise() syscall on the heap on each extension, which is unnecessary.
>>>>>>>>
>>>>>>>> First, pass the total heap size, instead of the extension size, to
>>>>>>>> madvise_thp: Linux does not care about the size passed, in case the
>>>>>>>> madvise() syscall is invoked with MADV_HUGEPAGE flag, because the flag
>>>>>>>> will be set on the entire VMA, no matter for what portion of the VMA
>>>>>>>> the syscall is invoked.
>>>>>>>>
>>>>>>>> This enables us to do the following: if the old heap size >= thp size,
>>>>>>>> we can guarantee that madvise() was invoked on one of the previous
>>>>>>>> extensions of the heap. So, avoid making the syscall in this case.
>>>>>>>>
>>>>>>>> The tricky part is computing the size of the heap, i.e the current program
>>>>>>>> break minus the initial program break. In case the first ever attempt at
>>>>>>>> extending the break fails, mp_.sbrk_base will be set to an mmapped address.
>>>>>>>> Therefore, we need some other way of remembering the initial location
>>>>>>>> of the program break. We can reuse some code for this: MORECORE (0), when
>>>>>>>> invoked for the first time ever, will give us the initial program break.
>>>>>>>> ---
>>>>>>>> The patch applies on 259adb087dd9. Built on Aarch64, all malloc tests pass.
>>>>>>> I am not sure if the calculation of
>>>>>>>
>>>>>>>>     malloc/malloc.c | 23 +++++++++++++++++++++--
>>>>>>>>     1 file changed, 21 insertions(+), 2 deletions(-)
>>>>>>>>
>>>>>>>> diff --git a/malloc/malloc.c b/malloc/malloc.c
>>>>>>>> index 0b21bdf1bd..277fb9e9ec 100644
>>>>>>>> --- a/malloc/malloc.c
>>>>>>>> +++ b/malloc/malloc.c
>>>>>>>> @@ -1938,6 +1938,12 @@ struct malloc_par
>>>>>>>>       /* First address handed out by MORECORE/sbrk.  */
>>>>>>>>       char *sbrk_base;
>>>>>>>>     +  /* The initial location of program break. This will most likely be equal
>>>>>>>> +     to sbrk_base; in case the first ever extension attempt of brk fails,
>>>>>>>> +     sbrk_base will point to an mmapped address (see sysmalloc_mmap_fallback),
>>>>>>>> +     in which case these two values will not be equal.  */
>>>>>>>> +  char *init_sbrk_base;
>>>>>>>> +
>>>>>>>>     #if USE_TCACHE
>>>>>>>>       /* Maximum number of small buckets to use.  */
>>>>>>>>       size_t tcache_small_bins;
>>>>>>>> @@ -2667,6 +2673,9 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>>>           if (__glibc_unlikely (mp_.thp_pagesize != 0))
>>>>>>>>         {
>>>>>>>>           uintptr_t lastbrk = (uintptr_t) MORECORE (0);
>>>>>>>> +      if (mp_.init_sbrk_base == NULL)
>>>>>>>> +        mp_.init_sbrk_base = (char *) lastbrk;
>>>>>>>> +
>>>>>>>>           uintptr_t top = ALIGN_UP (lastbrk + size, mp_.thp_pagesize);
>>>>>>>>           size = top - lastbrk;
>>>>>>>>         }
>>>>>>>> @@ -2682,8 +2691,18 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>>>           if ((ssize_t) size > 0)
>>>>>>>>             {
>>>>>>>>               brk = (char *) (MORECORE ((long) size));
>>>>>>>> -      if (brk != (char *) (MORECORE_FAILURE))
>>>>>>>> -        madvise_thp (brk, size);
>>>>>>>> +      if (brk != (char *) (MORECORE_FAILURE)) {
>>>>>>>> +        size_t old_size = (size_t) (brk - mp_.init_sbrk_base);
>>>>>>>> +
>>>>>>>> +        /*
>>>>>>>> +          If heap already marked with MADV_HUGEPAGE, skip madvise(). Note
>>>>>>>> +          that, we don't need to check mp_.init_sbrk_base != NULL; if it
>>>>>>>> +          is NULL, it implies that mp_.thp_pagesize == 0, in which case
>>>>>>>> +          madvise_thp() will not invoke madvise().
>>>>>>>> +         */
>>>>>>>> +        if (old_size < mp_.thp_pagesize)
>>>>>>>> +        madvise_thp (brk, old_size + size);
>>>>>>> I am not sure if this calculation is fully correct, with a simple testcase:
>>>>>>>
>>>>>>> $ cat t.c
>>>>>>> #include <stdlib.h>
>>>>>>> #include <pthread.h>
>>>>>>>
>>>>>>> static void *tf (void* arg)
>>>>>>> {
>>>>>>>      for (int i = 0; i < 1024; i++)
>>>>>>>        malloc (4096);
>>>>>>>      return NULL;
>>>>>>> }
>>>>>>>
>>>>>>> int main ()
>>>>>>> {
>>>>>>>      for (int i = 0; i < 1024; i++)
>>>>>>>        malloc (4096);
>>>>>>>
>>>>>>>      pthread_t t;
>>>>>>>      pthread_create (&t, NULL, tf, NULL);
>>>>>>>      pthread_join (t, NULL);
>>>>>>> }
>>>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 elf/ld.so --library-path . ./t
>>>>>>> madvise(0xaaaade400000, 3354624, MADV_HUGEPAGE) = -1 ENOMEM (Cannot allocate memory)
>>>>>>> madvise(0xf7fa28ed0000, 65536, 0x66 /* MADV_??? */) = 0
>>>>>>> strace: Process 313335 attached
>>>>>>> [pid 313335] madvise(0xf7fa24000000, 2097152, MADV_HUGEPAGE) = 0
>>>>>>> [pid 313335] madvise(0xf7fa28ed0000, 8314880, MADV_DONTNEED) = 0
>>>>>>> [pid 313335] +++ exited with 0 +++
>>>>>>> +++ exited with 0 +++
>>>>>>>
>>>>>>> Where without this patch:
>>>>>>>
>>>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 ./t
>>>>>>> madvise(0xc3da42fb9000, 2097152, MADV_HUGEPAGE) = 0
>>>>>>> madvise(0xc3da43200000, 2097152, MADV_HUGEPAGE) = 0
>>>>>>> strace: Process 313347 attached
>>>>>>> [pid 313347] madvise(0xfa19510a0000, 8314880, MADV_DONTNEED) = 0
>>>>>>> [pid 313347] +++ exited with 0 +++
>>>>>>> +++ exited with 0 +++
>>>>>> Thanks for testing! Although without the patch, I get this:
>>>>>>
>>>>>> madvise(0xfffff8000000, 2097152, MADV_HUGEPAGE) = 0
>>>>>> madvise(0xfffff8000000, 4194304, MADV_HUGEPAGE) = 0
>>>>>> madvise(0xfffff8000000, 6291456, MADV_HUGEPAGE) = 0
>>>>>> madvise(0xfffff7400000, 65536, 0x66 /* MADV_??? */) = -1 EINVAL (Invalid argument)
>>>>>> strace: Process 3162240 attached
>>>>>> [pid 3162240] madvise(0xfffff0000000, 2097152, MADV_HUGEPAGE) = 0
>>>>>> [pid 3162240] madvise(0xfffff7400000, 8314880, MADV_DONTNEED) = 0
>>>>>> [pid 3162240] +++ exited with 0 +++
>>>>>> +++ exited with 0 +++
>>>>>>
>>>>>> I don't know where the EINVAL is coming from, that shouldn't be happening, because we
>>>>>> have hardcoded MADV_HUGEPAGE everywhere. I played around and I think that when madvise() is not invoked
>>>>>> ever, then strace shows it returning EINVAL. I am not familiar with how strace actually
>>>>>> works and how to interpret the output.
>>>>> The EINVAL most likely come from the size not being multiple of the large
>>>>> page size, but I am not sure.  I could not understand why the flags are being
>>>>> shown as invalid (0x66), since I step on every madvise syscall trap and
>>>>> I did not see this valug.
>>>> For MADV_HUGEPAGE, the size does not matter, if you just do a madvise(brk, 1, MADV_HUGEPAGE)
>>>> it will work.
>>> The EINVAL from the 0x66 flags is actually the MADV_GUARD_INSTALL,
>>> and neither the kernel nor strace has support for it.
>>>
>>>>>> But, there is an obvious error in this patch. I shouldn't be making the madvise() call
>>>>>> on brk pointer, but on init_sbrk_base. Calling it on brk will cross the heap VMA, and
>>>>>> the target range may include a hole. I checked the kernel source and it returns -ENOMEM
>>>>>> if it encounters a gap.
>>>>>>
>>>>>> But I have a question. In madvise_thp, we align the pointer down if it is not aligned to
>>>>>> pagesize. If before that pointer, there is a gap, then madvise() will return -ENOMEM,
>>>>>> although, the madvise() operation will continue on the VMA which we have specified.
>>>>>> So, the operation will do what we want it to do, but return -ENOMEM. Do you think
>>>>>> this is something which needs to be fixed? The fix is simple: Align the pointer up
>>>>>> instead of down, and subtract from size instead of add. Then we can check again
>>>>>> whether size >= thp_pagesize, and continue with the syscall.
>>>>> Are you referring to the case where the initial heap is not largepage aligned, then
>>>>> the align down to a possible invalid address? If kernel returns ENOMEM in such case
>>>>> (I have not checked yet), your suggestion might work although I think it will not
>>>>> cover the initial page.
>>>> I am referring to this:
>>>>
>>>> /* Linux requires the input address to be page-aligned, and unaligned
>>>> inputs happens only for initial data segment. */
>>>> if(__glibc_unlikely(!PTR_IS_ALIGNED(p, GLRO(dl_pagesize))))
>>>> {
>>>> void*q = PTR_ALIGN_DOWN(p, GLRO(dl_pagesize));
>>>> size += PTR_DIFF(p, q);
>>>> p = q;
>>>> }
>>>>
>>>> In this case we should be aligning the pointer up and subtracting from size.
>>>> I'll send a patch to fix this - although it does not really matter, but I think
>>>> we should do this cleanly and not suppress the -ENOMEM even when the MADV_HUGEPAGE
>>>> operation will work.
>>> I think the ENOMEM from MADV_HUGEPAGE come from trying to madvise on unmapped regions.
>>>   From mm/madvise.c:
>>>
>>> 1609 /*
>>> 1610  * Walk the vmas in range [start,end), and call the madvise_vma_behavior
>>> 1611  * function on each one.  The function will get start and end parameters that
>>> 1612  * cover the overlap between the current vma and the original range.  Any
>>> 1613  * unmapped regions in the original range will result in this function returning
>>> 1614  * -ENOMEM while still calling the madvise_vma_behavior function on all of the
>>> 1615  * existing vmas in the range.  Must be called with the mmap_lock held for
>>> 1616  * reading or writing.
>>> 1617  */
>>> 1618 static
>>> 1619 int madvise_walk_vmas(struct madvise_behavior *madv_behavior)
>>> 1620 {
>>> [...]
>>> 1644         for (;;) {
>>> 1645                 /* Still start < end. */
>>> 1646                 if (!vma)
>>> 1647                         return -ENOMEM;
>>> 1648
>>> 1649                 /* Here start < (last_end|vma->vm_end). */
>>> 1650                 if (range->start < vma->vm_start) {
>>> 1651                         /*
>>> 1652                          * This indicates a gap between VMAs in the input
>>> 1653                          * range. This does not cause the operation to abort,
>>> 1654                          * rather we simply return -ENOMEM to indicate that this
>>> 1655                          * has happened, but carry on.
>>> 1656                          */
>>> 1657                         unmapped_error = -ENOMEM;
>>> 1658                         range->start = vma->vm_start;
>>> 1659                         if (range->start >= last_end)
>>> 1660                                 break;
>>> 1661                 }
>>>
>>> With the patch applied, on first madvise:
>>>
>>> (gdb) bt
>>> #0  0x0000fffff7e9774c in __GI_madvise () at ../sysdeps/unix/syscall-template.S:120
>>> #1  0x0000fffff7e3b11c in madvise_thp (p=<optimized out>, size=<optimized out>) at malloc.c:2085
>>> #2  madvise_thp (p=0xaaaaaac00000, size=<optimized out>) at malloc.c:2063
>>> #3  sysmalloc (nb=nb@entry=4112, av=av@entry=0xfffff7f70a50 <main_arena>) at malloc.c:2704
>>> #4  0x0000fffff7e3c160 in _int_malloc (av=av@entry=0xfffff7f70a50 <main_arena>, bytes=bytes@entry=4096) at malloc.c:4664
>>> #5  0x0000fffff7e3c358 in __libc_malloc2 (bytes=4096) at malloc.c:3479
>>> #6  0x0000fffff7f80948 in ?? ()
>>> #7  0x0000fffff7f80918 in ?? ()
>>> Backtrace stopped: not enough registers or memory available to unwind further
>>> (gdb) i r x0 x1 x2
>>> x0             0xaaaaaac00000      187649985871872
>>> x1             0x355000            3493888
>>> x2             0xe                 14
>>> (gdb) info proc mappings
>>> process 3827900
>>> Mapped address spaces:
>>>
>>>             Start Addr           End Addr       Size     Offset  Perms  objfile
>>>         0xaaaaaaaab000     0xaaaaaae00000   0x355000        0x0  rw-p   [heap]
>>>         0xfffff7d90000     0xfffff7f52000   0x1c2000        0x0  r-xp   [...]/libc.so
>>>         0xfffff7f52000     0xfffff7f6d000    0x1b000   0x1c2000  ---p   [...]/libc.so
>>>         0xfffff7f6d000     0xfffff7f70000     0x3000   0x1cd000  r--p   [...]/libc.so
>>>         0xfffff7f70000     0xfffff7f72000     0x2000   0x1d0000  rw-p   [...]/libc.so
>>>         0xfffff7f72000     0xfffff7f7e000     0xc000        0x0  rw-p
>>>         0xfffff7f80000     0xfffff7f81000     0x1000        0x0  r-xp   [...]/t
>>>         0xfffff7f81000     0xfffff7f9f000    0x1e000     0x1000  ---p   [...]/t
>>>         0xfffff7f9f000     0xfffff7fa0000     0x1000     0xf000  r--p   [...]/t
>>>         0xfffff7fa0000     0xfffff7fa1000     0x1000    0x10000  rw-p   [...]/t
>>>         0xfffff7fb0000     0xfffff7fd8000    0x28000        0x0  r-xp   [...]/ld.so
>>>         0xfffff7fee000     0xfffff7ff0000     0x2000    0x2e000  r--p   [...]/ld.so
>>>         0xfffff7ff0000     0xfffff7ff1000     0x1000    0x30000  rw-p   [...]/ld.so
>>>         0xfffff7ff1000     0xfffff7ff2000     0x1000        0x0  rw-p
>>>         0xfffff7ff8000     0xfffff7ffa000     0x2000        0x0  rw-p
>>>         0xfffff7ffa000     0xfffff7ffe000     0x4000        0x0  r--p   [vvar]
>>>         0xfffff7ffe000     0xfffff8000000     0x2000        0x0  r-xp   [vdso]
>>>         0xfffffffde000    0x1000000000000    0x22000        0x0  rw-p   [stack]
>>>
>>> And the patch tries to madvise the region of [0xaaaaaac00000,0xaaaaaaf55000], which is
>>> larger than the allocated heap.  I am not sure if kernel enables khugepage/THP page
>>> by page, meaning that it will have partial enablement if the memory region is is outside
>>> of the mapped range.
>> Missed to answer your doubt: as explained in the patch description, kernel *does not* enable
>> THP page by page. The granularity is VMA, not page. If the VMA which you are targetting is
>> the range R1 and you do a madvise on range R2, all you need to ensure is that R1's intersection
>> with R2 contains at least one page. Of course, to preserve sanity, if you have a VMA in the range
>> [4, 9], you will never write code doing a madvise call on the range [9, 20]: it does the job, but
>> looks insane, may return ENOMEM (which we ultimately ignore) if there is a gap after 9th page, and
>> may end up having the effect of madvise on neighbouring VMAs.
>>
> Right, so any intersection of the mapped VMA and the madvise VMA will be considered
> to be backed by MADV_HUGEPAGE, even if kernel returns ENOMEM?
>
> For instance, let's we have:
>
> process mapping:  |map1|gap|map2|___|
> madvise           |__|madvise-------|
>               
> Meaning that map1 has an inserction range, map2 is fully covered (and
> assuming the intersection is larger than the page size), and gap is an
> unmmaped range.
>
> Would kernel bail on 'gap' or will it continue on all ranges (including map2)?

I don't know this for a fact since I never had to look into this, but I am guessing it will
continue on all ranges. Looking at the kernel code you have pasted, nothing changes the
variable "last_end". We break from the unbounded for loop only if range->start cross last_end
(in which case we have completely covered our request range) or, we do madvise_vma_behaviour() on
the last possible range and then bail out because range->end crosses last_end. There is also a
return -ENOMEM on !vma, but that will only happen when the range->start we pass to find_vma_prev()
is an address beyond the last VMA of this process.

You will never get an error from madvise_vma_behaviour() for MADV_HUGEPAGE - that code path always
returns 0 - it is literally just setting a flag on the VMA. Although, it is possible for that code
path to fail to set the flag, in case (extreme implementation detail) we fail to allocate an internal
structure which is less than 64 bytes. In other words, the probability of madvise(range, MADV_HUGEPAGE)
failing is close to none (and probably which is why it is not worth to change kernel code to not suppress
this memory allocation failure, and also enlighten malloc code to check the return value).



More information about the Libc-alpha mailing list