[PATCH] malloc: Optimize the madvise behaviour on the main heap
Adhemerval Zanella Netto
adhemerval.zanella@linaro.org
Tue Nov 25 17:36:14 GMT 2025
On 25/11/25 14:25, Dev Jain wrote:
>
> On 25/11/25 8:20 pm, Adhemerval Zanella Netto wrote:
>>
>> On 22/11/25 07:58, Dev Jain wrote:
>>> On 20/11/25 10:51 pm, Adhemerval Zanella Netto wrote:
>>>> On 17/11/25 02:40, Dev Jain wrote:
>>>>> On 14/11/25 9:57 pm, Adhemerval Zanella Netto wrote:
>>>>>> On 09/11/25 04:29, Dev Jain wrote:
>>>>>>> Linux handles virtual memory in Virtual Memory Areas (VMAs). The
>>>>>>> madvise(MADV_HUGEPAGE) call works on a VMA granularity, which sets the
>>>>>>> VM_HUGEPAGE flag on the VMA. Therefore, if we can guarantee that a VMA
>>>>>>> has been marked with VM_HUGEPAGE already, then we do not need to call
>>>>>>> madvise() on that VMA again.
>>>>>>>
>>>>>>> For mp_.thp_pagesize != 0, currently we align the new brk to the thp size.
>>>>>>> This means that after the first extension, all such brk extensions are
>>>>>>> guaranteed to produce an extension size >= thp size: madvise_thp() will
>>>>>>> invoke the madvise() syscall only if size >= thp size, and the other
>>>>>>> condition is related to the sysctl setting, wherein mp_.thp_mode will be
>>>>>>> same throughout the lifetime of the process. Therefore, currently we invoke
>>>>>>> the madvise() syscall on the heap on each extension, which is unnecessary.
>>>>>>>
>>>>>>> First, pass the total heap size, instead of the extension size, to
>>>>>>> madvise_thp: Linux does not care about the size passed, in case the
>>>>>>> madvise() syscall is invoked with MADV_HUGEPAGE flag, because the flag
>>>>>>> will be set on the entire VMA, no matter for what portion of the VMA
>>>>>>> the syscall is invoked.
>>>>>>>
>>>>>>> This enables us to do the following: if the old heap size >= thp size,
>>>>>>> we can guarantee that madvise() was invoked on one of the previous
>>>>>>> extensions of the heap. So, avoid making the syscall in this case.
>>>>>>>
>>>>>>> The tricky part is computing the size of the heap, i.e the current program
>>>>>>> break minus the initial program break. In case the first ever attempt at
>>>>>>> extending the break fails, mp_.sbrk_base will be set to an mmapped address.
>>>>>>> Therefore, we need some other way of remembering the initial location
>>>>>>> of the program break. We can reuse some code for this: MORECORE (0), when
>>>>>>> invoked for the first time ever, will give us the initial program break.
>>>>>>> ---
>>>>>>> The patch applies on 259adb087dd9. Built on Aarch64, all malloc tests pass.
>>>>>> I am not sure if the calculation of
>>>>>>
>>>>>>> malloc/malloc.c | 23 +++++++++++++++++++++--
>>>>>>> 1 file changed, 21 insertions(+), 2 deletions(-)
>>>>>>>
>>>>>>> diff --git a/malloc/malloc.c b/malloc/malloc.c
>>>>>>> index 0b21bdf1bd..277fb9e9ec 100644
>>>>>>> --- a/malloc/malloc.c
>>>>>>> +++ b/malloc/malloc.c
>>>>>>> @@ -1938,6 +1938,12 @@ struct malloc_par
>>>>>>> /* First address handed out by MORECORE/sbrk. */
>>>>>>> char *sbrk_base;
>>>>>>> + /* The initial location of program break. This will most likely be equal
>>>>>>> + to sbrk_base; in case the first ever extension attempt of brk fails,
>>>>>>> + sbrk_base will point to an mmapped address (see sysmalloc_mmap_fallback),
>>>>>>> + in which case these two values will not be equal. */
>>>>>>> + char *init_sbrk_base;
>>>>>>> +
>>>>>>> #if USE_TCACHE
>>>>>>> /* Maximum number of small buckets to use. */
>>>>>>> size_t tcache_small_bins;
>>>>>>> @@ -2667,6 +2673,9 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>> if (__glibc_unlikely (mp_.thp_pagesize != 0))
>>>>>>> {
>>>>>>> uintptr_t lastbrk = (uintptr_t) MORECORE (0);
>>>>>>> + if (mp_.init_sbrk_base == NULL)
>>>>>>> + mp_.init_sbrk_base = (char *) lastbrk;
>>>>>>> +
>>>>>>> uintptr_t top = ALIGN_UP (lastbrk + size, mp_.thp_pagesize);
>>>>>>> size = top - lastbrk;
>>>>>>> }
>>>>>>> @@ -2682,8 +2691,18 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
>>>>>>> if ((ssize_t) size > 0)
>>>>>>> {
>>>>>>> brk = (char *) (MORECORE ((long) size));
>>>>>>> - if (brk != (char *) (MORECORE_FAILURE))
>>>>>>> - madvise_thp (brk, size);
>>>>>>> + if (brk != (char *) (MORECORE_FAILURE)) {
>>>>>>> + size_t old_size = (size_t) (brk - mp_.init_sbrk_base);
>>>>>>> +
>>>>>>> + /*
>>>>>>> + If heap already marked with MADV_HUGEPAGE, skip madvise(). Note
>>>>>>> + that, we don't need to check mp_.init_sbrk_base != NULL; if it
>>>>>>> + is NULL, it implies that mp_.thp_pagesize == 0, in which case
>>>>>>> + madvise_thp() will not invoke madvise().
>>>>>>> + */
>>>>>>> + if (old_size < mp_.thp_pagesize)
>>>>>>> + madvise_thp (brk, old_size + size);
>>>>>> I am not sure if this calculation is fully correct, with a simple testcase:
>>>>>>
>>>>>> $ cat t.c
>>>>>> #include <stdlib.h>
>>>>>> #include <pthread.h>
>>>>>>
>>>>>> static void *tf (void* arg)
>>>>>> {
>>>>>> for (int i = 0; i < 1024; i++)
>>>>>> malloc (4096);
>>>>>> return NULL;
>>>>>> }
>>>>>>
>>>>>> int main ()
>>>>>> {
>>>>>> for (int i = 0; i < 1024; i++)
>>>>>> malloc (4096);
>>>>>>
>>>>>> pthread_t t;
>>>>>> pthread_create (&t, NULL, tf, NULL);
>>>>>> pthread_join (t, NULL);
>>>>>> }
>>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 elf/ld.so --library-path . ./t
>>>>>> madvise(0xaaaade400000, 3354624, MADV_HUGEPAGE) = -1 ENOMEM (Cannot allocate memory)
>>>>>> madvise(0xf7fa28ed0000, 65536, 0x66 /* MADV_??? */) = 0
>>>>>> strace: Process 313335 attached
>>>>>> [pid 313335] madvise(0xf7fa24000000, 2097152, MADV_HUGEPAGE) = 0
>>>>>> [pid 313335] madvise(0xf7fa28ed0000, 8314880, MADV_DONTNEED) = 0
>>>>>> [pid 313335] +++ exited with 0 +++
>>>>>> +++ exited with 0 +++
>>>>>>
>>>>>> Where without this patch:
>>>>>>
>>>>>> $ strace -e madvise -f -E GLIBC_TUNABLES=glibc.malloc.hugetlb=1 ./t
>>>>>> madvise(0xc3da42fb9000, 2097152, MADV_HUGEPAGE) = 0
>>>>>> madvise(0xc3da43200000, 2097152, MADV_HUGEPAGE) = 0
>>>>>> strace: Process 313347 attached
>>>>>> [pid 313347] madvise(0xfa19510a0000, 8314880, MADV_DONTNEED) = 0
>>>>>> [pid 313347] +++ exited with 0 +++
>>>>>> +++ exited with 0 +++
>>>>> Thanks for testing! Although without the patch, I get this:
>>>>>
>>>>> madvise(0xfffff8000000, 2097152, MADV_HUGEPAGE) = 0
>>>>> madvise(0xfffff8000000, 4194304, MADV_HUGEPAGE) = 0
>>>>> madvise(0xfffff8000000, 6291456, MADV_HUGEPAGE) = 0
>>>>> madvise(0xfffff7400000, 65536, 0x66 /* MADV_??? */) = -1 EINVAL (Invalid argument)
>>>>> strace: Process 3162240 attached
>>>>> [pid 3162240] madvise(0xfffff0000000, 2097152, MADV_HUGEPAGE) = 0
>>>>> [pid 3162240] madvise(0xfffff7400000, 8314880, MADV_DONTNEED) = 0
>>>>> [pid 3162240] +++ exited with 0 +++
>>>>> +++ exited with 0 +++
>>>>>
>>>>> I don't know where the EINVAL is coming from, that shouldn't be happening, because we
>>>>> have hardcoded MADV_HUGEPAGE everywhere. I played around and I think that when madvise() is not invoked
>>>>> ever, then strace shows it returning EINVAL. I am not familiar with how strace actually
>>>>> works and how to interpret the output.
>>>> The EINVAL most likely come from the size not being multiple of the large
>>>> page size, but I am not sure. I could not understand why the flags are being
>>>> shown as invalid (0x66), since I step on every madvise syscall trap and
>>>> I did not see this valug.
>>> For MADV_HUGEPAGE, the size does not matter, if you just do a madvise(brk, 1, MADV_HUGEPAGE)
>>> it will work.
>> The EINVAL from the 0x66 flags is actually the MADV_GUARD_INSTALL,
>> and neither the kernel nor strace has support for it.
>>
>>>>> But, there is an obvious error in this patch. I shouldn't be making the madvise() call
>>>>> on brk pointer, but on init_sbrk_base. Calling it on brk will cross the heap VMA, and
>>>>> the target range may include a hole. I checked the kernel source and it returns -ENOMEM
>>>>> if it encounters a gap.
>>>>>
>>>>> But I have a question. In madvise_thp, we align the pointer down if it is not aligned to
>>>>> pagesize. If before that pointer, there is a gap, then madvise() will return -ENOMEM,
>>>>> although, the madvise() operation will continue on the VMA which we have specified.
>>>>> So, the operation will do what we want it to do, but return -ENOMEM. Do you think
>>>>> this is something which needs to be fixed? The fix is simple: Align the pointer up
>>>>> instead of down, and subtract from size instead of add. Then we can check again
>>>>> whether size >= thp_pagesize, and continue with the syscall.
>>>> Are you referring to the case where the initial heap is not largepage aligned, then
>>>> the align down to a possible invalid address? If kernel returns ENOMEM in such case
>>>> (I have not checked yet), your suggestion might work although I think it will not
>>>> cover the initial page.
>>> I am referring to this:
>>>
>>> /* Linux requires the input address to be page-aligned, and unaligned
>>> inputs happens only for initial data segment. */
>>> if(__glibc_unlikely(!PTR_IS_ALIGNED(p, GLRO(dl_pagesize))))
>>> {
>>> void*q = PTR_ALIGN_DOWN(p, GLRO(dl_pagesize));
>>> size += PTR_DIFF(p, q);
>>> p = q;
>>> }
>>>
>>> In this case we should be aligning the pointer up and subtracting from size.
>>> I'll send a patch to fix this - although it does not really matter, but I think
>>> we should do this cleanly and not suppress the -ENOMEM even when the MADV_HUGEPAGE
>>> operation will work.
>> I think the ENOMEM from MADV_HUGEPAGE come from trying to madvise on unmapped regions.
>> From mm/madvise.c:
>>
>> 1609 /*
>> 1610 * Walk the vmas in range [start,end), and call the madvise_vma_behavior
>> 1611 * function on each one. The function will get start and end parameters that
>> 1612 * cover the overlap between the current vma and the original range. Any
>> 1613 * unmapped regions in the original range will result in this function returning
>> 1614 * -ENOMEM while still calling the madvise_vma_behavior function on all of the
>> 1615 * existing vmas in the range. Must be called with the mmap_lock held for
>> 1616 * reading or writing.
>> 1617 */
>> 1618 static
>> 1619 int madvise_walk_vmas(struct madvise_behavior *madv_behavior)
>> 1620 {
>> [...]
>> 1644 for (;;) {
>> 1645 /* Still start < end. */
>> 1646 if (!vma)
>> 1647 return -ENOMEM;
>> 1648
>> 1649 /* Here start < (last_end|vma->vm_end). */
>> 1650 if (range->start < vma->vm_start) {
>> 1651 /*
>> 1652 * This indicates a gap between VMAs in the input
>> 1653 * range. This does not cause the operation to abort,
>> 1654 * rather we simply return -ENOMEM to indicate that this
>> 1655 * has happened, but carry on.
>> 1656 */
>> 1657 unmapped_error = -ENOMEM;
>> 1658 range->start = vma->vm_start;
>> 1659 if (range->start >= last_end)
>> 1660 break;
>> 1661 }
>>
>> With the patch applied, on first madvise:
>>
>> (gdb) bt
>> #0 0x0000fffff7e9774c in __GI_madvise () at ../sysdeps/unix/syscall-template.S:120
>> #1 0x0000fffff7e3b11c in madvise_thp (p=<optimized out>, size=<optimized out>) at malloc.c:2085
>> #2 madvise_thp (p=0xaaaaaac00000, size=<optimized out>) at malloc.c:2063
>> #3 sysmalloc (nb=nb@entry=4112, av=av@entry=0xfffff7f70a50 <main_arena>) at malloc.c:2704
>> #4 0x0000fffff7e3c160 in _int_malloc (av=av@entry=0xfffff7f70a50 <main_arena>, bytes=bytes@entry=4096) at malloc.c:4664
>> #5 0x0000fffff7e3c358 in __libc_malloc2 (bytes=4096) at malloc.c:3479
>> #6 0x0000fffff7f80948 in ?? ()
>> #7 0x0000fffff7f80918 in ?? ()
>> Backtrace stopped: not enough registers or memory available to unwind further
>> (gdb) i r x0 x1 x2
>> x0 0xaaaaaac00000 187649985871872
>> x1 0x355000 3493888
>> x2 0xe 14
>> (gdb) info proc mappings
>> process 3827900
>> Mapped address spaces:
>>
>> Start Addr End Addr Size Offset Perms objfile
>> 0xaaaaaaaab000 0xaaaaaae00000 0x355000 0x0 rw-p [heap]
>> 0xfffff7d90000 0xfffff7f52000 0x1c2000 0x0 r-xp [...]/libc.so
>> 0xfffff7f52000 0xfffff7f6d000 0x1b000 0x1c2000 ---p [...]/libc.so
>> 0xfffff7f6d000 0xfffff7f70000 0x3000 0x1cd000 r--p [...]/libc.so
>> 0xfffff7f70000 0xfffff7f72000 0x2000 0x1d0000 rw-p [...]/libc.so
>> 0xfffff7f72000 0xfffff7f7e000 0xc000 0x0 rw-p
>> 0xfffff7f80000 0xfffff7f81000 0x1000 0x0 r-xp [...]/t
>> 0xfffff7f81000 0xfffff7f9f000 0x1e000 0x1000 ---p [...]/t
>> 0xfffff7f9f000 0xfffff7fa0000 0x1000 0xf000 r--p [...]/t
>> 0xfffff7fa0000 0xfffff7fa1000 0x1000 0x10000 rw-p [...]/t
>> 0xfffff7fb0000 0xfffff7fd8000 0x28000 0x0 r-xp [...]/ld.so
>> 0xfffff7fee000 0xfffff7ff0000 0x2000 0x2e000 r--p [...]/ld.so
>> 0xfffff7ff0000 0xfffff7ff1000 0x1000 0x30000 rw-p [...]/ld.so
>> 0xfffff7ff1000 0xfffff7ff2000 0x1000 0x0 rw-p
>> 0xfffff7ff8000 0xfffff7ffa000 0x2000 0x0 rw-p
>> 0xfffff7ffa000 0xfffff7ffe000 0x4000 0x0 r--p [vvar]
>> 0xfffff7ffe000 0xfffff8000000 0x2000 0x0 r-xp [vdso]
>> 0xfffffffde000 0x1000000000000 0x22000 0x0 rw-p [stack]
>>
>> And the patch tries to madvise the region of [0xaaaaaac00000,0xaaaaaaf55000], which is
>> larger than the allocated heap. I am not sure if kernel enables khugepage/THP page
>> by page, meaning that it will have partial enablement if the memory region is is outside
>> of the mapped range.
>
> Missed to answer your doubt: as explained in the patch description, kernel *does not* enable
> THP page by page. The granularity is VMA, not page. If the VMA which you are targetting is
> the range R1 and you do a madvise on range R2, all you need to ensure is that R1's intersection
> with R2 contains at least one page. Of course, to preserve sanity, if you have a VMA in the range
> [4, 9], you will never write code doing a madvise call on the range [9, 20]: it does the job, but
> looks insane, may return ENOMEM (which we ultimately ignore) if there is a gap after 9th page, and
> may end up having the effect of madvise on neighbouring VMAs.
>
Right, so any intersection of the mapped VMA and the madvise VMA will be considered
to be backed by MADV_HUGEPAGE, even if kernel returns ENOMEM?
For instance, let's we have:
process mapping: |map1|gap|map2|___|
madvise |__|madvise-------|
Meaning that map1 has an inserction range, map2 is fully covered (and
assuming the intersection is larger than the page size), and gap is an
unmmaped range.
Would kernel bail on 'gap' or will it continue on all ranges (including map2)?
More information about the Libc-alpha
mailing list