[PATCH v2] malloc: Enable 2MB THP by default on Aarch64
Adhemerval Zanella Netto
adhemerval.zanella@linaro.org
Thu Dec 11 12:36:25 GMT 2025
On 05/12/25 11:13, Dev Jain wrote:
> Linux supports multi-sized Transparent Huge Pages (mTHP). For the purpose
> of this patch description, we call the block size mapped by a non-last
> level pagetable level, the traditional THP size (2M for 4K basepage,
> 512M for 64K basepage). Linux now also supports intermediate THP sizes
> mapped by the last level pagetable - we call that the mTHP size.
>
> The support for mTHP in Linux has grown to be better and stable over time -
> applications can benefit from reduced page faults and reduced kernel
> memory management overhead, albeit at the cost of internal fragmentation.
> We have observed consistent performance boosts with mTHP with little
> variance.
>
> As a result, enable 2M THP by default on Aarch64. This enables THP even if
> user hasn't passed glibc.malloc.hugetlb=1. If user has passed it, we avoid
> making the system call to check the hugepage size from sysfs, and override
> it with the hardcoded 2MB.
>
> There are two additional benefits of this patch, if the transparent
> hugepage sysctl is set to madvise or always:
>
> 1) The THP size is now hardcoded to 2MB for Aarch64. This avoids a
> syscall for fetching the THP size from sysfs.
>
> 2) On 64K basepage size systems, the traditional THP size is 512M, which
> is unusable and impractical. We can instead benefit from the size of
> 2M. Apart from the usual benefit of THPs/mTHPs as described above, Aarch64
> systems benefit from reduced TLB pressure on this mTHP size, commonly
> known as the "contpte" size. If the application takes a pagefault, and
> either the THP sysctl settings is "always", or the virtual memory area
> has been madvise(MADV_HUGEPAGE)'d along with sysctl being "madvise", then
> Linux will fault in a 2M mTHP, mapping contiguous pages into the pagetable,
> and painting the pagetable entries with the cont-bit. This bit is a hint to
> the hardware that the concerned pagetable entry maps a page which is part
> of a set of contiguous pages - the TLB then only remembers a single entry
> for this set of 2M/64K = 32 pages, because the physical address of any
> other page in this contiguous set is computable by the TLB cached physical
> address via a linear offset. Hence, what was only possible with the
> traditional THP size, is now possible with the mTHP size.
>
> We see a 6.25% performance improvement on SPEC.
>
> If the sysctl is set to never, no transparent hugepages will be created by
> the kernel. But, this patch still sets thp_pagesize = 2MB. The benefit is
> that on MORECORE() invocation, we extend the heap by 2MB instead of 4KB,
> potentially reducing the frequency of this syscall's invocation by 512x.
> Note that, there is no difference in cost between an sbrk(2M) and sbrk(4K);
> the kernel only does a virtual reservation and does not touch user physical
> memory.
The only thing bothering me is now there is no option users can *disable*
it for whatever reason. We should at least update the manual/tunables.texi
documentation that glibc.malloc.hugetlb=1 is the default for AArch64, and
that it uses a different Huge Page than the one defined by the kernel
(/sys/kernel/mm/transparent_hugepage/hpage_pmd_size).
Maybe it would be better to work toward adding an aarch64-specific tunable
knob to set glibc.malloc.hugetlb=1 as the default. At least it aligns better
with other architectures and gives aarch64 an option to disable this if
users want the previous semantics.
>
> Hacking the bench-malloc-thread.c benchmark:
>
> diff --git a/benchtests/bench-malloc-thread.c b/benchtests/bench-malloc-thread.c
> index 7759c61847..7eed33acb7 100644
> --- a/benchtests/bench-malloc-thread.c
> +++ b/benchtests/bench-malloc-thread.c
> @@ -51,8 +51,8 @@
> the typical amount allocated will be much smaller. */
> #define WORKING_SET_SIZE 1024
>
> -#define MIN_ALLOCATION_SIZE 4
> -#define MAX_ALLOCATION_SIZE 32768
> +#define MIN_ALLOCATION_SIZE 1 << 15
> +#define MAX_ALLOCATION_SIZE 1 << 16
>
> /* Get a random block size with an inverse square distribution. */
> static unsigned int
> @@ -135,13 +135,11 @@ malloc_benchmark_loop (void **ptr_arr)
> unsigned int offset_state = 0, block_state = 0;
> size_t iters = 0;
>
> - while (!timeout)
> + while (iters++ < 100000)
> {
> unsigned int next_idx = get_random_offset (&offset_state);
> unsigned int next_block = get_random_block_size (&block_state);
>
> - free (ptr_arr[next_idx]);
> -
> ptr_arr[next_idx] = TEST_FUNC (next_block);
>
> iters++;
>
> This ensures that allocations are big (but less than the mmap threshold),
> and since we do not free a chunk, the heap is forced to be extended.
>
> Running this benchmark with /sys/kernel/mm/transparent_hugepage/enabled
> = never and 100 threads, the result for vanilla malloc:
>
> "timing_type": "hp_timing",
> "functions": {
> "malloc": {
> "": {
> "duration": 5.46735e+11,
> "iterations": 1.00001e+07,
> "time_per_iteration": 54673,
> "max_rss": 1.99992e+07,
> "threads": 100,
> "min_size": 32768,
> "max_size": 65536,
> "random_seed": 88
> }
> }
> }
>
> whereas the result for patched malloc:
>
> "timing_type": "hp_timing",
> "functions": {
> "malloc": {
> "": {
> "duration": 5.65185e+10,
> "iterations": 1.00001e+07,
> "time_per_iteration": 5651.8,
> "max_rss": 2.00079e+07,
> "threads": 100,
> "min_size": 32768,
> "max_size": 65536,
> "random_seed": 88
> }
> }
> }
>
> The result is a 10x speed up on malloc.
> ---
> Patch applies on 3dd2cbfa35e0.
>
> v1->v2:
> - Invoke thp_init() also before we extend the brk for the first time
>
> malloc/malloc.c | 20 ++++++++++++++++++
> sysdeps/generic/malloc-hugepages.h | 4 ++++
> .../sysv/linux/aarch64/malloc-hugepages.h | 21 +++++++++++++++++++
> sysdeps/unix/sysv/linux/malloc-hugepages.c | 3 +++
> 4 files changed, 48 insertions(+)
> create mode 100644 sysdeps/unix/sysv/linux/aarch64/malloc-hugepages.h
>
> diff --git a/malloc/malloc.c b/malloc/malloc.c
> index bd92d5c396..f9384e06cc 100644
> --- a/malloc/malloc.c
> +++ b/malloc/malloc.c
> @@ -2053,10 +2053,17 @@ free_perturb (char *p, size_t n)
>
> /* ----------- Routines dealing with transparent huge pages ----------- */
>
> +static void thp_init (void);
> +
> static inline void
> madvise_thp (void *p, INTERNAL_SIZE_T size)
> {
> #ifdef MADV_HUGEPAGE
> +
> + /* Ensure thp_init () is invoked only once */
> + if (mp_.thp_pagesize < DEFAULT_THP_PAGESIZE)
> + thp_init ();
> +
> /* Only use __madvise if the system is using 'madvise' mode.
> Otherwise the call is wasteful. */
> if (mp_.thp_mode != malloc_thp_mode_madvise)
> @@ -2664,6 +2671,10 @@ sysmalloc (INTERNAL_SIZE_T nb, mstate av)
> previous calls. Otherwise, we correct to page-align below.
> */
>
> + /* Ensure thp_init () is invoked only once */
> + if (mp_.thp_pagesize < DEFAULT_THP_PAGESIZE)
> + thp_init ();
> +
> if (__glibc_unlikely (mp_.thp_pagesize != 0))
> {
> uintptr_t lastbrk = (uintptr_t) MORECORE (0);
> @@ -5660,6 +5671,15 @@ do_set_hugetlb (size_t value)
> return 0;
> }
>
> +static __always_inline void
> +thp_init (void)
> +{
> + /* thp_pagesize is set even if thp_mode is never. This reduces frequency
> + of MORECORE () invocation. */
> + mp_.thp_pagesize = DEFAULT_THP_PAGESIZE;
> + mp_.thp_mode = __malloc_thp_mode ();
> +}
> +
> int
> __libc_mallopt (int param_number, int value)
> {
> diff --git a/sysdeps/generic/malloc-hugepages.h b/sysdeps/generic/malloc-hugepages.h
> index 076f03d68a..9c034e86c8 100644
> --- a/sysdeps/generic/malloc-hugepages.h
> +++ b/sysdeps/generic/malloc-hugepages.h
> @@ -41,4 +41,8 @@ enum malloc_thp_mode_t __malloc_thp_mode (void) attribute_hidden;
> void __malloc_hugepage_config (size_t requested, size_t *pagesize, int *flags)
> attribute_hidden;
>
> +#ifndef DEFAULT_THP_PAGESIZE
> +# define DEFAULT_THP_PAGESIZE 0
> +#endif
> +
> #endif /* _MALLOC_HUGEPAGES_H */
> diff --git a/sysdeps/unix/sysv/linux/aarch64/malloc-hugepages.h b/sysdeps/unix/sysv/linux/aarch64/malloc-hugepages.h
> new file mode 100644
> index 0000000000..f432ff7f0c
> --- /dev/null
> +++ b/sysdeps/unix/sysv/linux/aarch64/malloc-hugepages.h
> @@ -0,0 +1,21 @@
> +/* Huge Page support. Aarch64 Linux implementation.
> + Copyright (C) 2021-2025 Free Software Foundation, Inc.
> + This file is part of the GNU C Library.
> +
> + The GNU C Library is free software; you can redistribute it and/or
> + modify it under the terms of the GNU Lesser General Public License as
> + published by the Free Software Foundation; either version 2.1 of the
> + License, or (at your option) any later version.
> +
> + The GNU C Library is distributed in the hope that it will be useful,
> + but WITHOUT ANY WARRANTY; without even the implied warranty of
> + MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU
> + Lesser General Public License for more details.
> +
> + You should have received a copy of the GNU Lesser General Public
> + License along with the GNU C Library; see the file COPYING.LIB. If
> + not, see <https://www.gnu.org/licenses/>. */
> +
> +# define DEFAULT_THP_PAGESIZE 1UL << 21
> +
> +#include_next <malloc-hugepages.h>
> diff --git a/sysdeps/unix/sysv/linux/malloc-hugepages.c b/sysdeps/unix/sysv/linux/malloc-hugepages.c
> index e23cdfb6b7..9776ba4c35 100644
> --- a/sysdeps/unix/sysv/linux/malloc-hugepages.c
> +++ b/sysdeps/unix/sysv/linux/malloc-hugepages.c
> @@ -25,6 +25,9 @@
> unsigned long int
> __malloc_default_thp_pagesize (void)
> {
> + if (DEFAULT_THP_PAGESIZE != 0)
> + return DEFAULT_THP_PAGESIZE;
> +
> int fd = __open64_nocancel (
> "/sys/kernel/mm/transparent_hugepage/hpage_pmd_size", O_RDONLY);
> if (fd == -1)
More information about the Libc-alpha
mailing list