[PATCH v2] Allow shrinking of arena heaps using mmap based on overcommit settings when available

Siddhesh Poyarekar siddhesh@redhat.com
Tue Aug 21 18:37:00 GMT 2012


On Mon, 13 Aug 2012 02:32:36 -0400, KOSAKI wrote:
> (Again, MADV_DONTNEED has good multi thread performance than
> PROT_NONE.)

Right, so I did another program[1] that uses malloc/free rather than the
system calls directly to get an idea of the kind of performance
improvement we would see in a multithreaded scenario and as it turns
out, there is a significant difference as you said.

The program I wrote, spawns 8 threads (conservative number for my 4
core box), each of them allocating 64MB in splits of 16MB
(MMAP_THRESHOLD adjusted to put these on arenas) and freeing them in
reverse order, resulting in a heap trim at each free call. Each of
those threads did 1024 such iterations, giving me 4096 samples. The
results were as follows:

With madvise:

RUNS=4096, TOTAL=9264596.000000, AVG=2261.864258
RUNS=4096, TOTAL=9541428.000000, AVG=2329.450195
RUNS=4096, TOTAL=9381367.000000, AVG=2290.372803
RUNS=4096, TOTAL=9495634.000000, AVG=2318.270020
RUNS=4096, TOTAL=9654729.000000, AVG=2357.111572
RUNS=4096, TOTAL=9202857.000000, AVG=2246.791260
RUNS=4096, TOTAL=9444178.000000, AVG=2305.707520
RUNS=4096, TOTAL=9105736.000000, AVG=2223.080078

With mmap:

RUNS=4096, TOTAL=10211110.000000, AVG=2492.946777
RUNS=4096, TOTAL=10429254.000000, AVG=2546.204590
RUNS=4096, TOTAL=10463869.000000, AVG=2554.655518
RUNS=4096, TOTAL=10599614.000000, AVG=2587.796387
RUNS=4096, TOTAL=10318008.000000, AVG=2519.044922
RUNS=4096, TOTAL=10509473.000000, AVG=2565.789307
RUNS=4096, TOTAL=10325819.000000, AVG=2520.951904
RUNS=4096, TOTAL=10497513.000000, AVG=2562.869385

which gives a roughly 300us advantage per call on average, which is
quite significant. However, since the madvise will still break when
vm.overcommit_memory == 2, I tried a different approach for the patch.
Instead of removing that check, I introduced some Linux-specific code
that checks the contents of /proc/sys/vm.overcommit_memory and if it is
2, uses MAP_FIXED to drop the PTEs. I have attached the patch for
review.

Regards,
Siddhesh

ChangeLog:

	* malloc/arena.c: Include malloc-sysdep.h.
	(shrink_heap): New static variable may_shrink_heap to decide if
	madvise is sufficient to shrink the heap or an unmap is needed.
	* sysdeps/generic/Makefile (sysdep_headers): Add
	malloc-sysdep.h as a dependency when building malloc.
	* sysdeps/generic/malloc-sysdep.h: New file.  Define
	new function check_may_shrink_heap.
	* sysdeps/unix/sysv/linux/Makefile (sysdep_headers): Add
	malloc-sysdep.h as dependency when building malloc.
	* sysdeps/unix/sysv/linux/malloc-sysdep.h: New file.  Define
	new function check_may_shrink_heap.


[1] The benchmark code:

#include <pthread.h>
#include <stdlib.h>
#include <string.h>
#include <stdio.h>
#include <malloc.h>
#include <assert.h>

#define NTHR 8
#define BLOCKSIZE 16*1024*1024
#define COUNT 1024

pthread_mutex_t printf_lock = PTHREAD_MUTEX_INITIALIZER;

void *
thr (void *num)
{
  double total = 0.0;
  int i;

  for (i = 0; i < COUNT; i++)
    {
      /* 4 blocks of 16M each to fill my 64M heap.  */
      int j = 0;
      void *m[4];

      for (j = 0; j < 4; j++)
	{
	  m[j] = malloc (BLOCKSIZE);
	  assert (m[j] != NULL);
	  memset (m[j], 0, BLOCKSIZE);
	}

      for (j = 3; j >= 0; j--)
	{
	  struct timespec before, after;
	  double elapsed;

	  clock_gettime (CLOCK_MONOTONIC_RAW, &before);
	  free (m[j]);
	  clock_gettime (CLOCK_MONOTONIC_RAW, &after);
	  elapsed = (after.tv_sec - before.tv_sec) * 1000000;
	  elapsed += (after.tv_nsec - before.tv_nsec) / 1000;
	  total += elapsed;
	  pthread_mutex_lock (&printf_lock);
	  printf ("%d:%d:%d: %lf\n", (long)num, i, j, elapsed);
	  pthread_mutex_unlock (&printf_lock);
	}
    }
  pthread_mutex_lock (&printf_lock);
  for (i = 0; i < 80; i++)
    putchar ('=');
  putchar ('\n');
  printf ("RUNS=%d, TOTAL=%lf, AVG=%lf\n", COUNT * 4,
	  total, total / (COUNT * 4));
  pthread_mutex_unlock (&printf_lock);
}

int
main ()
{
  long i;
  pthread_t t[NTHR];

  mallopt (M_MMAP_THRESHOLD, 31 * 1024 * 1024);

  for (i = 0; i < NTHR; i++)
    {
      pthread_create (&t[i], NULL, thr, (void *)i);
    }

  for (i = 0; i < NTHR; i++)
    {
      pthread_join (t[i], NULL);
    }
  return 0;
}
-------------- next part --------------
A non-text attachment was scrubbed...
Name: heap-shrink.patch
Type: text/x-patch
Size: 5318 bytes
Desc: not available
URL: <http://sourceware.org/pipermail/libc-alpha/attachments/20120821/030b7d6d/attachment.bin>


More information about the Libc-alpha mailing list