[PATCH 2/3] Mutex: Only read while spinning
Adhemerval Zanella
adhemerval.zanella@linaro.org
Thu Apr 5 20:55:00 GMT 2018
On 30/03/2018 04:14, Kemi Wang wrote:
> The pthread adaptive spin mutex spins on the lock for a while before going
> to a sleep. While the lock is contended and we need to wait, going straight
> back to LLL_MUTEX_TRYLOCK(cmpxchg) is not a good idea on many targets as
> that will force expensive memory synchronization among processors and
> penalize other running threads. For example, it constantly floods the
> system with "read for ownership" requests, which are much more expensive to
> process than a single read. Thus, we only use MO read until we observe the
> lock to not be acquired anymore, as suggusted by Andi Kleen.
>
> Test machine:
> 2-sockets Skylake paltform, 112 cores with 62G RAM
>
> Test case: Contended pthread adaptive spin mutex with global update
> each thread of the workload does:
> a) Lock the mutex (adaptive spin type)
> b) Globle variable increment
> c) Unlock the mutex
> in a loop until timeout, and the main thread reports the total iteration
> number of all the threads in one second.
>
> This test case is as same as Will-it-scale.pthread_mutex3 except mutex type is
> modified to PTHREAD_MUTEX_ADAPTIVE_NP.
> github: https://github.com/antonblanchard/will-it-scale.git
>
> nr_threads base head(SPIN_COUNT=10) head(SPIN_COUNT=1000)
> 1 51644585 51307573(-0.7%) 51323778(-0.6%)
> 2 7914789 10011301(+26.5%) 9867343(+24.7%)
> 7 1687620 4224135(+150.3%) 3430504(+103.3%)
> 14 1026555 3784957(+268.7%) 1843458(+79.6%)
> 28 962001 2886885(+200.1%) 681965(-29.1%)
> 56 883770 2740755(+210.1%) 364879(-58.7%)
> 112 1150589 2707089(+135.3%) 415261(-63.9%)
In pthread_mutex3 it is basically more updates in a global variable synchronized
with a mutex, so if I am reading correct the benchmark, a higher value means
less contention. I also assume you use the 'threads' value in this table.
I checked on a 64 cores aarch64 machine to see what kind of improvement, if
any; one would get with this change:
nr_threads base head(SPIN_COUNT=10) head(SPIN_COUNT=1000)
1 27566206 28778254 (4.211680) 28778467 (4.212389)
2 8498813 7777589 (-9.273105) 7806043 (-8.874791)
7 5019434 2869629 (-74.915782) 3307812 (-51.744839)
14 4379155 2906255 (-50.680343) 2825041 (-55.012087)
28 4397464 3261094 (-34.846282) 3259486 (-34.912805)
56 4020956 3898386 (-3.144122) 4038505 (0.434542)
So I think this change should be platform-specific.
>
> Suggested-by: Andi Kleen <andi.kleen@intel.com>
> Signed-off-by: Kemi Wang <kemi.wang@intel.com>
> ---
> nptl/pthread_mutex_lock.c | 23 +++++++++++++++--------
> 1 file changed, 15 insertions(+), 8 deletions(-)
>
> diff --git a/nptl/pthread_mutex_lock.c b/nptl/pthread_mutex_lock.c
> index 1519c14..c3aca93 100644
> --- a/nptl/pthread_mutex_lock.c
> +++ b/nptl/pthread_mutex_lock.c
> @@ -26,6 +26,7 @@
> #include <atomic.h>
> #include <lowlevellock.h>
> #include <stap-probe.h>
> +#include <mutex-conf.h>
>
> #ifndef lll_lock_elision
> #define lll_lock_elision(lock, try_lock, private) ({ \
> @@ -124,16 +125,22 @@ __pthread_mutex_lock (pthread_mutex_t *mutex)
> if (LLL_MUTEX_TRYLOCK (mutex) != 0)
> {
> int cnt = 0;
> - int max_cnt = MIN (MAX_ADAPTIVE_COUNT,
> - mutex->__data.__spins * 2 + 10);
> + int max_cnt = MIN (__mutex_aconf.spin_count,
> + mutex->__data.__spins * 2 + 100);
> do
> {
> - if (cnt++ >= max_cnt)
> - {
> - LLL_MUTEX_LOCK (mutex);
> - break;
> - }
> - atomic_spin_nop ();
> + if (cnt >= max_cnt)
> + {
> + LLL_MUTEX_LOCK (mutex);
> + break;
> + }
> + /* MO read while spinning */
> + do
> + {
> + atomic_spin_nop ();
> + }
> + while (atomic_load_relaxed (&mutex->__data.__lock) != 0 &&
> + ++cnt < max_cnt);
> }
> while (LLL_MUTEX_TRYLOCK (mutex) != 0);
>
>
More information about the Libc-alpha
mailing list