Deadlock in glibc rwlock on arm32

Adhemerval Zanella adhemerval.zanella@linaro.org
Mon Jun 27 19:30:55 GMT 2022



> On 27 Jun 2022, at 16:26, Mathieu Desnoyers via Libc-alpha <libc-alpha@sourceware.org> wrote:
> 
> Hi Carlos,
> 
> I am hitting a spurious hang in my liburcu CI infrastructure on arm32 boards.
> It happens in the "rwlock" tests, which is really just a baseline to test the
> speed of the glibc rwlock. I therefore suspect either a glibc or a Linux kernel
> bug.
> 
> Some info about the reproducer setup I have:
> 
> Architecture:                    armv7l
> Byte Order:                      Little Endian
> CPU(s):                          4
> On-line CPU(s) list:             0-3
> Thread(s) per core:              1
> Core(s) per socket:              4
> Socket(s):                       1
> Vendor ID:                       ARM
> Model:                           10
> Model name:                      Cortex-A9
> Stepping:                        r2p10
> CPU max MHz:                     996.0000
> CPU min MHz:                     396.0000
> 
> glibc: 2.31-13+deb11u3
> 
> failing test: liburcu [1] branch stable-0.13 (b8359af6329d387a6ebb411acb85924b4d1be625)
> configured with: ./configure --enable-rcu-debug
> test program: tests/benchmark/test_rwlock
> 
> Running the test in a loop got the issue to reproduce approximately
> each 100 iterations.
> 
> Here are the backtraces when looking at the stuck program with gdb:
> 
> (gdb) thread apply all bt
> 
> Thread 5 (Thread 0xb5691450 (LWP 7024) "test_rwlock"):
> #0  __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1  0xb6f9f05c in futex_abstimed_wait (private=0, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2  __pthread_rwlock_wrlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:731
> #3  __GI___pthread_rwlock_wrlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_wrlock.c:27
> #4  0x004f0fe0 in thr_writer (_count=0x119f1d0) at test_rwlock.c:205
> #5  0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6  0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
> 
> Thread 4 (Thread 0xb5e92450 (LWP 7023) "test_rwlock"):
> #0  __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1  0xb6f9eea4 in futex_abstimed_wait (private=0, abstime=0x0, clockid=0, expected=2, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2  __pthread_rwlock_wrlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:830
> #3  __GI___pthread_rwlock_wrlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_wrlock.c:27
> #4  0x004f0fe0 in thr_writer (_count=0x119f1c8) at test_rwlock.c:205
> #5  0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6  0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
> 
> Thread 3 (Thread 0xb6693450 (LWP 7022) "test_rwlock"):
> #0  __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1  0xb6f9e7be in futex_abstimed_wait (private=<optimized out>, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2  __pthread_rwlock_rdlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:460
> #3  __GI___pthread_rwlock_rdlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_rdlock.c:27
> #4  0x004f11ba in thr_reader (_count=0x119f1b8) at test_rwlock.c:157
> #5  0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6  0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
> 
> Thread 2 (Thread 0xb6e94450 (LWP 7021) "test_rwlock"):
> #0  __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1  0xb6f9e7be in futex_abstimed_wait (private=<optimized out>, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2  __pthread_rwlock_rdlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:460
> #3  __GI___pthread_rwlock_rdlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_rdlock.c:27
> #4  0x004f11ba in thr_reader (_count=0x119f1b0) at test_rwlock.c:157
> #5  0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6  0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
> 
> Thread 1 (Thread 0xb6fe1d60 (LWP 7020) "test_rwlock"):
> #0  __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1  0xb6f9bafc in __pthread_clockjoin_ex (threadid=3068740688, thread_return=thread_return@entry=0xbeb9ba8c, clockid=clockid@entry=0, abstime=abstime@entry=0x0, block=block@entry=true) at pthread_join_common.c:145
> #2  0xb6f9b8ec in __pthread_join (threadid=<optimized out>, thread_return=thread_return@entry=0xbeb9ba8c) at pthread_join.c:24
> #3  0x004f0b20 in main (argc=<optimized out>, argv=0xbeb9bc14) at test_rwlock.c:367
> 
> Here is the rwlock state:
> 
> (gdb) print *rwlock
> $5 = {__data = {__readers = 19, __writers = 0, __wrphase_futex = 3, __writers_futex = 3, __pad3 = 0, __pad4 = 0, __flags = 0 '\000', __shared = 0 '\000', 
>    __pad1 = 0 '\000', __pad2 = 0 '\000', __cur_writer = 0}, __size = "\023\000\000\000\000\000\000\000\003\000\000\000\003", '\000' <repeats 18 times>, __align = 19}
> 
> Reader threads { 2, 3 } are stuck on:
> 
>          int err = __futex_abstimed_wait64 (&rwlock->__data.__wrphase_futex,
>                                             1 | PTHREAD_RWLOCK_FUTEX_USED,
>                                             clockid, abstime, private);
> 
> One writer thread { 5 } is stuck on:
> 
>          int err = __futex_abstimed_wait64 (&rwlock->__data.__writers_futex,
>                                             1 | PTHREAD_RWLOCK_FUTEX_USED,
>                                             clockid, abstime, private);
> The other writer thread { 4 } is stuck on:
> 
>          int err = __futex_abstimed_wait64 (&rwlock->__data.__wrphase_futex,
>                                             PTHREAD_RWLOCK_FUTEX_USED,
>                                             clockid, abstime, private);
> 
> I notice that you authored commit faf8c066d ("rwlock: Fix explicit hand-over (bug 21298)") in the same area where the hangs happen.
> 
> There are 2 reader and 2 writer threads blocked on futex_abstimed_wait. The __wrphase_futex = 3 which means I would expect the
> writer thread {4} __pthread_rwlock_wrlock_full to be unblocked since the parameter passed to futex(5) is the value to expect
> (and block for). Somehow this writer thread is not awakened.
> 
> It's weird that __pthread_rwlock_rdlock_full can set the state to 1 | PTHREAD_RWLOCK_FUTEX_USED with atomic_compare_exchange_weak_relaxed
> without doing any futex wake:
> 
>      while (((wpf = atomic_load_relaxed (&rwlock->__data.__wrphase_futex))
>              | PTHREAD_RWLOCK_FUTEX_USED) == (1 | PTHREAD_RWLOCK_FUTEX_USED))
>        {
>          int private = __pthread_rwlock_get_private (rwlock);
>          if (((wpf & PTHREAD_RWLOCK_FUTEX_USED) == 0)
>              && (!atomic_compare_exchange_weak_relaxed
>                  (&rwlock->__data.__wrphase_futex,
>                   &wpf, wpf | PTHREAD_RWLOCK_FUTEX_USED)))
>            continue;
> 
> Thoughts ?

Maybe you are hitting https://sourceware.org/bugzilla/show_bug.cgi?id=24774 ?
Could you check if https://patchwork.sourceware.org/project/glibc/patch/20210929191430.884057-1-adhemerval.zanella@linaro.org/ helps?

> 
> FYI my testing infrastructure will be down in the coming days due to work on the AC plumbing, so
> my ability to test changes will be somewhat delayed.
> 
> Thanks,
> 
> Mathieu
> 
> [1] http://liburcu.org/
> 
> -- 
> Mathieu Desnoyers
> EfficiOS Inc.
> http://www.efficios.com



More information about the Libc-alpha mailing list