Deadlock in glibc rwlock on arm32
Adhemerval Zanella
adhemerval.zanella@linaro.org
Mon Jun 27 19:30:55 GMT 2022
> On 27 Jun 2022, at 16:26, Mathieu Desnoyers via Libc-alpha <libc-alpha@sourceware.org> wrote:
>
> Hi Carlos,
>
> I am hitting a spurious hang in my liburcu CI infrastructure on arm32 boards.
> It happens in the "rwlock" tests, which is really just a baseline to test the
> speed of the glibc rwlock. I therefore suspect either a glibc or a Linux kernel
> bug.
>
> Some info about the reproducer setup I have:
>
> Architecture: armv7l
> Byte Order: Little Endian
> CPU(s): 4
> On-line CPU(s) list: 0-3
> Thread(s) per core: 1
> Core(s) per socket: 4
> Socket(s): 1
> Vendor ID: ARM
> Model: 10
> Model name: Cortex-A9
> Stepping: r2p10
> CPU max MHz: 996.0000
> CPU min MHz: 396.0000
>
> glibc: 2.31-13+deb11u3
>
> failing test: liburcu [1] branch stable-0.13 (b8359af6329d387a6ebb411acb85924b4d1be625)
> configured with: ./configure --enable-rcu-debug
> test program: tests/benchmark/test_rwlock
>
> Running the test in a loop got the issue to reproduce approximately
> each 100 iterations.
>
> Here are the backtraces when looking at the stuck program with gdb:
>
> (gdb) thread apply all bt
>
> Thread 5 (Thread 0xb5691450 (LWP 7024) "test_rwlock"):
> #0 __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1 0xb6f9f05c in futex_abstimed_wait (private=0, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2 __pthread_rwlock_wrlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:731
> #3 __GI___pthread_rwlock_wrlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_wrlock.c:27
> #4 0x004f0fe0 in thr_writer (_count=0x119f1d0) at test_rwlock.c:205
> #5 0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6 0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
>
> Thread 4 (Thread 0xb5e92450 (LWP 7023) "test_rwlock"):
> #0 __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1 0xb6f9eea4 in futex_abstimed_wait (private=0, abstime=0x0, clockid=0, expected=2, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2 __pthread_rwlock_wrlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:830
> #3 __GI___pthread_rwlock_wrlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_wrlock.c:27
> #4 0x004f0fe0 in thr_writer (_count=0x119f1c8) at test_rwlock.c:205
> #5 0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6 0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
>
> Thread 3 (Thread 0xb6693450 (LWP 7022) "test_rwlock"):
> #0 __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1 0xb6f9e7be in futex_abstimed_wait (private=<optimized out>, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2 __pthread_rwlock_rdlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:460
> #3 __GI___pthread_rwlock_rdlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_rdlock.c:27
> #4 0x004f11ba in thr_reader (_count=0x119f1b8) at test_rwlock.c:157
> #5 0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6 0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
>
> Thread 2 (Thread 0xb6e94450 (LWP 7021) "test_rwlock"):
> #0 __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1 0xb6f9e7be in futex_abstimed_wait (private=<optimized out>, abstime=0x0, clockid=0, expected=3, futex_word=<optimized out>) at ../sysdeps/nptl/futex-internal.h:287
> #2 __pthread_rwlock_rdlock_full (abstime=0x0, clockid=0, rwlock=0x5020d8 <lock>) at pthread_rwlock_common.c:460
> #3 __GI___pthread_rwlock_rdlock (rwlock=rwlock@entry=0x5020d8 <lock>) at pthread_rwlock_rdlock.c:27
> #4 0x004f11ba in thr_reader (_count=0x119f1b0) at test_rwlock.c:157
> #5 0xb6f9a98e in start_thread (arg=0x421f5686) at pthread_create.c:477
> #6 0xb6f35bec in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
> Backtrace stopped: previous frame identical to this frame (corrupt stack?)
>
> Thread 1 (Thread 0xb6fe1d60 (LWP 7020) "test_rwlock"):
> #0 __libc_do_syscall () at ../sysdeps/unix/sysv/linux/arm/libc-do-syscall.S:46
> #1 0xb6f9bafc in __pthread_clockjoin_ex (threadid=3068740688, thread_return=thread_return@entry=0xbeb9ba8c, clockid=clockid@entry=0, abstime=abstime@entry=0x0, block=block@entry=true) at pthread_join_common.c:145
> #2 0xb6f9b8ec in __pthread_join (threadid=<optimized out>, thread_return=thread_return@entry=0xbeb9ba8c) at pthread_join.c:24
> #3 0x004f0b20 in main (argc=<optimized out>, argv=0xbeb9bc14) at test_rwlock.c:367
>
> Here is the rwlock state:
>
> (gdb) print *rwlock
> $5 = {__data = {__readers = 19, __writers = 0, __wrphase_futex = 3, __writers_futex = 3, __pad3 = 0, __pad4 = 0, __flags = 0 '\000', __shared = 0 '\000',
> __pad1 = 0 '\000', __pad2 = 0 '\000', __cur_writer = 0}, __size = "\023\000\000\000\000\000\000\000\003\000\000\000\003", '\000' <repeats 18 times>, __align = 19}
>
> Reader threads { 2, 3 } are stuck on:
>
> int err = __futex_abstimed_wait64 (&rwlock->__data.__wrphase_futex,
> 1 | PTHREAD_RWLOCK_FUTEX_USED,
> clockid, abstime, private);
>
> One writer thread { 5 } is stuck on:
>
> int err = __futex_abstimed_wait64 (&rwlock->__data.__writers_futex,
> 1 | PTHREAD_RWLOCK_FUTEX_USED,
> clockid, abstime, private);
> The other writer thread { 4 } is stuck on:
>
> int err = __futex_abstimed_wait64 (&rwlock->__data.__wrphase_futex,
> PTHREAD_RWLOCK_FUTEX_USED,
> clockid, abstime, private);
>
> I notice that you authored commit faf8c066d ("rwlock: Fix explicit hand-over (bug 21298)") in the same area where the hangs happen.
>
> There are 2 reader and 2 writer threads blocked on futex_abstimed_wait. The __wrphase_futex = 3 which means I would expect the
> writer thread {4} __pthread_rwlock_wrlock_full to be unblocked since the parameter passed to futex(5) is the value to expect
> (and block for). Somehow this writer thread is not awakened.
>
> It's weird that __pthread_rwlock_rdlock_full can set the state to 1 | PTHREAD_RWLOCK_FUTEX_USED with atomic_compare_exchange_weak_relaxed
> without doing any futex wake:
>
> while (((wpf = atomic_load_relaxed (&rwlock->__data.__wrphase_futex))
> | PTHREAD_RWLOCK_FUTEX_USED) == (1 | PTHREAD_RWLOCK_FUTEX_USED))
> {
> int private = __pthread_rwlock_get_private (rwlock);
> if (((wpf & PTHREAD_RWLOCK_FUTEX_USED) == 0)
> && (!atomic_compare_exchange_weak_relaxed
> (&rwlock->__data.__wrphase_futex,
> &wpf, wpf | PTHREAD_RWLOCK_FUTEX_USED)))
> continue;
>
> Thoughts ?
Maybe you are hitting https://sourceware.org/bugzilla/show_bug.cgi?id=24774 ?
Could you check if https://patchwork.sourceware.org/project/glibc/patch/20210929191430.884057-1-adhemerval.zanella@linaro.org/ helps?
>
> FYI my testing infrastructure will be down in the coming days due to work on the AC plumbing, so
> my ability to test changes will be somewhat delayed.
>
> Thanks,
>
> Mathieu
>
> [1] http://liburcu.org/
>
> --
> Mathieu Desnoyers
> EfficiOS Inc.
> http://www.efficios.com
More information about the Libc-alpha
mailing list