[PATCH v3] elf: Release dl_load_lock before running dlopen constructors (BZ 15686)
Artem Proskurnev
temap@mail.ru
Tue Jul 21 10:22:05 GMT 2026
* Florian Weimer
> The bug is calling gdk_pixbuf_new_from_file from an ELF constructor?
> Alternatively, the backends probably shouldn't spawn threads unless the
> process is already multi-threaded.
>
> Adhemerval's fix seems simple enough that we can merge it whether it
> helps this particular scenario or not.
>
On the last point, agreed. My earlier reply to Adhemerval in this
thread says explicitly to take his lock-free __cxa_thread_atexit_impl
patch independently of v3. It is a strict improvement and orthogonal
to v3; both should land.
On the first two points, I think the framing is too narrow.
== Constructors that do real work are legitimate
ELF constructors are not limited to type registration and static
initialisers. Across the GLib/GTK/GNOME ecosystem and plugin systems
generally, constructors that perform real work -- loading config,
registering with a session bus, initialising an image loader, setting
up a sandboxed subprocess -- are a long-standing supported pattern.
GObject type registration itself is driven by constructors that can
reach arbitrary code paths.
The glycin case specifically: gdk-pixbuf routes image loading through
a sandboxed loader when one is installed. Calling
gdk_pixbuf_new_from_file() from a plugin ctor is not exotic; it is
what every image-handling plugin does at load time, and it has worked
for years. The regression appeared with Rust 1.94, where
std::thread::spawn gained ChildSpawnHooks backed by a new thread_local,
and that thread_local is what eventually hit __cxa_thread_atexit_impl
on the spawned worker. Pinning the bug on "the ctor does I/O"
rewrites a stable API contract that the deployed ecosystem already
relies on.
And even if we decided ctors shouldn't do I/O, the deployed surface
already does. Telling users to rewrite their plugins is not a
release-time answer, and the same class of bug stays one refactor
away from recurring in whatever replaces the ctor-I/O pattern.
== "Backends shouldn't spawn threads" doesn't fit the glycin case
Two problems with the backend-side framing.
(1) There is no portable, race-free way for a library to check
whether the process is already multi-threaded at the moment it
considers spawning. pthread_is_multithreaded_np does not exist;
inspecting /proc/self/status is racy and not portable. So this
would require glibc to expose a new API, and then require every
thread-spawning library in the ecosystem to adopt it. That is a
larger surface change than v3, not a smaller one.
(2) The glycin case isn't actually "backend spawned a thread, thread
raced with ctor". It is "backend spawned a worker for sandbox setup;
the worker reached getgrouplist -> NSS -> __libc_dlopen_mode ->
_dl_open -> dl_load_lock". The same blocking shape fires for any
code path that touches NSS during the ctor window, with or without a
spawned thread -- a ctor that calls getpwnam, getgrnam, getaddrinfo,
gethostbyname, or any of the get*ent family hits it the same way.
Spawning is one shape; the underlying issue is that dl_load_lock is
held while running arbitrary user code that may legitimately need to
re-enter the loader.
If the answer is "applications should pre-cache NSS results before
calling dlopen", then every plugin host in the ecosystem needs to
pre-call every NSS-touching function any plugin might reach during
construction. Nothing does this today, and neither POSIX nor the
glibc manual currently asks it to.
== Why this points at v3
The two framings above share a property: they accept that dl_load_lock
must be held during ctor execution, and they shift the burden onto
applications and libraries to avoid whatever code paths happen to need
it. That is the per-case patching strategy that produced the BZ 15686
history -- each new manifestation is a new bug, a new narrow fix, and
a new round of downstream regressions.
v3 removes the underlying invariant. Whatever a ctor (or a thread it
spawns) does during the ctor window becomes a legitimate operation,
because the lock those operations compete for is no longer held by
the thread running the ctor. That covers __cxa_thread_atexit_impl,
the NSS path, dlsym, _dl_addr, _dl_find_dso_for_object, recursive
dlopen, and the rest of the surface. Adhemerval's patch cleans up
one of those paths more elegantly than v3 does; it does not remove
the class.
The A/B result earlier in this thread shows this concretely:
azanella/bz15686 deadlocks on the glycin reproducer because the NSS
path is what fires there, not __cxa_thread_atexit_impl. v3 does not.
I would rather we land both patches than choose between them, but if
the choice is forced, v3 is the one that fixes the user-visible bug.
Thanks,
Artem
More information about the Libc-alpha
mailing list