[PATCH] posix: Do not use WNOHANG in waitpid call for Linux posix_spawn
Florian Weimer
fweimer@redhat.com
Mon Oct 23 10:40:00 GMT 2017
On 10/23/2017 12:38 PM, Szabolcs Nagy wrote:
> On 23/10/17 07:32, Florian Weimer wrote:
>> On 10/22/2017 10:51 PM, Adhemerval Zanella wrote:
>>> As shown in some buildbot issues on aarch64 and powerpc, calling
>>> clone (VFORK) and waitpid (WNOHANG) does not guarantee the child
>>> is ready to be collected. This patch changes the call back to 0
>>> as before fe05e1cb6d64 fix.
>>
>> I see it on x86-64, too. It does look like a kernel bug.
>>
>>> This change can lead to the scenario 4.3 described in the commit,
>>> where the waitpid call can hang undefinitely on the call. However
>>> this is also a very unlikely and also undefinied situation where
>>> both the caller is trying to terminate a pid before posix_spawn
>>> returns and the race pid reuse is triggered. I don't see how to
>>> correct handle this specific situation within posix_spawn.
>>
>> Agreed. I wish we could do better here, but it seems we can't.
>>
>
> musl writes a close-on-exec pipe in the child on error,
> reading it in the parent tells if the child died before
> exec or not. (so the waitpid can be made precise)
How so? A wildcard wait can still collect the PID, which is the problem
we were trying to defend against with the WNOHANG flag.
Thanks,
Florian
More information about the Libc-alpha
mailing list