[RFC] Toward Shareable POSIX Signals

Daniel Colascione dancol@dancol.org
Sun Mar 11 18:56:00 GMT 2018


On 03/11/2018 11:07 AM, Zack Weinberg wrote:
> On Thu, Mar 8, 2018 at 12:52 PM, Daniel Colascione <dancol@dancol.org> wrote:
>> I've written up a proposal for improving the application signal APIs,
>> written below. Might there be any interest in prototyping this work
>> in glibc?
> 
> I want to say first of all that I think you have identified a real
> problem and I appreciate your having taken the time to write up a
> proposed solution.  

Thanks for taking a look.

> However, along with most of the other posters in
> this thread, I don't like the proposed solution -- and not just
> because I don't like signals (although, indeed, I do not like signals)
> but because I think the basic mechanism you suggest, chained handlers,
> is inherently unreliable and will cause more problems than it solves.

Why? You're ignoring the present reality that people _already_ use 
chained signal handlers. They're not going to stop. When libc 
maintainers reject widespread use cases as illegitimate, all they're 
doing is forestalling any sort of improvement.

A C runtime needs to consider realistic proposals to address real 
problems of real software --- not hold out instead for some idealized 
alternative family of APIs that will never materialize.

> We have had nothing but bad luck with mechanisms that rely on several
> user space components, not all maintained by the same people,
> cooperating in access to a shared resource. 

Resource arbitration is hard. It's even harder with the hacks (like 
ART's libsigchain) that people are forced to use today because libc 
authors don't consider signals legitimate somehow. Here, I propose an 
API that ensures that the right thing happens as long as everyone 
follows the rules, and that's far better than the lawless waste that 
exists today.

> Adding signal_register()
> to a universe that already has signal() also introduces a nasty
> compatibility problem: suppose library A uses the new API to register
> a handler for SIGINT (for example), but library B, or the application,
> calls signal(SIGINT, SIG_IGN), or sighold(SIGINT): what do you do?

My proposal specifically addresses this subject. Signal mask behavior is 
unchanged. Legacy signal hander installation behavior is unchanged. 
signal and sigaction affect the legacy signal handler slot, even when 
called with SIG_IGN.

> I also think you haven't gone deep enough into the root cause of the
> problem you're trying to solve.  You set out to make it possible to
> have more than one signal handler per process for each signal, but
> _why_ is that an undesirable limitation?  In most cases, it's because
> _signals are too coarse_.  When you get a SIGCHLD or a SIGIO or a
> SIGSEGV, you don't know which of many possible child processes / file
> descriptors / memory addresses is relevant. 

This claim is technically incorrect. The siginfo structure passed to the 
sigaction handler (and, in my proposal, to registered handlers) provides 
the necessary specificity. There are issues with merging asynchronous 
non-queued signals, but no such issues for synchronous signals like 
SIGSEGV, which cannot be queued.

> If we had a mechanism for
> dispatching _specific_ events in these categories directly to the code
> that cared about them, then we wouldn't need to have SIGwhatever
> handlers in the first place, and we also wouldn't need to worry about
> buggy or malevolent handlers eating events that were not for them.
> _That_ should be your goal.

You still need some way to register handlers of interest for specific 
events. You still need some kind of catch-all mechanism in case no 
specific handler is applicable. The overall shape of the API starts to 
resemble the one I proposed.

> With that in mind, let's run down the list of signals with their uses:
> 
> CHLD, PIPE, POLL/IO, URG, RTMIN through RTMAX -- These all represent
> I/O events.  In most cases it is already possible to receive a
> notification tied to the specific file descriptor that's relevant,
> instead.  The biggest gap I know about is that child processes are not
> represented by file descriptors, and this would be solved by adopting
> pdfork() (with some improvements).  

I would prefer process handle file descriptors. Linux upstream has 
specifically rejected process handle file descriptors on several 
occasions. I see no realistic path to solving this problem at that level.

> HUP, INT, QUIT, TERM, TSTP, TTIN, TTOU, WINCH, USR1, USR2, XCPU, PWR,
> ALRM, VTALRM, PROF -- Often is right to conceptualize these as I/O
> events as well, and many of them can already be turned into normal I/O
> (e.g. by putting the tty in raw mode, or by using timer_create instead
> of alarm), and for those that can't it should be made possible.  But
> another valid way to look at them is that they represent _broadcast_
> notifications that are already as fine-grained as they can be. 

I would consider SIGCHLD such a broadcast as well. "One of your child 
processes has died" is a perfectly bit of news to provide to the process 
as a whole.

ALRM, VTALRM, PROF --- as a completely separate matter, additional 
arbitration for coordinating timer deadlines would be useful.

> So,
> for these, I could be persuaded to support a multi-handler approach --
> but one in which all of the registered handlers are always called, no
> matter what.

Thanks for being receptive in this area. I understand your motivation 
for ensuring all such handlers are called for these broadcast signals. I 
think API uniformity matters more than ensuring that all handlers are 
called, especially since I'm certain that we need multi-handler support 
for synchronous signals as well as asynchronous ones, and synchronous 
signals need to be cancelable.

> I would need to hear a compelling answer to the
> coordination problem I mentioned above, though (what do you do if
> there are registered handlers and then someone else uses the legacy
> API to ignore the signal?)

Then no legacy handler is called, but the registered handler is. Any 
execved child inherits the SIG_IGN entry in the legacy slot, just like 
today, and none of the registered handlers.

> ILL, ABRT, FPE, SEGV, BUS, SYS, TRAP, IOT, EMT, STKFLT -- Synchronous
> signals arising from processor faults deserve a specialized mechanism
> all their own.  The notion I currently like, at the kernel level, is
> just-in-time instantiation of a ptrace monitor

That approach doesn't solve the arbitration issue and would make 
performance significantly worse than present. Not every instance of one 
of these synchronous signals is a crash. Spawning a process to handle 
them is far too expensive (and unreliable!) for something like a Java 
runtime's null pointer checks.

>, because that avoids
> the problem of recovering from memory corruption from within the
> corrupted address space.

Not every instance of these signals results in corrupted process state. 
In certain contexts, continuing to execute after receiving these signals 
is perfectly safe.

> At the C-library level, there are several
> plausible strategies for deciding whose responsibility a processor
> fault is: special ELF sections that label regions of code with
> handlers (like the except_table in the Linux kernel); dynamically
> registered annotations on memory regions; SEH; etc. 

While I would approve of adding SEH, I don't think it's a realistic 
option at the moment.

An except_table approach might solve part of the problem, but you'd need 
to provide a dynamic registration facility for the sake of JIT systems. 
I also don't think that keying handler _purely_ on program counter value 
is sufficient --- one might want to handle faults to a particular memory 
region or instruction type independent of precise code identity --- and 
for these use cases, a PC-keyed lookup table is inadequate. Besides, you 
still need a registered to be able to defer to the global process signal 
handler in case a SIGSEGV it receives really does represent a crash.

Also, think of how a table lookup would work at a mechanical level. The 
kernel would still push a SIGSEGV frame onto some stack and transfer 
control flow to the handler. Whether the handler is a set of chained 
user handlers as I propose or a libc-internal table lookup, you have the 
issues with asynchronous safety and memory corruption. My approach is 
just as safe and provides much greater flexibility.

You could teach the kernel to do the table lookup, but that's a much 
bigger task.

userfaultfd isn't adequate because it doesn't work on unmapped memory 
ranges, because it doesn't provide the values of registers in the 
faulting thread, and because there's no hope of making it a portable 
interface (because it's too powerful for limited systems).

> But notice that
> all of those can be built on top of "instantiate a ptrace monitor
> instead of delivering a fatal signal."  Someone would need to do
> something about how hard it is to write ptrace monitors, but that is
> technically a separate issue.
ptrace is far too slow. Besides, automatic ptrace-monitor creation 
suffers from the problem of unreliability in case we can't spawn a 
process (which can happen for any number of reasons) and conflicts with 
other processes ptracing the parent.

SIGSEGV, by contrast, with sigaltstack, is reliable and does not 
interfere with debugging (except to the extent that the debugger needs 
to be configured to ignore SIGSEGV).



More information about the Libc-alpha mailing list