Unicode security
Joseph Myers
joseph@codesourcery.com
Mon Jan 10 21:07:33 GMT 2022
On Mon, 10 Jan 2022, Reini Urban via Binutils wrote:
> So sooner or later some ELF/COFF/bla header field will be needed to state
> the obvious:
> name is UTF-8.
I think that's a matter for the ELF gABI document, where it describes the
*C binding to ELF* (where it says "External C symbols have the same names
in C and object files' symbol tables." - which says nothing about
encoding, since the point of that statement in its historical context was
probably to imply "no leading underscores added"). That is, ELF symbols
are arbitrary 0-terminated octet sequences (ELF is not limited to C
objects, symbols only need to be interpreted when included in diagnostics,
the assembler and linker should allow you to work with objects with
arbitrary 0-terminated octet sequences for symbols if you want to, just as
you can use ASCII characters in ELF symbols that aren't valid in C
identifiers), but when ELF is used for objects compiled from C, those
octet sequences for C identifiers with external linkage need to be
interpreted in a particular way.
Maybe such a change could be proposed on the generic-abi list once Cary's
public repository is available.
I attempted to get such a statement about UTF-8 encoding of ELF symbols
for C identifiers with external linkage into the gABI in January 2005, and
was directed to the ia64-abi@unix-os.sc.intel.com mailing list for that
purpose. At that time, the gABI maintainers on that mailing list weren't
willing to accept such a change, but maybe the current maintainers on the
current mailing list would be. If it's still not acceptable for the gABI,
then an operating-system-specific ABI would be the place to go.
DWARF (version 3 and later) does have DW_AT_use_UTF8.
--
Joseph S. Myers
joseph@codesourcery.com
More information about the Binutils
mailing list