wctomb() accepts out-of-range character in C-locale

Bruno Haible bruno@clisp.org
Mon Mar 25 11:26:00 GMT 2024


Hi Corinna,

> Jun T wrote:
> > ---------------------------------------
> > #include <stdio.h>
> > #include <stdlib.h>
> > #include <locale.h>
> > 
> > int main() {
> >     char buf[MB_CUR_MAX];
> >     setlocale(LC_ALL, "C");
> >     printf("%d\n", wctomb(buf, 0x80));
> >     return 0;
> > }
> > ---------------------------------------
> > 
> > On Linux it outputs '-1'.

"On Linux" is ambiguous:
  - In glibc, it outputs -1 because of this glibc bug:
    https://sourceware.org/bugzilla/show_bug.cgi?id=19932
    https://sourceware.org/bugzilla/show_bug.cgi?id=29511
  - In musl libc, it outputs -1 because the "C" locale (like all locales)
    uses UTF-8 encoding and the lone byte "\x80" is not an entire character
    in UTF-8.

> > But a wide character >= 0x80 can't be converted into a valid
> > character in C-loccale (7bit), I think.

Err. "C" locale, a.k.a. "POSIX" locale, is not 7-bit but 8-bit.
Quoting https://pubs.opengroup.org/onlinepubs/9699919799.2018edition/basedefs/V1_chap06.html#tag_06_02 :
  "The POSIX locale shall contain 256 single-byte characters ..."

> During testing I found that gnulib was replacing various functions built
> into Cygwin for several reasons, and one of them was that the conversion
> of wide char to multibyte in the "C" locale was not transparently
> converting chars from 0x80 up to 0xff.

What you did is to make Cygwin POSIX compliant in this aspect, which is
good.

> I'm actually puzzled right now that this doesn't work in GLibc either.

It's the aforementioned glibc bug.

> Do you have an idea what gnulib configure test might have been the
> trigger for the above revert?

It's the "checking whether the C locale is free of encoding errors..." test
(macro gl_MBRTOWC_C_LOCALE in m4/mbrtowc.m4).

Bruno





More information about the Newlib mailing list