Improved check-localedef script
Mike FABIAN
mfabian@redhat.com
Fri Aug 4 09:50:00 GMT 2017
Rafal Luzynski <digitalfreak@lingonborough.com> wrote:
> 4.08.2017 11:14 Mike FABIAN <mfabian@redhat.com> wrote:
>> But even though U+20AC cannot be converted to ISO-8859-1, the
>> ca_ES.ISO-8859-1 locale still works because it is transliterated:
>>
>> $ LC_ALL=ca_ES locale -k currency_symbol charmap
>> currency_symbol="EUR"
>> charmap="ISO-8859-1"
>>
>> So this does not cause an actual problem.
>
> So the "â¬" character is actually representable in ISO-8859-1 because
> we can convert it to "EUR". Looks like a false positive then.
Yes.
>> The ca_ES source file is not ASCII, it has
>>
>> % catalÃ
>> lang_name "<U0063><U0061><U0074><U0061><U006C><U00E0>"
>>
>> So maybe I could just convert the file to UTF-8
>> and change â% Charset: ISO-8859-1â into â% Charset: UTF-8â
>> to get rid of the check-localedef warning.
>>
>> Would that be OK?
>
> I think that no, it's not OK. If I understand correctly the
> "source file is ASCII" sentence means that the individual characters:
> '<', '2', '0', 'A', 'C', '>' are ASCII.
Yes.
> They may describe something more complex like <U00E0>. But even this
> is not UTF-8 because UTF-8 would be <C3> <A0> (UTF-8 is 8-bit). The
> closest charset would be UCS-2 or simply a generic Unicode.
My understanding at the moment is that the â% Charset: ...â comment
indicates the encoding used to write the source file. So something like
â<U20AC>â is definitely ASCII. Non-ASCII stuff in locale source files
seems to exist only in comments at the moment.
--
Mike FABIAN <mfabian@redhat.com>
More information about the Libc-alpha
mailing list