How do I configure wcwidth?
Carlos O'Donell
carlos@redhat.com
Fri Apr 10 21:44:32 GMT 2026
On 4/9/26 9:36 AM, Alan Mackenzie via Libc-help wrote:
> Hello, Libc-help.
>
> I'm trying to enhance the Linux console to handle full UTF8. At the
> moment, I'm using the Unifont 8x16 font. It has both half-width glyphs
> (e.g. normal ASCII) and full-width glyphs (e.g. Asian languages and many
> symbols).
>
> I can currently output these to the console, but the readline system
> (used by bash) corrupts some glyphs due to not knowing reliably how wide
> they are.
>
> For example, U+2614 (UMBRELLA WITH RAIN DROPS) is a half-width glyph in
> Unifont, but readline thinks it is full-width. By contrast, U+2622
> (RADIOACTIVE SIGN) is full-width in Unifont, but half-width in readline.
>
> I think readline uses the glibc function wcwidth, the function that
> returns the width in column positions of its utf32 argument. I would
> like to configure this function for Unifont. I'd appreciate help on
> doing this.
>
> I've glanced through some of the glibc source code, in particular
> ..../glibc-2.42/locale/programs/ld-ctype.c and I think the wcwidth
> configuration is part of the locale. Can this be?
>
> If so, the info manual ‘The GNU C Library Reference Manual’ describes
> how to select a locale and the effect of this, but doesn't seem to say
> how to modify a locale, or create a new one from scratch.
>
> In summary, I'd like to get readline recognising the correct widths for
> Unifont glyphs. Am I on the right track?
For UTF-8 the width data is derived directly from the Unicode release data.
In glibc we're using Unicode 17.0.0 right now.
glibc/localedata/charmaps/UTF-8
55568 % Character width according to Unicode 17.0.0.
55569 % Width is determined by the following rules, in order of decreasing precedence:
55570 % - U+00AD SOFT HYPHEN has width 1, as a special case for compatibility (https://archive.is/b5Ck).
55571 % - U+115F HANGUL CHOSEONG FILLER has width 2.
55572 % This character stands in for an intentionally omitted leading consonant
55573 % in a Hangul syllable block; as such it must be assigned width 2 despite its lack
55574 % of visible display to ensure that the complete block has the correct width.
55575 % (See below for more information on Hangul syllables.)
55576 % - Combining jungseong and jongseong Hangul jamo have width 0; generated from
55577 % "grep '^[^;]*;[VT]' HangulSyllableType.txt".
55578 % One composed Hangul "syllable block" like 퓛 is made up of
55579 % two to three individual component characters called "jamo".
55580 % The complete block must have total width 2;
55581 % to achieve this, we assign a width of 2 to leading "choseong" jamo,
55582 % and of 0 to medial vowel "jungseong" and trailing "jongseong" jamo.
55583 % - Non-spacing and enclosing marks have width 0; generated from
55584 % "grep -E '^[^;]*;[^;]*;(Mn|Me);' UnicodeData.txt".
55585 % - "Default_Ignorable_Code_Point"s have width 0; generated from
55586 % "grep '^[^;]*;\s*Default_Ignorable_Code_Point' DerivedCoreProperties.txt".
55587 % - Double-width characters have width 2; generated from
55588 % "grep '^[^;]*;[WF]' EastAsianWidth.txt".
55589 % - Default width for all other characters is 1.
The processing of the Unicode release data is in:
glibc/localedata/unicode-gen/*
I think the next steps would be to ascertain if the Unicode data from the
release matches your expectations, and if it doesn't, why not?
Does that help?
--
Cheers,
Carlos.
More information about the Libc-help
mailing list