How do I configure wcwidth?

Carlos O'Donell carlos@redhat.com
Fri Apr 10 21:44:32 GMT 2026


On 4/9/26 9:36 AM, Alan Mackenzie via Libc-help wrote:
> Hello, Libc-help.
> 
> I'm trying to enhance the Linux console to handle full UTF8.  At the
> moment, I'm using the Unifont 8x16 font.  It has both half-width glyphs
> (e.g. normal ASCII) and full-width glyphs (e.g. Asian languages and many
> symbols).
> 
> I can currently output these to the console, but the readline system
> (used by bash) corrupts some glyphs due to not knowing reliably how wide
> they are.
> 
> For example, U+2614 (UMBRELLA WITH RAIN DROPS) is a half-width glyph in
> Unifont, but readline thinks it is full-width.  By contrast, U+2622
> (RADIOACTIVE SIGN) is full-width in Unifont, but half-width in readline.
> 
> I think readline uses the glibc function wcwidth, the function that
> returns the width in column positions of its utf32 argument.  I would
> like to configure this function for Unifont.  I'd appreciate help on
> doing this.
> 
> I've glanced through some of the glibc source code, in particular
> ..../glibc-2.42/locale/programs/ld-ctype.c and I think the wcwidth
> configuration is part of the locale.  Can this be?
> 
> If so, the info manual ‘The GNU C Library Reference Manual’ describes
> how to select a locale and the effect of this, but doesn't seem to say
> how to modify a locale, or create a new one from scratch.
> 
> In summary, I'd like to get readline recognising the correct widths for
> Unifont glyphs.  Am I on the right track?

For UTF-8 the width data is derived directly from the Unicode release data.

In glibc we're using Unicode 17.0.0 right now.

glibc/localedata/charmaps/UTF-8

55568 % Character width according to Unicode 17.0.0.
55569 % Width is determined by the following rules, in order of decreasing precedence:
55570 % - U+00AD SOFT HYPHEN has width 1, as a special case for compatibility (https://archive.is/b5Ck).
55571 % - U+115F HANGUL CHOSEONG FILLER has width 2.
55572 %   This character stands in for an intentionally omitted leading consonant
55573 %   in a Hangul syllable block; as such it must be assigned width 2 despite its lack
55574 %   of visible display to ensure that the complete block has the correct width.
55575 %   (See below for more information on Hangul syllables.)
55576 % - Combining jungseong and jongseong Hangul jamo have width 0; generated from
55577 %   "grep '^[^;]*;[VT]' HangulSyllableType.txt".
55578 %   One composed Hangul "syllable block" like 퓛 is made up of
55579 %   two to three individual component characters called "jamo".
55580 %   The complete block must have total width 2;
55581 %   to achieve this, we assign a width of 2 to leading "choseong" jamo,
55582 %   and of 0 to medial vowel "jungseong" and trailing "jongseong" jamo.
55583 % - Non-spacing and enclosing marks have width 0; generated from
55584 %   "grep -E '^[^;]*;[^;]*;(Mn|Me);' UnicodeData.txt".
55585 % - "Default_Ignorable_Code_Point"s have width 0; generated from
55586 %   "grep '^[^;]*;\s*Default_Ignorable_Code_Point' DerivedCoreProperties.txt".
55587 % - Double-width characters have width 2; generated from
55588 %   "grep '^[^;]*;[WF]' EastAsianWidth.txt".
55589 % - Default width for all other characters is 1.

The processing of the Unicode release data is in:
glibc/localedata/unicode-gen/*

I think the next steps would be to ascertain if the Unicode data from the
release matches your expectations, and if it doesn't, why not?

Does that help?

-- 
Cheers,
Carlos.



More information about the Libc-help mailing list