[PATCH] cs_CZ locale: fix collation [BZ #22336]
Mike FABIAN
mfabian@redhat.com
Mon Nov 27 06:49:00 GMT 2017
[BZ #22336]
* localedata/locales/cs_CZ (LC_COLLATE): Use âcopy "iso14651_t1"â
and implement the collation rules for cs from CLDR on top of that.
* Makefile: Add cs_CZ.UTF-8 to test-input and to the list
of locales to be built for testing.
* cs_CZ.UTF-8.in: New file with test data to test the Czech sorting.
Difference in sorting of the added test file (added in the patch)
before and after applying the sorting changes of the patch:
$ diff -u cs_CZ.UTF-8.in.old cs_CZ.UTF-8.in
--- cs_CZ.UTF-8.in.old 2017-11-24 16:47:52.688348458 +0530
+++ cs_CZ.UTF-8.in 2017-11-24 16:42:08.364221944 +0530
@@ -1,7 +1,3 @@
-È¥
-Ȥ
-Ê
-Æ·
a
a
a
@@ -65,7 +61,6 @@
cenných
cenným
cenou
-cH
cvrÄek
cz
cZ
@@ -94,8 +89,9 @@
H
hruška
ch
-CH
+cH
Ch
+CH
chÅestýšům
ChÅestýšům
chÅipka
@@ -188,6 +184,8 @@
Z
ź
Ź
+È¥
+Ȥ
za
Za
źa
@@ -209,6 +207,8 @@
Žb
žluva
Žluva
+Ê
+Æ·
0
1
1
I think "cH" was sorted completely wrong before and "CH" slightly
wrong as well. So this patch seems to not only base the Czech
LC_COLLATE implementation on the iso14651_t1 file as requested in this
bug but also improves the sorting of the uppercase/lowercase variants
of the ch digraph.
And of course it improves the sorting of some non-Czech characters
like Ê and È¥ because these were not handled at all in the old
Czech LC_COLLATE implementation.
--
Mike FABIAN <mfabian@redhat.com>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: 0001-cs_CZ-locale-Base-collation-on-iso14651_t1-BZ-22336.patch
Type: text/x-patch
Size: 88965 bytes
Desc: not available
URL: <http://sourceware.org/pipermail/libc-alpha/attachments/20171127/d276c500/attachment.bin>
More information about the Libc-alpha
mailing list