Hi, according to this discussion, this task has been created. Please visit any category on ckbwiki which has member(s) with starting "ئ" (i.e. پۆل:ئەورووپا). And see the label on that category; "ء" is appeared instead of "ئ" while pages under that label are start with "ئ" (not "ء"). However, the character "ء" is not one of the characters in my language. We couldn't find a way to replace that character. So,we thought it is a bug in the system. Thanks!
Description
Details
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Resolved | BUG REPORT | Bawolff | T310051 Incorrect category header "ء" needs to be "ئ" instead (on ckbwiki) | ||
| Resolved | PRODUCTION ERROR | jhsoby | T390142 PHP Deprecated: Use of CollationCkb::__construct was deprecated in MediaWiki 1.44. [Called from Wikimedia\ObjectFactory\ObjectFactory::getObjectFromSpec] |
Event Timeline
@AramBakir: Thanks for reporting this. Please always use the bug report form (linked from the top of the task creation page) to create a bug report, and fill in the sections in the template, so it is way easier to read what happens where and what you expected to happen instead. Thanks a lot.
ء is a Hamza (glottal stop) (U+0621), usually to be combined with another letter (in some languages). I wonder if the parser somehow has issues with combining the Hamza with یresulting in ئ.
@Aklapper : I didn’t really think the problem was big enough to fill out the entire form so I wrote the report in paragraphs.
Please see تصنيف:بلدان أوروبية غربية on arwiki; Hamza poses no difficulty in classifying the members of that category.
https://gerrit.wikimedia.org/g/mediawiki/core/+/HEAD/includes/collation/CollationCkb.php might be relevant here: We fall back to Farsi.
Hmm, I don't see any bullet point that starts with ئ on that arwiki page? If you see a specific string, please post the specific spring and always be explicit - thanks!
@Aklapper, I don't think there's a problem with my example. Let me explain further. I didn't say that members of the Arabic Wikipedia category should start with ئ or that it should start like this. I also mentioned that example from our discussion on Wikipedia:Village pump (miscellaneous) because a user thought that the letter Hamza ء was mixed with ئ and for some reason only the Hamza was appeared. And I said if that's a bug in the system, then Arabic Wikipedia has the letters ا, إ, أ and should appear ء there instead of those letters, but as you can see each article is neatly placed under its own label (not ء). I didn't say the Arabic category label should start with ئ. If you still don't understand why I mentioned that Arabic Wikipedia category as an example, please read the last three comments here. Although I doubted that fawiki would be involved. However, this is a problem and must be fixed. Thank you for taking care of this task.
Makes sense; thanks for the explanation. The bigger question is if this also happens in Farsi (because of the collation fallback) instead of Arabic.
@Aklapper, Thank you for your understanding. I think there is no problem with fawiki. See an example: رده:افراد زنده
Hi @Aklapper, I don't know if you remember this task, but I installed MediaWiki software on my computer yesterday. I noticed that what we see on the Central Kurdish Wikipedia in this case I did not see at all on my installed Mediawiki. That is, what I saw in my personal Mediawiki shows me exactly what I expected.
I thought the problem was with ckbwiki itself, not MediaWiki, so I searched a lot in ckbwiki's local scripts, but couldn't find anything affected this case.
Why this MediaWiki permanent link (https://www.mediawiki.org/w/index.php?title=Manual:$wgCategoryCollation&oldid=5047484) specifically discusses ckb?
I really feeling confused! What do you thinking?
Without clear steps to reproduce, version information, operating system information, browser and browser version, any reply would only be random guessing.
Why this MediaWiki permanent link (https://www.mediawiki.org/w/index.php?title=Manual:$wgCategoryCollation&oldid=5047484) specifically discusses ckb?
Maybe "View History" of that page has some hints.
Sorry, I don't know exactly what to do. I did a search through Phabricator and Gerrit and the @matmarex came up a lot around this case. I hope he has information on this.
The CKB collation is inheriting from the farsi collation.
In the farsi collation, U+0621 Arabic Letter Hamza ء and U+0626 Arabic Letter Yeh with Hamza Above ئ are considered to be only secondarily different (Similarly, أ ؤ إ ئ ىٕ are also only considered secondary different). For a group of characters that are only secondarily different, we can only have one character represent the group as a header for that section in the category listing.
In farsi, the representitive character of these group of characters is ء (U+0621). We could make it be ئ (U+0626) instead for ckb, however that would apply to all members of this group. Is that what is desired here?
[Looking back at the history of this, it looks like I was the one who made this choice originally in a075f0de2873 by guessing... I definitely don't speak any of these languages]
cc'ing @Ladsgroup if as a farsi speaker he has any thoughts on what the proper behaviour here is on the farsi side of things
What bothers me about using HAMZA as a category heading is that it never was part of the Persian alphabet, Yeh is part of the alphabet maybe all of them should go under Yeh? I'd summon @Ebrahim who knows much more about unicode than I do.
Words starting with ئ (which none exist in Persian and this word is just incorrect and imaginary) are categorized under ء also in Persian Wikipedia,
So if ckb inherits from fa, let's see what we want for Persian first of all in hope our fix can help ckb collation eventually also and if that fails to go forward then a break of ckb-fa inheritance should be proposed perhaps,
According https://collation-charts.org/icu442/icu442-fa.html the following group exists for ء
"ئ" won't happen in Persian in words initials so what I'm going to suggest won't change much with Persian but assume instead of just categorizing the first letter, we wanted to create a tree for all the words, letter by letter, and when we reach to the third letter don't we want to put پائیز and پاییز and مؤسسه and موسسه nearby (they are semantically the same and used interchangeably for better or worse as one Persian native speaker can attest), instead of considering non existing words like پاءیز and مءسسه?
So I think it makes sense to break this group «ء أ ٲ إ ٳ ؤ یٔ ىٔ ئ» even for Persian and move «ی» like letters of this group to «ی» group and «ؤ» to «و» group given real world examples and that's fortunately also what OP needs and has requested for ckb also and if this plan fails to proceed, the fallback would be break fa-ckb inheritance I guess.
We don't really have the ability to break up groups (or at least, not very easily). We mostly just have the ability to chose which letter represents the group. (Or potentially a string like "ٲ - ئ" if that is better, as long as it starts with one of the letters in question)
[For historical context - the discussion on the cldr end on why they do it like they do is at https://unicode-org.atlassian.net/browse/CLDR-4207 ]
I thought no one cared about this—thank you, everyone! While ckb and fa are somewhat similar, there are still differences. By the way, ckb has the same issue with the letter Heh. For example, the word «ھەور» (cloud) begins with the letter «ھ» but is classified under the letter «ه».
I don’t see this issue on ckb Wiktionary, where the first letter directly becomes the header. However, the Wiktionary world may be different since it is character-sensitive.
So, assuming https://en.wikipedia.org/wiki/Kurdish_alphabets#Kurdo-Arabic_alphabet is accurate, and those are all the letters in ckb (Please confirm that is true), it looks like there is the following problems:
- ئـ (U+0626 Arabic Letter Yeh with Hamza Above followed by U+0640 Arabic tatweel) is filed under the header ء (U+0621 Arabic Letter Hamza[), which is what started this all
- The english wikipedia article does not list the letter ئ (U+0626 Arabic Letter Yeh with Hamza Above) as one of the Kurdish letters. Instead it lists the letter as ئـ (U+0626 Arabic Letter Yeh with Hamza Above followed by U+0640 Arabic tatweel). To clarify should the section header be just the ئ or should it be ئـ ?
- The digraph وو (U+0648 followed by U+0648) should be treated as a letter separate from و (U+0648) that comes between و (U+0648) and ۆ (U+06C6) with its own heading
- ی (U+06CC) and ێ (U+06CE) are currently considered variants of the same letter that both sort under ی (U+06CC), but should be considered separate letters with separate headings
- ھ (U+06BE Arabic Letter Heh Doachashmee) is sorted under heading ه (U+0647 Arabic Letter Heh) which is wrong. It should be under the heading ھ
- ە (U+06D5 Arabic Letter Ae) is sorted under ه (U+0647 Arabic Letter Heh) [These two characters look very similar in my font] which is incorrect, it should be sorted under ە (U+06D5 Arabic Letter Ae)
@Aram Is that a correct summary of the sorting issues on ckb?
So testing things out, i think the problem here is falling back to farsi. It seems like we get all the behvaiour we want if we fall back to the root collation (Except وو won't be treated as a digraph, but i assume that is minor. As a point of clarity, the root collation uses ئ as the header for ئـ and ئ. ) [We would still need to ensure that digitTransformLanguage is ckb and not "en" like it normally would be]
The only issue i see here is that it sorts ڕ before ز which seems like the wrong order according to the english wikipedia article. However it does give both their own section with their own heading. I'm not sure if that is a big deal or not.
@Aram In terms of a plan, does it sound good to you if we switch ckb wikipedia to use the root collation (uca-default) and then see where things stand?
p.s. Also, ckb wiki is currently set to sort numbers lexicographical and not numerically (e.g. The numbers 12, 5, 208 would be sorted 5, 208, 12 ). Should i also change this to numeric while i'm changing things?
I don’t see this issue on ckb Wiktionary, where the first letter directly becomes the header. However, the Wiktionary world may be different since it is character-sensitive.
ckb Wikipedia is set to use farsi ordering, where ckbwiktionary is set to use unicode codepoint ordering (Note on confusing terminology, this is not the same as uca-default which i mentioned above). It looks like this dates back to the original request at T54015 which was taken to be only for ckb wikipedia and not other ckb language projects. The farsi sort order is in particular noticeably different when it comes to the ordering of the characters پ چ ژک گ ه which probably have very different orderings in ckbwiktionary then ckb wikipedia.
@Bawolff, sorry for the late reply.
Yes, the Kurdo-Arabic alphabet section has all ckb letters and sorts them correctly. You can rely on it. Per this discussion, the community decided:
- ئ should not have U+0640 (Tatweel ـ).
- و should be used even if the title has وو.
- وو exists, but there's no single Unicode for it. We just use و twice, and it never starts a word.
- ێ should be under ی. ی and ێ are separate. ێ is a vowel and never starts a word.
- ھ should be under ھ, not ه. ه itself should go under ھ.
- ە (U+06D5 Arabic Ae) is a vowel and never starts a word. We're unsure how to handle it. If you agree, we can follow the default behavior.
ە (U+06D5) and ه (U+0647) look the same but are different. The community prefers ھ over ه for clarity, but some still use ه instead of ھ.
You're right: ڕ should come before ز, but it's currently wrong. Both need their own headers. For sorting, follow the English Wikipedia list but ignore ـ (Tatweel) after ئ.
p.s. Also, ckb wiki is currently set to sort numbers lexicographical and not numerically (e.g. The numbers 12, 5, 208 would be sorted 5, 208, 12 ). Should i also change this to numeric while i'm changing things?
Yes! The community wants numbers sorted as digits, not as text. [I always wondered why some numbers were wrong—this explains it!]
@Aram In terms of a plan, does it sound good to you if we switch ckb wikipedia to use the root collation (uca-default) and then see where things stand?
Honestly, I don’t know the outcome, so I can’t say for sure.
ckb Wikipedia is set to use farsi ordering, where ckbwiktionary is set to use unicode codepoint ordering (Note on confusing terminology, this is not the same as uca-default which i mentioned above). It looks like this dates back to the original request at T54015 which was taken to be only for ckb wikipedia and not other ckb language projects
So far, I’ve seen no issues on ckbwiktionary.
The farsi sort order is in particular noticeably different when it comes to the ordering of the characters پ چ ژک گ ه which probably have very different orderings in ckbwiktionary then ckb wikipedia.
Good point! I don’t know how Persian Wikipedia handles this, but I noticed something strange in Special:AllPages. Pages starting with پ, چ, or ژ and may be missed some letter yet—which should be earlier—are listed near the end. If this is relevant, we should fix ckb Wikipedia’s character order.
Please let me know if I missed anything, and thanks!
Change #1127472 had a related patch set uploaded (by Brian Wolff; author: Brian Wolff):
[mediawiki/core@master] collation: Add new collation uppercase-ckb for Central Kurdish
Test wiki created on Patch demo by Bawolff using patch(es) linked to this task:
http://patchdemo.wmcloud.org/wikis/15b06fb5e6/w/
Ok. Based on that description, i think its best we use a custom collation instead of the UCA based one (The custom one gives us more flexibility. The UCA one is more complex and might give better results for characters from other languages but it doesn't allow us to customize it as much. The big difference is that the UCA one allows more options for breaking ties, but i don't think that is super important for ckb. UCA might sort some obscure letters and foreign letters better).
I know these sorts of things are hard to talk about in the abstract, so I have created a test site where you can test this out, to see if it looks good https://patchdemo.wmcloud.org/wikis/15b06fb5e6/wiki/%D9%BE%DB%86%D9%84:Test Feel free to create some pages and add them to some categories to test out the new order.
Here is a screenshot (click to make bigger) of the new sort order.
Please note, for custom collations (which we are using) it only supports making the letters be equal in the same section or to be in separate sections. So if I put ێ in the same section as ی (which is what I did in the demo above) that means the sorting algorithm will treat both characters as having the same spot in the ordering. If one needs to come after the other in the sort order, we would need to put them in separate sections. (Of course if no words start with that letter, then you wouldn't see a section for it, but it still makes a difference if the character appears later in the word.)
When it comes for things like وو - lots of languages have character combinations where they sort differently if two specific characters are together, so we do have support for that. In the demo we are not treating وو as a single letter, but if they community changes their mind on that, let me know. If they do get treated as a separate letter, it would mean they would have their own section. It would also affect sorting slightly for words containing those letters. e.g. it would change if وو or وی comes first.
Good point! I don’t know how Persian Wikipedia handles this, but I noticed something strange in Special:AllPages. Pages starting with پ, چ, or ژ and may be missed some letter yet—which should be earlier—are listed near the end. If this is relevant, we should fix ckb Wikipedia’s character order.
Unfortunately we can't customize sort order on special pages.
For comparision sake, here are some of the other sorting options:
This is the root-UCA option. This is the other main choice.
This is what unicode code point order would be. This is what wiktionary uses (except also doing number sorting)
And for comparision, this is what is currently on ckb wikipedia
Anyhow, let me know what you think.
@Bawolff Thanks for creating the demo and for all your explanations. I tested the demo and noticed the following:
- Sorting numbers numerically works perfectly. However, in the current version, each number seems to have its own header, whereas in the demo, all numbers are grouped under a single header (٠-٩). I don’t see this as a major issue, but if the community prefers the previous format, how can we revert to it?
- On ckbwiki, headers currently appear in this order: numbers → all ckb alphabet → Latin alphabet (which aligns with the community's preference). However, in the demo, the order is: numbers → Latin → Arabic → ckb letters. We need ckb letters to be prioritized (after numbers and symbols), followed by Arabic and Latin alphabets If we can do so.
- Please check this query. These letters are vowels and never start a word, but they have their own articles. Do you think we need headers for these letters? If not, how should these four articles be categorized?
Here is a screenshot (click to enlarge) of the new sort order:
- In the screenshot, under the ھ header, ھ appears first, followed by ه, but in the demo, the order is reversed.
Unfortunately, we can't customize the sort order on special pages.
Unfortunately.
That's all from me—thanks again!
Its tied to number sorting. If we sort numbers in numerical order, then we have to put them under one header. We can change what the header is, but all the numbers have to go together when using number sorting. The only way to have separate headings for each number is to disable number sorting and use lexicographic order for numbers.
- On ckbwiki, headers currently appear in this order: numbers → all ckb alphabet → Latin alphabet (which aligns with the community's preference). However, in the demo, the order is: numbers → Latin → Arabic → ckb letters. We need ckb letters to be prioritized (after numbers and symbols), followed by Arabic and Latin alphabets If we can do so.
We can move specific letters later, but any letters that are left to their default ordering come before customized letters (So we can move the normal latin letters later, but variants like Á would likely still come first).
This is one of the downsides that the "custom" collations have that the UCA collations don't.
- Please check this query. These letters are vowels and never start a word, but they have their own articles. Do you think we need headers for these letters? If not, how should these four articles be categorized?
We can make these letters be part of their own section or as part of a different section (as is done in the demo).
I think making the vowels their own section probably makes the most sense. The only time it would come up would be for articles about them since they otherwise never start a word. However we can do whichever is preferred.
Here is a screenshot (click to enlarge) of the new sort order:
- In the screenshot, under the ھ header, ھ appears first, followed by ه, but in the demo, the order is reversed.
Right now we mark them as being the same. So if there are no other letters in the word to break the tie, then the order they come in is random (Its probably whichever article was created first comes first, but I'm not 100% sure).
Its tied to number sorting. If we sort numbers in numerical order, then we have to put them under one header. We can change what the header is, but all the numbers have to go together when using number sorting. The only way to have separate headings for each number is to disable number sorting and use lexicographic order for numbers.
I think we’re good with number sorting and the ٠-٩ header for all numbers.
We can move specific letters later, but any letters that are left to their default ordering come before customized letters (So we can move the normal latin letters later, but variants like Á would likely still come first).
This is one of the downsides that the "custom" collations have that the UCA collations don't.
For now, we don’t have better options for sorting non-ckb letters, so let’s keep the default behavior as it is without rearranging any letters.
We can make these letters be part of their own section or as part of a different section (as is done in the demo).
I think making the vowels their own section probably makes the most sense. The only time it would come up would be for articles about them since they otherwise never start a word. However we can do whichever is preferred.
Thanks for your input on vowels. Since we've decided to give them their own headers, let's move ێ into its own section.
Right now we mark them as being the same. So if there are no other letters in the word to break the tie, then the order they come in is random (Its probably whichever article was created first comes first, but I'm not 100% sure).
I initially thought the indexes determined the priority of letter order, but it seems they don’t affect the order within the same array. Since it's not about priority, we’ll leave them as they are.
I think we're about to wrap up this task. We truly appreciate your patience and all the effort you've put in. Thank you!
Test wiki created on Patch demo by Bawolff using patch(es) linked to this task:
http://patchdemo.wmcloud.org/wikis/dc0677df79/w/
Thanks for your input on vowels. Since we've decided to give them their own headers, let's move ێ into its own section.
Just to clarify, with putting vowels under their own header now, are we including وو in a separate header?
With the new changes, the only pairs of characters that are equal weight are: ک & ك and ھ & ه
Here is a new version of the test site you can test out https://patchdemo.wmcloud.org/wikis/dc0677df79/wiki/%D9%BE%DB%86%D9%84:Test (This is with وو as a separate header. Let me know if it should still be under و)
Just to clarify, with putting vowels under their own header now, are we including وو in a separate header?
No, please. وو should be under و.
With the new changes, the only pairs of characters that are equal weight are: ک & ك and ھ & ه
Exactly!
Test wiki created on Patch demo by Bawolff using patch(es) linked to this task:
http://patchdemo.wmcloud.org/wikis/e6153e3910/w/
Test wiki created on Patch demo by Bawolff using patch(es) linked to this task:
http://patchdemo.wmcloud.org/wikis/518f6d5d06/w/
Sounds good.
Here is a demo site with the final version of the sorting algorithm: https://patchdemo.wmcloud.org/wikis/518f6d5d06/wiki/%D8%AF%DB%95%D8%B3%D8%AA%D9%BE%DB%8E%DA%A9
From here, we have to wait for some Wikimedia internal approvals before it will be live on ckb wikipedia (This will probably take about 2 weeks, but could be more or could be less). Right now we are planning to only put the new sorting order on ckb.wikipedia.org, but if ckb.wiktionary also wants it, please let me know.
@Bawolff, I believe everything is ready, and we're just waiting for the update. As for ckb.wiktionary, I don't think it needs this configuration, as it’s more sensitive to individual letters/characters. Regarding number sorting, I’m not entirely sure, but I don’t think we need it either for now.
Thank you so much! I had no idea we had to make all these file edits to fix this!
Change #1127472 merged by jenkins-bot:
[mediawiki/core@master] collation: Add new collation uppercase-ckb for Central Kurdish
Change #1131651 had a related patch set uploaded (by Jon Harald Søby; author: Jon Harald Søby):
[operations/mediawiki-config@master] Change category collation for ckbwiki
Change #1131651 merged by jenkins-bot:
[operations/mediawiki-config@master] Change category collation for ckbwiki
Now that the change has been deployed to ckbwiki, I think there's nothing more to do here. Thanks @Bawolff!





