Wikidata/Notes/Normalization: Difference between revisions
Appearance
Content deleted Content added
Закрываем проект Tag: Reverted |
restore rev 19958050 (2020-04-06T04:47:57Z) by ~riley Tag: Undo |
||
| Line 1: | Line 1: | ||
| ⚫ | It seems like there are problems with several charsets that will be part of the project. At least there are problems with some comparisons in [[w:en:Malayalam|Malayalam]] (note the url and title at [//ml.wikipedia.org/wiki/%E0%B4%B9%E0%B5%88%E0%B4%A1%E0%B5%8D%E0%B4%B0%E0%B4%9C%E0%B4%A8%E0%B5%8D%E2%80%8D this page] and compare to url and title at [//ml.wikipedia.org/wiki/%E0%B4%B9%E0%B5%88%E0%B4%A1%E0%B5%8D%E0%B4%B0%E0%B4%9C%E0%B5%BB this page]) and [[w:en:Arabic language|Arabic]], with sort orders in Arabic, Persian and Hebrew, and with composition in Bangla. |
||
[[File:Woodpecker10.jpg|5000px]]<br> |
|||
| ⚫ | |||
There are several places where we do run into trouble due to this |
There are several places where we do run into trouble due to this |
||
Latest revision as of 17:54, 5 March 2026
It seems like there are problems with several charsets that will be part of the project. At least there are problems with some comparisons in Malayalam (note the url and title at this page and compare to url and title at this page) and Arabic, with sort orders in Arabic, Persian and Hebrew, and with composition in Bangla.
There are several places where we do run into trouble due to this
- Adding and removing sitelinks as the page names must match the page names reported by the sites and also must match the strings in the items
- Adding and removing aliases in specific languages in the items as the strings must match
- Adding and removing label-description pairs in specific languages in the items as the strings must match
- Lookup of items according to site-page pairs in sitelinks in secondary storage
- Lookup of aliases in specific languages in secondary storage
- Lookup of label-description pairs in specific languages in secondary storage
See also mw:Unicode normalization considerations.
Directory /includes/normal/
[edit]This directory contains some Unicode normalization routines. See includes/normal/README for more information.