We want to verify the effort/cost of collecting training and evaluation data for the top 20 languages.
The top 20 languages are English (en), French (fr), German (de), Spanish (es), Japanese (ja), Russian (ru), Portuguese (pt), Italian (it), Chinese (zh), Persian (fa), Polish (pl), Arabic (ar), Dutch (nl), Ukrainian (uk), Romanian (ro), Hebrew (he), Indonesian (id), Norwegian (no), Turkish (tr), Czech (cs).
In T386645, we use two approaches to collect fine-grained paragraph/sentence-level labels for peacock behavior data for English:
- finding reverted edits that mention [[WP:PEACOCK]] (a link to the Wikipedia: Manual of Style#Puffery page) in their reverting comments, then extracting the changed paragraphs from these edits.
- finding sentences marked with template:Peacock_inline.
To verify whether these approaches work for non-English languages and assess the available data, we need to check:
- if each language has its own Wikipedia: Manual of Style#Puffery page, and what are the abbreviations/shortcuts (e.g., "WP:PEACOCK"), and if communities actively use these abbreviations when reverting edits containing peacock language.
- if template:Peacock_inline exists in each language and what name is used for it.