Despite its volume and value, navigating, retrieving and re-using visual content on Wikipedia is hard, due to the lack of labels, categories and metadata. Classification of this content for research and editing purposes is becoming increasingly important. Unfortunately, the value offered by its uniqueness comes with the disadvantage that common off-the-shelf classification models based on ImageNet give unsatisfactory results, requiring a custom solution.
The goal of this task is two-fold:
- develop a classification taxonomy to label images on Wikipedia and
- develop a model for image classification and embedding.
Meta page: https://meta.wikimedia.org/wiki/Research:Automated_Categorization_of_Wikipedia_Images