Context. Efforts by the PSI team to assess the feasibility of various models for detecting abusive content run into a consistent issue: that we lack purpose-built datasets that can be used to ground-truth their output. To this end, @cwylo is leading a sub-hypothesis under WE 4.12 to come up with a repeatable framework that would expedite the creation of future ground truth datasets, in various areas of anti-abuse work.
Description.
This framework is expected to:
- Provide guidance on defining the types of edit to sample (e.g. by editor parameter; by edit type; by namespace; etc.)
- Provide guidance on recommended size of the dataset
- Lay out guidelines for who and how dataset evaluation is conducted
- Provide recommendations for hosting, format etc. to facilitate evaluation
- Outline recommended uses for the dataset as well as what not to use it for
Core deliverables will be:
- A reusable framework to guide ground-truth dataset creation
- A ground-truth dataset for vandalism on English Wikipedia, made using the framework above
Estimated Effort.
High due to fast turnaround. I'm told I can expect additional support from PSI team members as-needed.
Priority
High; a Q4 hypothesis. Completion of this KR will directly support other Q4 efforts in automating detection of abusive content, and serve as a foundation for projected work in FY26-27 Q1 and Q2. @kostajh who is the KR owner ideally wants this completed by the end of April.