Page MenuHomePhabricator

WE4.12.1 Establish a repeatable framework for establishing a “ground truth” for categories of abusive content, to support future automation efforts.
Open, HighPublic

Description

Context. Efforts by the PSI team to assess the feasibility of various models for detecting abusive content run into a consistent issue: that we lack purpose-built datasets that can be used to ground-truth their output. To this end, @cwylo is leading a sub-hypothesis under WE 4.12 to come up with a repeatable framework that would expedite the creation of future ground truth datasets, in various areas of anti-abuse work.

Description.
This framework is expected to:

  • Provide guidance on defining the types of edit to sample (e.g. by editor parameter; by edit type; by namespace; etc.)
  • Provide guidance on recommended size of the dataset
  • Lay out guidelines for who and how dataset evaluation is conducted
  • Provide recommendations for hosting, format etc. to facilitate evaluation
  • Outline recommended uses for the dataset as well as what not to use it for

Core deliverables will be:

  1. A reusable framework to guide ground-truth dataset creation
  2. A ground-truth dataset for vandalism on English Wikipedia, made using the framework above

Estimated Effort.
High due to fast turnaround. I'm told I can expect additional support from PSI team members as-needed.

Priority
High; a Q4 hypothesis. Completion of this KR will directly support other Q4 efforts in automating detection of abusive content, and serve as a foundation for projected work in FY26-27 Q1 and Q2. @kostajh who is the KR owner ideally wants this completed by the end of April.

Details

Due Date
Apr 30 2026, 5:00 AM

Event Timeline

DKumar-WMF assigned this task to cwylo.
DKumar-WMF triaged this task as High priority.
DKumar-WMF moved this task from Backlog to FY2025-26-Research-April-June on the Research board.
cwylo renamed this task from WE4 hypothesis on defining oversighting and vandalism to inform future work on automated detection to WE4.12.1 Establish a repeatable framework for establishing a “ground truth” for categories of abusive content, to support future automation efforts..Apr 3 2026, 4:46 PM
cwylo updated the task description. (Show Details)
cwylo set Due Date to Apr 30 2026, 5:00 AM.
cwylo added a subscriber: kostajh.
DKumar-WMF added a subscriber: cwylo.

We determined this was work for Applied Science. Miriam, reassigning to you as I'm not sure what the result of your discussion was with the WE4 team was on this, and whether to resolve this task or repurpose it.

kostajh added a subscriber: Tchanders.

Note that @Tchanders is leading the hypothesis for WE4.12 Content policy model evaluation and can help coordinate what happens next for this task (which will probably need to be rewritten)