Page MenuHomePhabricator

Qualitative evaluation of the quality of recommendations in the Moderator Dashboard
Closed, ResolvedPublic

Description

Context
Under WE 1.3 New moderator homepage T402632, Diego has built T417442 WE1.3.5 Article similarity model (ASM). The Moderator Tools team now wants to see if patrollers find the recommendations from the article similarity model better (more engaging, more interesting, more urgent) than their existing methods for finding edits for review, in order to guide how the team implements the model and presents recommended edits for review, to end-users of the Personal Dashboard.

Description
This will be a moderated participant interview with a stimulus, in the form of a list of links to relevant edits, generated by the ASM. Our anticipated participant profile is:

  • Active patroller of their wiki (an editor who seeks out edits by others to review them for quality)
  • Speaks English

We plan to aim for about 4 participants from enwiki and possibly 4 participants from wikis other than enwiki; the latter group is dependent on our ability to generate a list of recommendations for non-enwiki projects.

Each session should take the form of a short interview, asking participants how they currently find edits for review and their experience with these workflows, then providing the participant with a list of ASM-generated edits and asking participants to compare their existing method against the suggestions.

The primary research goal of this study is to find out whether or not patrollers consider ASM-generated suggestions better than their existing workflow for finding edits to review. We also want to know what we should do in order to make ASM-generated recommendations "better" for patrollers (as compared to their existing workflows).

Final deliverable should be product recommendations for the Moderator Tools Team, along with feature suggestions for the ASM.

Estimated Effort.

  • For now, @Samwalton9-WMF + team will lead recruitment efforts.
  • @diego will need to provide a set of edits per-participant session
  • @cwylo will be responsible for designing the study, conducting it, and analyzing results.

Priority
High. The team anticipates the outcome of this work is needed to achieve their Q4 objective; these recommendations make up a large portion of the dashboard, and if users are not interested in it or don't view it as useful, they are unlikely to return to the dashboard.

Timeline
(See due date.)

Details

Due Date
May 15 2026, 5:00 AM

Event Timeline

cwylo updated the task description. (Show Details)
cwylo set Due Date to May 15 2026, 5:00 AM.
cwylo added a subscriber: Samwalton9-WMF.

Update:

  • I met with Sam and talked to Diego asynchronously to provide more details to the main ticket.

Update:

  • Discussion guide and outreach messages complete
  • We have started recruitment on Discord, with seven users contacted and three scheduled interviews for the next week
  • We have coordinated prototype/stimuli creation. Current aim is to provide 2-3 days lead time for recommendation creation
  • We may extend the interview by testing a second prototype during the session, using the same protocol. I don't expect this to require significant alteration of the discussion guide, but it's dependent on the progress of that prototype's creation. This is also a "nice-to-have", if it is unavailable we should proceed with the sessions regardless.

Update:

  • For this project I have completed 4 interviews with 3 more scheduled, so we are on track to complete 7 interviews this week
  • Follow-up emails to two other participants who are partially through the interview process
  • Recruitment closed for now as we've hit our targets for both populations

Update: The report is complete and has been shared to stakeholders. Waiting on a final meeting about next steps, scheduled for next week, before closing this ticket.

Broader share-out scheduled for next Wednesday; I'm marking this particular ticket as "resolved".