Pull together thoughts following the GLAM pilot of combining a purpose-driven edit-a-thon with newcomer task recommendations and discussions around ML + Equity.
Description
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Resolved | Isaac | T293516 Recommender Systems + Content Equity | |||
| Resolved | Isaac | T307254 Recommendation Equity: Findings from GLAM pilot and ML Equity Strategy |
Event Timeline
weekly updates: haven't worked on this much but did meet with PG from Language to discuss future of content translation recommendations and potential collaborations there
weekly updates: most of thinking in this space has been preparation for Mo's summer work around personalized edit recommendation and content equity (focus on SuggestBot). also in re-reading Diego's proposal for AI + Knowledge Integrity, he's covered the need to discuss data generation strategies with ML Platform and Product stakeholders so the planning for that (which we'd identifed as the main priority of next year in this space) can likely happen in collaboration with him (perhaps using vandalism detection as the case study, which is something I would have proposed anyhow).
weekly updates: met with JT and MR to discuss knowledge gaps / ml equity + product. shared general takeaways:
- For measurement of content impact, gender and geography gives good coverage: gender because interventions have been shown to be affective for getting editors to edit content about women so design choices can have a real impact. geography because interventions are less effective (editors are more likely to edit content with which they are familiar and geographic familiarity is a large component of this) so without measuring individual editor demographics, tracking content geography gives some insight into the diversity of the editor community and encourages long-term investments in supporting a more diverse editor community.
- For design: individual filters (e.g., topics, countries) are good and should continue to receive development but individual action won't close knowledge gaps. for that, we need collective action of the type organized by campaigns/edit-a-thons. so long-term, connecting recommender systems with campaigns feels like the much more effective approach.
aspects that came up and are good to continue to think about:
- how much vandalism do we see with these recommender systems? in what ways can we limit this, especially if we're encouraging connections with campaigns
- can we think of structured tasks that have equity impacts too -- e.g., alt-text for low-vision readers
- countries are generally how Product thinks about audiences so content data at the country-level is very useful
Weekly updates:
- Put together initial thoughts on next phase of this project around data gaps: https://meta.wikimedia.org/wiki/User:Isaac_(WMF)/Content_tagging/Data_gaps
- Tracking SuggestBot analysis/experiment work under this task: T310379
weekly updates:
- did a lot of thinking about suggestbot experimental design and revising our offline analyses to be more in-line with the proposed experiment. in particular, stuck on how to transform data we have for each editor receiving suggestbot recommendations (their edits that are attributable to suggestbot recs and broader edit history) to appropriately capture their flexibility when it comes to editing about various topics -- e.g., if they predominantly edit articles about men, would that affect the likelihood that they'd accept a recommendation to edit a biography of a woman? unfortunately even with gender (which is relatively simple), there are several challenges:
- not all articles are biographies so e.g., a feature that captures whether a recommendation matches an editor's preferences around biography gender (as gathered via edit history) doesn't distinguish between ambivalence about the gender of the biography and not editing biographies.
- for an e.g., editor that edits 40% women and 60% men, should we be more surprised if they accept a recommendation for a man or a woman biography? presumably this example editor prefers to edit about women but maybe they're just editing about a topic that has more women and they don't actually have a preference (or they're editing about sports and have a strong preference)?
- solution might be to not model it but try to capture via descriptive stats, which would probably also more easily capture editor variability in the flexibility of their preferences
- also working on summarizing current state of project with Mo before we decide on parameters for experiment
resolving: I updated slide deck with links to GLAM pilot. takeaways were country filters enabled events that otherwise couldn't have happened easily (importance of organizers focusing on content within their country) but the country-tagged articles in some countries were barely sufficient in Spanish Wikipedia when crossed with add-an-image recs to keep a large group busy without edit conflicts. definitely motivates a more extended country article set as keyword-based country filters (e.g., articles that mention Argentina) were far too broad. work continues on suggestbot task, in continued conversations with Product, and next year under ML Equity work.