Summary
Request for a one-off analysis updating and extending the experiment impact analysis completed in June (T429589). This covers four areas: updated core metrics with the latest data, a drop-off funnel visualization, a redefined abandonment rate, and editor retention by outcome group. All results should be broken down by editor experience level (junior vs. experienced) and platform (mobile vs. desktop) and pilot wiki.
Background
The previous one-off analysis (T429589) resolved several interpretation issues with the automated dashboard and gave us a clean read of the experiment results filtered to post-May 27 data. Since then, the experiment has continued running and new metrics for junior editors on mobile have been added to the dashboard. This analysis picks up where that one left off, adding funnel, abandonment, and retention work that was flagged as better suited to a one-off than the automated dashboard.
What we need
1. Core metrics
A repeat of the June one-off analysis with current data:
- Absolute article counts per group: total articles created and total surviving articles (control vs. treatment)
- Average articles created per user and average surviving articles per user
- 30-day article survival rate with statistical significance
- These are the main ones, but If we can replicate the metrics that are in the test kitchen automatic analysis, that'd be great.
As fo July 30th:
Survival rate shows no longer statistically significant in the test kitchen analysis.
30-day: +8.0%, p=0.068, 96.44% chance to win. Dropped from p=0.011.
15-day: +7.0%, p=0.105. No longer significant.
7-day: +6.5%, p=0.119. No longer significant.
This is the first time since the experiment stabilised that none of the survival rate windows are significant. The point estimates have also come down noticeably (30-day was +13.6% last week, now +8.0%). We would like to understand what could be driving both the significance drop as well as the survival rate drop since this started happening kind of suddenly around 27th of July.
2. Drop-off funnel visualization
A visual representation of user progression through the Article Guidance workflow steps:
- Red link > title is populated to Article guidance > Searching for a matching wikidata item → wikidata item shown
- Wikidata item (topic) selected / "pick a type instead" (manual topic selection)
- Sources validation step (user is blocked because they don't have valid sources, user adds valid soruces)
- Guidance tips shown → clicks on start writing
- Interaction with pre-populated structure (T432599)
- Article published
We want to understand where users are voluntarily dropping and where the flow is blocking them from continuing rightfully.
3. Abandonment rate redefined
Recalculate abandonment anchored to the exposure event (the earliest event logged for both groups, at red link click) rather than editing_started, to eliminate the denominator mismatch between treatment and control identified in T429589.
Exclude sessions blocked by notability restrictions from the abandonment count, as these represent the workflow functioning as intended rather than genuine editor abandonment (per Pau's comment in T429589).
Notability restrictions are defined by communities in the outlines, and are triggered to the user when they are hitting a notability restriction on their selected outline:
- for that topic, you can only continue if it has an existing wikidata item
- only if you provide 2 non-discourage sources
- only if it's available in other language wikis
4. Editor retention by outcome group
How likely are editors to return in the 30 days following comparison between control and treatment. Regardless of how many attempts, the success of their article. I'd consider exluding editors who have dropped off voluntarily from treatment for this, what do you think?
Currently in the automated test kitchen analysis we see a negative trend for retention and we'd need to understand it fully:
As of July 30th: -10.2%, p=0.021, CI (-19.0%, -1.5%). This is the first time retention has reached significance. It seems treatment editors are less likely to publish a second article within 30 days. Are we interpreting this result correctly? could we be missing something here?
Notes
All data should be broken down by:
- Editor experience (junior vs. experienced)
- Platform (mobile vs. desktop)
- Pilot wiki
- All results should use post-May 27 data to exclude the instrumentation gap period.
- Consider that PT wiki was paused in June 4th. that TR and Simple English experiment is exposed to experienced and junior editors and that FR wiki is only exposed to Junior editors.



