Page MenuHomePhabricator

One-off experiment analysis, funnel visualization, abandonment rate, updated impact metrics, and editor retention
Open, Needs TriagePublic

Description

Summary

Request for a one-off analysis updating and extending the experiment impact analysis completed in June (T429589). This covers four areas: updated core metrics with the latest data, a drop-off funnel visualization, a redefined abandonment rate, and editor retention by outcome group. All results should be broken down by editor experience level (junior vs. experienced) and platform (mobile vs. desktop) and pilot wiki.

Background

The previous one-off analysis (T429589) resolved several interpretation issues with the automated dashboard and gave us a clean read of the experiment results filtered to post-May 27 data. Since then, the experiment has continued running and new metrics for junior editors on mobile have been added to the dashboard. This analysis picks up where that one left off, adding funnel, abandonment, and retention work that was flagged as better suited to a one-off than the automated dashboard.

What we need

1. Core metrics

A repeat of the June one-off analysis with current data:

  • Absolute article counts per group: total articles created and total surviving articles (control vs. treatment)
  • Average articles created per user and average surviving articles per user
  • 30-day article survival rate with statistical significance
  • These are the main ones, but If we can replicate the metrics that are in the test kitchen automatic analysis, that'd be great.

As fo July 30th:

Survival rate shows no longer statistically significant in the test kitchen analysis.

30-day: +8.0%, p=0.068, 96.44% chance to win. Dropped from p=0.011.
15-day: +7.0%, p=0.105. No longer significant.
7-day: +6.5%, p=0.119. No longer significant.

This is the first time since the experiment stabilised that none of the survival rate windows are significant. The point estimates have also come down noticeably (30-day was +13.6% last week, now +8.0%). We would like to understand what could be driving both the significance drop as well as the survival rate drop since this started happening kind of suddenly around 27th of July.

2. Drop-off funnel visualization

A visual representation of user progression through the Article Guidance workflow steps:

  1. Red link > title is populated to Article guidance > Searching for a matching wikidata item → wikidata item shown
  2. Wikidata item (topic) selected / "pick a type instead" (manual topic selection)
  3. Sources validation step (user is blocked because they don't have valid sources, user adds valid soruces)
  4. Guidance tips shown → clicks on start writing
  5. Interaction with pre-populated structure (T432599)
  6. Article published

We want to understand where users are voluntarily dropping and where the flow is blocking them from continuing rightfully.

3. Abandonment rate redefined

Recalculate abandonment anchored to the exposure event (the earliest event logged for both groups, at red link click) rather than editing_started, to eliminate the denominator mismatch between treatment and control identified in T429589.

Exclude sessions blocked by notability restrictions from the abandonment count, as these represent the workflow functioning as intended rather than genuine editor abandonment (per Pau's comment in T429589).

Notability restrictions are defined by communities in the outlines, and are triggered to the user when they are hitting a notability restriction on their selected outline:

  • for that topic, you can only continue if it has an existing wikidata item
  • only if you provide 2 non-discourage sources
  • only if it's available in other language wikis

4. Editor retention by outcome group

How likely are editors to return in the 30 days following comparison between control and treatment. Regardless of how many attempts, the success of their article. I'd consider exluding editors who have dropped off voluntarily from treatment for this, what do you think?

Currently in the automated test kitchen analysis we see a negative trend for retention and we'd need to understand it fully:

As of July 30th: -10.2%, p=0.021, CI (-19.0%, -1.5%). This is the first time retention has reached significance. It seems treatment editors are less likely to publish a second article within 30 days. Are we interpreting this result correctly? could we be missing something here?

Notes

All data should be broken down by:

  • Editor experience (junior vs. experienced)
  • Platform (mobile vs. desktop)
  • Pilot wiki
  • All results should use post-May 27 data to exclude the instrumentation gap period.
  • Consider that PT wiki was paused in June 4th. that TR and Simple English experiment is exposed to experienced and junior editors and that FR wiki is only exposed to Junior editors.

Event Timeline

Note:

Editor retention > we can keep this simply to editors that have published through article guidance and returned for now. Control vs. treatment. Instead of the detailed breakdown in the task description.

@GGalofre-WMF
To refine the funnel analysis and edit abandonment rate calculation, I’d like to clarify which steps intentionally block the user from creating the article.

(1) Notability Check (action: notability_check_shown): Does showing this check block the user from proceeding, or is there a pathway to bypass it and continue to the next step?

(2) Other Blocking Points: Are there any other specific steps in the workflow where a user is explicitly blocked from moving forward?

How funnel drop-offs from notability and source restrictions work

There are two distinct mechanisms that can block a user from continuing the article creation flow. Both apply only to junior editors.

Mechanism 1: Notability restriction (Wikidata-based)

Some outlines have a community-set restriction that requires the topic to exist in Wikidata before a junior editor can proceed.

  • User searches for a topic with no Wikidata entry (e.g. "Gerard Galofre")
  • No results appear, they tap "pick a type instead" and choose "Actor"
    Captura de pantalla 2026-08-07 a la(s) 11.48.53.png (431×71 px, 10 KB)
    Captura de pantalla 2026-08-07 a la(s) 11.49.34.png (789×233 px, 25 KB)
  • The flow checks: does the Actor outline have a Wikidata notability restriction? Is the user a junior editor?
  • If both conditions are true, the user is blocked and cannot continue
  • This block is hard — there is no workaround
    Captura de pantalla 2026-08-07 a la(s) 11.47.19.png (787×359 px, 49 KB)

An experienced editor choosing the same outline for the same topic is not blocked, because they only fail one of the two conditions.

Mechanism 2: Source restriction

Some outlines require junior editors to add a minimum number of valid sources before they can continue.

  • User searches for a topic that is in Wikidata (e.g. "Bellvei," a human settlement)
  • They select the result and enter the flow
  • The source check requires them to add valid sources before the "Continue" button enables
  • If they add sources not recognised as valid, the button stays disabled
  • The block is intentional — they need real sources to proceed

Captura de pantalla 2026-08-07 a la(s) 12.00.15.png (756×786 px, 111 KB)

  • In practice this restriction is easy to bypass, because the list of explicitly invalid sources is small. Most sources are simply not in the system either way, so the check passes.

An experienced editor is not subject to this restriction at all.

Key distinction between the two

  • Notability block: hard.
  • Source block: soft. It can be bypassed by adding sources that are not flagged as invalid.

@GGalofre-WMF I've completed a one-off analysis updating and extending the experiment impact analysis completed in T429589. This provides some additional insights to supplement the overall experiment metrics available in the Test Kitchen automated analytics dashboard.

Core Metrics with Breakdowns by Wiki, Platform and Editor Experience

I recalculated all experiment metrics, including absolute article counts and breakdowns by wiki, platform and editor experience. Full results and filters are available in this google sheet.

Note: Many segmented views, particularly those filtered by mobile platform or specific wikis, have low sample sizes and have not reached statistical power. Please refer to Columns H (p-value) and I (95% Confidence Interval) in the spreadsheet to evaluate statistical significance before drawing conclusions based on these smaller subsets. We typically look for p < 0.05 to confirm statistical significance.

Results reflect experiment data recorded from 27 May 2026 through 10 August 2026.

This spreadsheet also includes the following updated metrics:

  • Article Abandonment Rate (Redefined): I updated the article abandonment rate to be anchored to the exposure event and to exclude sessions blocked by notability and/or source restrictions (as described in T432601#12194715).
    • With this updated definition, the control has an article abandonment rate of 72.16% and the treatment has an article abandonment rate of 76.10%. A statistically significant 5.5% increase in article abandonment rate. Even after removing sessions that were blocked, the workflow's additional steps create some friction that may cause users to drop off before saving their articles. The funnel analysis includes more insights on where these drop-off points occur.
  • Article Guidance editor retention (30 days; no survival check): Defined as the proportion of editors who publish a second article within 30 days of publishing their first. Does not take into account if the articles were deleted within 30 days of publication.
    • Note: This Aligns with how retention rate is currently defined in the Test Kitchen dashboard but I added some fixes to address issues I identified in the original query.
    • With these updates, we still see a decrease in editor retention (-5.94% relative change; 37.95% -> 35.70%).
    • This retention rate decrease may be because the new workflow slows down high-volume, low-intent repeat creation in favor of higher initial article quality.
  • Article Guidance editor retention (30 days; with survival check): Defined as the proportion of editors who publish a second article within 30 days of publishing their first. Requires that both articles are not deleted within 30 days of being published.
    • When we limit to non-deleted articles, the differences in retention rates between the control and treatment are minimal. There is a slight but non-statistically significant increase in overall editor retention rate when we limit to the retention of users who published an article that survived 30 days.
Drop-Off Funnel Visualization and Article Survival Rate Trends

Please see this notebook for details and findings from the drop-off funnel analysis. This notebook also includes a review of article survival rate (30-day) over time.

@MNeisler * thasks for this extensive report. I'll dedicate Friday to go through it.

Some initial questions:

1. Treatment vs. control funnel comparison
We want to compare the write_start to article_saved conversion rate between treatment and control — editors who arrive in VE via Article Guidance with pre-populated structure vs. editors who arrive at a blank page. This would help us isolate whether the article structure itself is helping people complete and save their article, separate from the filtering effect earlier in the funnel. Broken down by junior/experienced and mobile/desktop.

2. article saved vs. articles created discrepancy — and potentially adding survival to the funnel
We noticed a discrepancy between the funnel and the spreadsheet: the funnel shows ~600 articles saved for junior editors in treatment, but the spreadsheet shows 2,592 total articles created for the same segment. Since one red link click leads to one article creation attempt, we'd expect these numbers to be closer. Can you help us understand what each is counting and why they differ? Additionally, we could add a "survived 30 days" layer to the funnel so we can trace from articles saved through to survival outcome? Or any other way to understand how many of those 600 survive.

The select topic stage is the highest drop-off point for both Junior and Experienced Editors.

We see "exit at title" is the highest drop off point. Could we break down this drop between "title search" and "topic selection" by the search outcome people received, for example:

  • no wikidata results at all >>> manual topic selection VS. drop off
  • wikidata results that lead to "article already exists" >>> improve existing article / drop off
  • wikidata results with "no outline/guidance available" >>> drop off VS. "start without guidance" >>> VE >>> "publish" / abandon

We’re asking to learn whether the drop is because the tool can’t offer a suitable path, or because people aren’t choosing from the paths it offers.

@GGalofre-WMF I've provided some initial responses below but we can discuss on Monday as well.

  1. Treatment vs. control funnel comparison

We want to compare the write_start to article_saved conversion rate between treatment and control — editors who arrive in VE via Article Guidance with pre-populated structure vs. editors who arrive at a blank page. ...

This is conceptually very similar to how article abandonment rate is currently defined on the Superset dashboard.
In the dashboard, article abandonment rate is defined as the editing_start (the user entering the editor) to article_saved conversion rate.

(Note: It's not clear from the documentation if there is much of a difference between editing_start and the write_start step (user proceeded from the guidance flow into the article editor), except that we use write_start to define the treatment-specific-only funnel metric.)

Per the dashboard, 70.4% of control users abandon an article creation attempt after entering the editor vs. only 28.5% of treatment users (-59.5% relative decrease)

However, comparing these rates directly introduces significant selection bias (as mentioned in T429589#12063460).

  • Treatment Editors: To reach the editor via Article Guidance with a pre-populated structure, these users have already navigated multiple workflow steps. By the time they hit write_start, we have effectively filtered for high-intent editors.
  • Control Editors: Users arriving at the standard blank page have not passed through these filtering steps, meaning this group naturally includes lower-intent users (e.g., accidental clicks or curious browsing).

To accurately isolate the impact of pre-populated content itself, we would need an experiment setup where both groups reach the editor interface at the same moment or both groups going through the same prior steps.

  1. article saved vs. articles created discrepancy — and potentially adding survival to the funnel

Can you help us understand what each is counting and why they differ?

Yes, thanks for raising this! The difference here is due to a different approach needed to complete the funnel analysis, which resulted in needing to use a filtered sample of events.

In short, I'd recommend the funnel charts to understand relative drop-off locations across steps, but rely on the spreadsheet numbers (eg. 2,592 articles created by junior editors) for total volume results.

Here are the details on why they differ:

The article_saved event was not instrumented with a funnel entry token (which is what I used to link all events within a single article creation attempt). To mitigate this gap, I looked for article_saved actions performed by the same user ID in the treatment group within 2 hours of the write_start event. As a result, users who completed an article creation outside the 2-hour window were filtered out of the funnel analysis.

In addition, the funnel analysis reflects only attempts with a valid token. I needed to exclude some attempts where token tracking was broken, possibly due to client-side errors, timing issues, or unexpected browser navigation (e.g. page refreshes; browser back buttons).

I'll add a note to the report to clarify this discrepancy.

Additionally, we could add a "survived 30 days" layer to the funnel so we can trace from articles saved through to survival outcome? Or any other way to understand how many of those 600 survive.

Yes, this is feasible. However, because the article_saved event lacks a funnel_entry_token, we cannot directly link individual article creation attempts down to specific 30-day article survival events.

If it would help visualize drop-off, I can look to see what proportion of those ~600 articles survived 30 days.

However, I'd recommend using the Article Survival Rates and Total Survived Articles results from the spreadsheet for reporting, since those represent the most accurate counts.

As a recap of today's meeting, there seems to be 2 areas of work that could help us move forward:

1. Understanding the big drop of after title populated:

This relates to the question in this thread:

The select topic stage is the highest drop-off point for both Junior and Experienced Editors.

We see "exit at title" is the highest drop off point. Could we break down this drop between "title search" and "topic selection" by the search outcome people received, for example:

  • no wikidata results at all >>> manual topic selection VS. drop off
  • wikidata results that lead to "article already exists" >>> improve existing article / drop off
  • wikidata results with "no outline/guidance available" >>> drop off VS. "start without guidance" >>> VE >>> "publish" / abandon

We’re asking to learn whether the drop is because the tool can’t offer a suitable path, or because people aren’t choosing from the paths it offers.

@MNeisler you suggested to look at existing instrumentation here. I went through it and marked in green the actions I think are relevant that were not included in the funnel, it would look like these if I'm not mistaken:

  • subject_covered_shown: "article already exists" communicated to user > drop off
  • subject_covered_shown > subject_covered_action > improve
  • title_conflict_shown > drop off
  • title_conflict_shown > title_conflict_action > use_suggestion > continue
  • title_conflict_shown > title_conflict_action > view_existing
  • unsupported_subject_shown > drop off
  • unsupported_subject_shown > unsupported_subject_action > request_support
  • unsupported_subject_shown > unsupported_subject_action > start_writing

2. Data on deletion reasons for articles created via the AG workflow:

For articles created via article guidance that were deleted, a log of:

  • Article title
  • Outline used (if that info is available)
  • Deletion reason / code
  • x wiki
  • x junior / experienced

For articles created via article guidance that survived, a log of:

  • Article title
  • Outline used (if that info is available)
  • x wiki
  • x junior / experienced

I think both of these are big enough for separate tasks. If you confirm these are feasible I will create them separately :)

Drop-Off Funnel Visualization and Article Survival Rate Trends

Please see this notebook for details and findings from the drop-off funnel analysis. This notebook also includes a review of article survival rate (30-day) over time.

This analysis is great!
Sharing some notes with my initial impressions (sorry if some of this has been already discussed):

  • Survival rate metric seems to be more volatile than I was anticipating. We may need to (a) understand why and (b) find a more reliable way to measure it. For example, look at results by type of topic may give some clues. Given that outlines cover a specific set of topics, and the control group covers any topic, I wonder how numbers look for the same types of articles.
  • For French, the notability check seems to apply much less (0.6%) than for other languages (~10%), and the survival rate results are worse for French. We may want to check if those aspects are connected: identify which are notability restrictions on key article types that are missing in French compared to other languages, and consider adjusting the outlines accordingly to observe if there is any change in the results.
  • “Junior editors experience substantial drop-off during the drafting process, with 40.7% of editors abandoning their draft.” This seems to be a useful reference metric to improve with the “Guidance while editing” interventions.
  • We may want to distinguish "intended drop-off" from the not intended one. Catching notability or reference issues early is not a problem and serves the goal of avoiding potentially irrelevant articles to be created on the wiki.
  • The main drop-off takes place at topic selection. We may want to research and/or experiment by looking at it from multiple angles since there can be multiple reasons behind this. I wonder which would be the effect of different interventions like: (a) increase outline coverage for more topics, (b) surface “types of articles” as part of the list of results too, (c) allow for searching for “similar” articles, (d) generic outline…

Additional notes and comments (minor observations):

  • “Only about 6% of all sessions were shown the notability check.” Does this mean that 6% of articles are prevented from creation due to not meeting the notability criteria? It may be useful to identify the types of articles. This is information that could be surfaced in the outline pages to illustrate the impact of a given restriction.
  • “Experienced editors are nearly 1.7x more likely to complete the workflow and save their article.” I wonder how this proportion looks like for the default "starting from scratch" experience.
  • “Junior editors represent the vast majority of total workflow entry volume (6,366 sessions vs. 2,791 sessions. This is expected as the French Wikipedia experiment is only exposed to Junior editors.” On French, is the control group also only counting junior editors? Otherwise we may be comparing junior editors (experiment) vs. all editors (control).
  • “There 118 sessions by experienced users where a notability check was logged as being shown. This does not seem expected…” Why is not expected? This depends on the specific notability restrictions used in the outlines. Outlines can scope the restrictions to only juniors by adding the “junior” keyword, but this may not be the case in all outlines.
  • “Once users reach Write Start, 69.5% of Desktop writers go on to save their article (913 / 1,314), compared to 60.9% on Mobile (142 / 233),” I wonder if mobile completion correlates with the length of the initial contents of the outline. That is, having the editor pre-filled with an amount of content that feels too much work to edit on mobile. This can inform how to support a more gradual addition of contents when providing guidance while editing.
  • “junior editors add a valid source in 16.4% of their article attempts (vs. 11.8% for Experienced), but also encounter higher error/drop rates at this step (1.3% invalid sources, 2.6% exits).” Does the “1.3% invalid sources” represent sources that were listed as discouraged by the community? If that is the case, we should present this as a success, not an error/problem. The community captures commonly problematic references and the system identifies them early and educates users about them before they reach the article.