User story & summary:
As a member of the Growth team, I want to quickly understand whether newcomers are able to complete the Revise Tone task successfully so that I can evaluate task difficulty and identify areas for early improvement.
~2 weeks after the start of the Revise Tone A/B Test, we will review the following leading indicators.
We will use these indicators to determine if any immediate adjustments or follow-up investigations are needed before continuing to full evaluation. We will use this task to document the analysis and next steps.
Leading indicators
(DRAFT: these may be adjusted)
| ID | Indicator | Owner | Metric(s) for Evaluation | Plan of Action | |
|---|---|---|---|---|---|
| 1. | Task completion rate | Product Analytics | Proportion of users who start a Revise Tone task and successfully publish an edit | If completion is lower than the Copyedit Suggested Edit, investigate drop-off points in the workflow and adjust the UI or guidance to clarify next steps. | |
| 2. | Edit Revert rate | Product Analytics | Proportion of published Revise Tone edits reverted within 48 hours | If revert rate is significantly higher than baseline (unstructured copyedits), review reverted edits to identify recurring quality issues and adjust model thresholds or task guidance. | |
| 3. | Rejection rate | Product Analytics | Proportion of cases where users reject the suggestions | If high (>50%), it may indicate suggestions are irrelevant or too difficult. Review rejection reasons and determine if model improvements are needed, or if the task difficulty level needs to be changed. | |
| 4. | Community evaluation | Growth PM | Proportion of completed edits considered "good" by experienced editors | Since "unreverted" doesn't mean an edit is "good" we should gather community feedback on edit quality. If Revise Tone edits are considered bad more than 75% of the time, we should consider shifting the task to "Medium" so it's less likely to be completed by brand new editors. | |
| 5. | Model service availability | ML | Service Availability SLO: 95% of all requests return a 200/300/400 response | If SLO falls below 95%, prioritize infrastructure improvements or investigate failure points. | |