User Details
- User Since
- Oct 1 2018, 2:19 PM (410 w, 2 d)
- Availability
- Available
- IRC Nick
- isaacj
- LDAP User
- Isaac Johnson
- MediaWiki User
- Isaac (WMF) [ Global Accounts ]
Today
The numbers look good to me. The table contains only page_ids not found in the event stream or in the event stream but content is null
@Snwachukwu ahh thank you! That clarifies then. So just an issue on my end. Back to the drawing board...
@Alaexis if you have any other details to add, feel free to do so here. If you could add your toolsadmin user name (Wikimedia Developer Account username under https://toolsadmin.wikimedia.org/profile/settings/accounts/), that might be helpful too for debugging.
Yesterday
@Snwachukwu I had been doing my own work with an Enterprise snapshot and found that it was missing a large number of articles (2.6M instead of the >7M that are on English Wikipedia). I just looked at yours and it also seems to be missing a lot (4.3M, which is still about 3M shy though better than mine). Before I investigate, I just wanted to check that you were expecting to have the full snapshot of English Wikipedia, right (and hadn't applied some filtering already to reduce it down)?
Looks good to me -- thanks @DDeSouza ! I'll follow up here when I post an announcement to wiki-research-l but resolving for now.
Thu, Aug 6
Also chemistry infoboxes are being triggered too because chembox is one of the classes: https://en.wikipedia.org/wiki/?curid=36155711
I'm realizing the the disambiguation boxes added via https://en.wikipedia.org/wiki/Template:Disambiguation are showing up as messageboxes too. They should just be notes. Might need to add a condition based on the role attribute to try to exclude them? I should look into how this works in other languages and with other ones like article-stub boxes.
Thank you! Just pasting a few images so it's easy for folks to quickly see what it would look like right now (when expanded).
Test wiki created on Patch demo by CMedelius-WMF using patch(es) linked to this task: https://cb6b54e86f.catalyst.wmcloud.org/wiki/Regent's_Park
Exciting @medelius ! How would I trigger the functionality?
Tue, Aug 4
Yikes, so many wiki pages out there! For a backfill, the Snapshot API from Enterprise should make that a lot faster and reduce the burden on the Parsoid APIs. See the list of current wikis + namespaces that they support: P95873. Essentially won't help with Commons/Wikidata or many user/discussion namespaces but good coverage of article namespaces otherwise. Unfortunately I don't think these snapshots are currently available on our infrastructure but you can see some tickets requesting this (T403298 and T305688) as well as a job that Search runs that downloads a smaller set of wikis from a different Enterprise API (T414066).
Thu, Jul 30
I was referencing the readme for the editing suggestions api schema. Is that what you need? I'm happy to document elsewhere, just gotta figure out where that would live.
That works. I can share the response format, which has the fields that would be needed. And then we can continue to work through the details of what target will look like in this context.
Wed, Jul 29
To the open questions, I'd also add one about how to handle situations in which it's believed that suggestions have been invalidated in bad-faith -- e.g., someone going through and rejecting every suggestion they see across many articles.
Noting the a more thorough research summary of WE1.9 KR is now published on Meta: https://meta.wikimedia.org/wiki/Research:Understanding_barriers_to_newcomer_retention
This is a really open-ended task so I'm going to close in favor of creating subtasks for the more specific needs. We already have a few documented: https://phabricator.wikimedia.org/maniphest/query/k43oTtaGizSY/#R
thanks for the additional context @AAlhazwani-WMF !
Tue, Jul 28
Excited to see this move forward! First question I guess is what makes a good recommendation. Given the focus on retention, I'd say three factors to pay attention to.
- Motivated to do: these are newcomers so editing is hard/confusing. They need some reason to make an attempt. This is going to be easier if they are familiar with the topic and even more so the closer they feel to being a relative expert of sorts on it. For example, folks know a lot about the country they live in, but so do many other people. Much better to find the most niche article that they care about (like their town/neighborhood), as they will have relatively more expertise to contribute and that can boost confidence. Things like trending/most-viewed can help motivate folks but these are extrinsic factors and it's going to be much more effective to get a newcomer to a topic that they intrinsically believe is important.
- Capable of completing: these are Suggested Edits so there's less of a burden on the newcomer to spot an issue themselves, but the smaller the article, the less context they need to work through to understand a potential edit. This also gets to the familiarity piece -- the more familiar you are with the topic, the easier to assess whether the suggested edit is a good one.
- Convince them to stick around: this is a product of how motivated they were to do the edit in the first place, how impactful the edit feels (there is a reason why add-an-image significantly boosts retention whereas add-a-link seems to have almost no effect despite boosting activation), and what happens after the edit. Regarding the last point, articles that are already higher quality or have more editors contributing to them are also probably places where a newcomer is more likely to trigger a revert. Lower-quality articles with fewer edits tend to be safer places to start.
Mon, Jul 27
Thanks for creating this task! Might be useful to try to think through all of the potential ways of representing a diff and potential models upfront so we can split up the work a little? My gut feeling is that we should first see how we can do from content alone as it should hopefully be enough alone to determine the larger category of edit that's being made. I'm sure the user features will boost offline model performance a bit as there are definitely user patterns, but they will also will build in some long-term fragility to the model I suspect because those are things that could easily change.
Fri, Jul 24
Loving the progress and I'm more than happy to help with evaluation post-launch! Some additional data / links for context:
- Details on what currently generates notifications: https://en.wikipedia.org/wiki/Special:DisplayNotificationsConfiguration
- Unfortunately the Special:ValidationStatistics page generated by FlaggedRevs only seems to show average time-to-review for the protection-mode wikis, which is often heavily skewed by a single unreviewed edit. But some data on median time-to-review from the larger FlaggedRevs wikis shows ranges from 5 minutes to 14 days. So in most cases, newcomers would have to revisit the article they edited (presumably a few times) before they'd notice their edit taking effect. I'll see if I can dig up more complete data though.
- I keep asserting that newcomers likely aren't putting pages on their watchlist when they edit existing pages because it's not the default action. I finally actually computed some data on that and looked at the first ten edits from editors on Wikipedia who created their account in the first three months of this year (query below). Only 25% of the time was the page that they edited on their current watchlist. This data allows for folks adding the page to their watchlist well after the edit or having added it but with an expiration that has passed, but I assume that's minimal. 25% of these newcomer edits were reverted (so the person got a notification regardless of watchlist add) but that still leaves 55% of edits that were neither reverted nor added to a watchlist, so editing and presumably hearing nothing about it :(
While the feature has been successful in reinforcing contributor motivation, its current implementation reflects its origins as a Growth MVP.
Tue, Jul 21
Fri, Jul 17
Thanks for this update @medelius !
Thu, Jul 16
Whoops -- just realized that they are in the HTML, they're just not being rendered. For example in the example with Chicago, the following lines appear:
Jul 13 2026
I forget what our policy on this is but there are three arXiv preprints (none of them via a conference though or intended to at this point as far as I know):
Jul 7 2026
Closing this task as FY25-26 has ended and the WE1.9 KR that housed much of this work has been accepted. One note: I don't include the work by Anbar for WE1.9 here around affiliates + mobile onboarding because this task is specifically about Research-team contributions, but it was quite important to the WE1.9 report as well.
Jun 26 2026
Jun 22 2026
Forgot to add this here but Applied Science component completed this past Thursday (snippets drafted and went through a round of review with the relevant team members) and handed back to Leila. I also did my best to add all the potential UX Research updates where I encountered them @DKumar-WMF so hopefully your job is a bit easier (though I'm sure I wasn't exhaustive).
Jun 18 2026
We never enabled access to the Homepage for legacy users, so the property should still be zero for the control group users (unless they enabled it themselves, ofc). Maybe we should do that sometime, but so far it didn't happen (and after all, if you survived despite no Homepage, it probably wouldn't give you too much usefulness).
Makes sense and useful for this data analysis too. Then all looks good from my perspective!
Jun 17 2026
Just removing myself as who this is assigned to as I haven't been working in this area for a bit.
Jun 16 2026
Just chiming in to say that I took a pass on the code and everything looks good with one clarifying question below. Thanks @Urbanecm_WMF for the work on this and @Bobocicada for close engagement/questions!
Jun 15 2026
@AJayadi-WMF there's a page with more details/examples of FeaturedFeeds: https://www.mediawiki.org/wiki/Extension:FeaturedFeeds/WMF_deployment. It looks like they might already have a feed for on-this-day (OTD) so perhaps already have a local Wikimedian who understands how the system works and could set it up for featured articles too? If not, you might try asking around internally at WMF to see if we have anyone who could advise on next steps for setting up.
Jun 12 2026
Even if the UI only ever uses one suggestion/action, modeling a 1 to many relationship will make it possible to choose.
Personally for the image/rewording example, this feels like part of the 10-20% that I don't think we should try to fit into the data model. That level of complexity would presumably require custom VE code to parse/display it anyways so why not just keep the data model simple and put that complexity in the code where it partially has to live anyways? So for these one-to-many type relationships, couldn't there just be multiple recommendations of the same type and we leave it to VE on how to merge them? To that end, Michael's suggestion about a confidence score field sounds reasonable and could allow for the data to convey their appropriate ranking while staying pretty flat/simple.
Jun 10 2026
that means each suggestion type would have its own API endpoint
I assume many API calls isn't good as you'd need to know which types of suggestions to fetch which feels like a lot of burden on the end-user clients?
Jun 4 2026
FYI this was completed under T289532: Add more languages to Wikipedia Clickstream! See https://dumps.wikimedia.org/other/clickstream/ starting with 2024-08.
Thanks for the suggestion Kai and putting together the task @MGerlach! I was curious so did some very quickly exploration of impact on the dataset. Code below for looking at pages that would meet the privacy threshold (10 referrals) in April on English Wikipedia: there were 76,261 distinct page titles, which is ~1% of the 5,840,736 total pages that made it into clickstream that month for enwiki. So not most at this point but still a reasonably large set of pages that in theory will continue to grow...
May 29 2026
verified -- thank you very much for the fix and also the two examples with cross-namespace and within-namespace!!
May 20 2026
Huh! In my revert spelunking in T423583, I learned that mw-reverted is only added on the 15 revisions prior to the reverting revision (or something like that).
Oh what a great task! Excellent details in there about revert mechanics (thank you). Good point that Mediawiki's heuristics are also semi-arbitrary so we're essentially choosing between two things right now:
- Mediawiki: Only look 15 edits back to find which edits have been reverted.
- Mediawiki History: Look back some amount of time to find potentially-reverted revisions. Currently all of time but now we're considering 90 days as an additional approach. I had advocated for 48 hours (the original proposal).
Thanks @xcollazo for this data and the details around back-filling the revision tag fields in T425573: mediawiki_history_incremental_v1: schema specification for stakeholder review. I mention that here because I think the ability to filter on mw-reverted tag will be necessary for the metrics that folks are working on. The long-tail of revert times have a lot of false positives due to bugs in how we calculate reverts -- i.e. null-content revisions at the very edge of snapshots that match up against any previously-deleted revisions in the page's history and page moves/protections. I checked a sample from English Wikipedia and it was 50% false positives after 48 hours. Rerunning your queries but with the WHERE ARRAY_CONTAINS(revision_tags, 'mw-reverted') as a much more high-fidelity signal of what edits actually have been reverted since 2020 gives us:
- The 48h window captures only 55% of all reverted revisions globally. Nearly half the signal is lost. becomes 70% within 48 hours (and 78% if you look just at Wikipedias). So just a quarter lost.
- 35.4% of all reverts cross a month boundary becomes 19% cross-month (12% if you look just at Wikipedias).
May 18 2026
no strong feelings -- it's true that e.g., mediawiki history drops the namespace prefix and so that seems to be the standard we've adopted in the data world. In that sense, just using page_title without additional qualifiers would seem to map to the non-prefixed version in most places. I guess my question would be: why include the namespace name if it's not relevant for any data joins? I guess it allows you to quickly construct a full URL if that's what's desired but that seems to be in the meta field already?
May 15 2026
@diego I looked through the report/ticket but didn't see anywhere that someone has actually looked at the reverts that get picked up when you move from a 48 hour to 7 day threshold. In the ticket, I do see some beliefs expressed that the reverts that happen after 48 hours might be trickier in nature. Which I think is a reasonable assumption but I took a look at some edits on English Wikipedia from April that were reverted between 2 and 7 days after the edit (just adjusting the logic slightly in P91457) and what I found was mostly a number of minor edits that just took a while until someone noticed them and decided to revert. The reverts that I imagine are probably of interest to DE4 are the ones that relate to sockpuppetry/block evasion, but those are a minority it would seem of these long-tail reverts and presumably there are more direct ways of capturing block/sockpuppet times (or the query could be narrowed to focus on time-to-revert for folks who are later blocked)?
May 13 2026
Ok, I had started some responses to the above but modifying a bit to summarize my thoughts post-discussion and try to separate out between things that I really hope we do long-term and things that are hopefully helpful for us getting this across the line in the short term.
May 12 2026
I gave this task a rather controversial title in order to help prompt discussion because I really think a change is needed because we no longer have easy access to survey results
Love it -- it got me to think deeply about what I actually wanted here as opposed to just defending the Welcome Survey as status quo.
Big thanks to all working on this and clearly laying it out! I know we're going to meet to discuss but I wanted to share some feedback in advance in case that helps with scoping the conversation:
May 11 2026
[unasked-for advice as part of my quest for better data] Do we have a more precise indicator of "made through structured editing experiences" than "≥1 Edit Suggestion was seen during the course of the edit session"? I'm assuming that means that someone clicked on the Suggestion to open up the details panel? But I wonder if we could go a bit further as I imagine folks will open up the Suggestions just out of curiosity. I know at least some of the suggestions have an action button and no button - could there be edit tags based on interaction with these instead? Or perhaps a more implicit "you opened the Suggestion and edited the element that was attached to it?"
May 7 2026
Oh and forgot to say, but the reason I'm not concerned about making the change around constructive edits quickly is that it doesn't look like it'll have a large impact on the metric. This chart comparing the two approaches (internal doc) shows that Retained Editors and Retained Constructive Editors have very similar trends, just that Retained Constructive Editors is about 5% less editors.
Hey all -- @Mayakp.wiki and I discussed this metric (and the related Active Editors metric). My summary below:
- Overview: to me, this metric seems like it's in a pretty good shape. I have some suggestions for improvement, which I'll lay out below, but I don't see them as blockers to "success".
- Should there be an error / confidence estimate for this metric? Maya and I agree that there shouldn't be. This is a count (not an estimate) and we don't know of any significant source of data loss so the number itself has no natural confidence interval. With any metric, there's a question of how to interpret changes in the trend (what's "significant"). I think that should be handled via proper contextualization / slices -- i.e. showing year-over-year numbers to handle seasonality, allowing for breakdowns by wiki -- which seem ready from what Maya was showing me.
- Any major exogenous/confounding events that could impact this metric?: these are counts of editors (not edits) and deduplicated across wikis, so for there to be a confounding factor here, you'd have to have A LOT of individual accounts created/reactivated and used to edit in a recurring way. That seems highly unlikely to happen in a way that's not intended to be captured by the metric. The only thing I can think of is a bot farm. The metric does exclude "bots", but that relies on the user account having "bot" in its username or being added to the bot usergroup, both of which only happen in cases where the bot account is being appropriately declared. For example, as far as I can tell, User:TomWikiAssist (the AI Agent that edited Wikipedia) is not in the bot group because they never went through the appropriate process for requesting permission and while they semi-declared their AI-agentness via "Assist", they didn't follow the "bot" norm. So how to exclude bots that don't follow policy (as a potential confounding factor that would move the metric without being the impact we're actually intending)? I think that gets into how we handle reverts, which I touch on below. All of that said, this is a theoretical thing at this point. As far as I know, we haven't seen mass bot account creation so I think it's worth addressing but again I don't see it as a blocker to a "successful" metric.
May 5 2026
Apr 29 2026
@KStoller-WMF this one makes me sad but I don't think you're necessarily wrong. One thought: rather than just removing it, why don't we think about it less as a survey and more as a simple onboarding flow? So instead of "Why did you create your account today?" with a dropdown and this being purely extractive, we have a "What do you plan to do?" with buttons maybe like:
- To create a new Wikipedia article: if click, redirect to something like https://en.wikipedia.org/wiki/Help:Your_first_article?
- To edit a specific Wikipedia article: maybe we provide an embedded search bar for them to find it?
- To learn how to edit: maybe we redirect to Newcomer Homepage?
- To read Wikipedia: if clicks, redirects to Main page?
- I'm not sure what I want to do: if clicks, maybe there's a general community hub we could redirect to like https://en.wikipedia.org/wiki/Help:Introduction?
And then the log data on these button clicks would functionally act in a similar way to the Welcome Survey data. Caveat that it's been a long time since I've created an account so I don't fully understand the existing flow and the above perhaps makes no sense.
Apr 28 2026
Hey @the-leeky-cauldron, thanks for reporting this and the patience. Some background:
Just a note that LTR vs. RTL has no effect on LLM performance (it's purely a reflection of how browsers etc. display the characters but they're encoded the same way from the model standpoint). The script/alphabet is the big one that actually matters and one major script family that's missing would be an Indo-Aryan language.
Apr 27 2026
Apr 24 2026
Just an FYI that @MGerlach will be out until May 6th -- let me know if you'd like faster validation and I can step in and handle.
Apr 10 2026
For project tracking purposes, I'm going to resolve this ticket. That said, @Alaexis feel free to continue to comment on it (at least until we have a new ticket opened for the next steps and then we can move discussion there).
Apr 9 2026
Just quickly chiming in to add another use-case for gpt-oss-safeguard-20b. I'll be reporting in more detail shortly, but we've found it be quite effective in T414816: [WE1.7.3] Exploration of automated verifiability checks. Namely, I ran our dataset of 119 claims+sources through the model and latency was average of <3 seconds and performance was quite high. Representative example of the prompt can be found in P90328 and working code here: https://gitlab.wikimedia.org/repos/research/source-verification/-/blob/main/notebooks/02b_%5Bround_2%5D_liftwing_pipeline.ipynb?ref_type=heads
Apr 7 2026
If the Micro-Task Generator can quickly generate the topics/tasks/quality associated with the article, why can't we?
I can help with that one -- we didn't have many users given that it was a prototype so we found that it worked fine to just hit the LiftWing API with a bunch of parallel requests (up to 20 I believe) and not do any caching. The topic/quality models are relatively lightweight so I think it didn't create too much load. I don't know what your usage assumptions are but what @MHorsey-WMF said with We cannot request a bulk prediction of topics for an article so we would need to run this asynchronously and store the result. feels right to me (and same for quality) for a production-level tool. My guess is that while "background job" sounds slow, it should be quite quick in practice and you all will probably just want to be better citizens than we were and check with the ML Team about load.
Apr 6 2026
At this point, I think the key things re: Structured to keep in mind:
- At this point, we don't have a Product OKR requirement that depends on StructuredEditTypes but I could easily expect one in the future -- e.g., I'm currently using it in a research hypothesis at T414816: [WE1.7.3] Exploration of automated verifiability checks to detect new citations within edits. Fabian needed it for another product exploration in T406827 (NDA so might not be able to view).
- It has some shared utilities within Simple so we just need to keep that in mind when making changes.
- I would agree with Aisha of keeping the focus on Simple for now. StructuredEditTypes are more complicated and prone to breaking if you make the wrong change. I would just ask that we document the approach well as hopefully whatever process we use for Simple can be extended to Structured with relatively little effort if we decide to do profiling there.
Apr 3 2026
Weekly updates:
- We have a dataset with 119 annotated claims (~40 of which might be "unclear" for the model) for evaluation. The breakdown of labels is as follows:
- No - Significant Omission: 29
- Yes - Direct Evidence: 19
- "Unclear (other language, too complex, etc.)": 15 (we may eventually update some of these labels based on the AI outputs and further exploration)
- "No - Landing Page (e.g., abstract only; generic website)": 13
- Yes - Accurate Paraphrasing: 12
- No - Over-Extrapolation: 11
- Yes - Implicit Consistency: 7
- No - Contradiction: 5
- No - Insignificant Omission: 5
- No - Unrelated Page: 3
- Some results coming in from @Trokhymovych suggest high overall performance (F1 > 0.8) for the question of Yes vs. No, which is good. They also seem to show that providing the full paragraph with the claim helps (to your point above Alaexis) and is an improvement over even providing just the previous sentence to the claim. This is all for gpt-oss-safeguard-20b right now and once we have fixed a few more parameters, we'll start bringing in other models.
- The ability to pinpoint the right explanation is a bit more mixed so we'll consider either iterating on the examples/descriptions for each or perhaps merging some of them. For instance, I could imagine merging the "Yes" categories together. Same with perhaps moving "Unrelated Page" and "Landing Page" into "Significant Omission" as they're essentially a subcategory of that.
- Code and some more detailed early results here: https://gitlab.wikimedia.org/repos/research/source-verification/-/blob/main/notebooks/01_%5Bround_1%5D_openai_pipeline.ipynb?ref_type=heads
Apr 2 2026
Agreed with everything @Daimona is saying here. Hearing this reflected back as well as seeing the progress that you all are making is leading me to shift some of my thinking away from "we should extend PageAssessments" (cc @ifried): one of the downsides for WikiProjects (at least for newer editors who aren't really familiar with them) is that there is no sign-up. Which means it's kinda weird to be dropped in when there's no immediate "action" to be taken and it's more of a "here's a place to ask questions and explore topic-specific documentation". And while there are some ways of tracking active vs. in-active WikiProjects, we don't have a great proxy for which ones are in a good place to "accept" newcomers either. On the other hand, I've been really impressed with the Events infrastructure that you all have built out and if @cmelo's work on worklists moves forward, you've brought Events to a place where they have all the aspects that make WikiProjects attractive in the sense of connecting editors with a broader community of editors interested in their topic:
- Worklist tracking
- Sign-up capabilities
- Discoverability -- this really hit home in the worklist demo where an editor who isn't signed up for an Event is given a "do you want to join this Event?" when they edit an article on its worklist. I think this is a lot more direct than trying to build a bridge between Events and WikiProjects via topics. If the worklists feel too narrow to be effective here, you can always work on recommending new articles to add to a worklist (which is something we've developed a lot of prototypes for).
- Off-ramp from individual editing to collaboration with like-minded editors -- this is what the Collaboration List currently is so I guess the key thing is how to make it easy for WikiProjects to have a corresponding "Event" that has the worklist/sign-up capabilities and therefore much higher discoverability.
Apr 1 2026
FYI in case this is helpful @MNeisler, this is the query I put together for the first round of continuous research recruitment: https://gitlab.wikimedia.org/isaacj/miscellaneous-wikimedia/-/blob/master/continuous-research/editors-to-contact.ipynb?ref_type=heads
Mar 27 2026
Updates:
- Many thanks to @MGerlach, @Miriam, @Trokhymovych, @TAndic, and @YLiou_WMF for help with annotating this week! We're up to 113 annotations but unfortunately that's led to still only 68 instances where the claim and the scraped content are sufficiently complete to enable evaluation. I'll continue to expand (and others are welcome to continue to chip in) though 68 is a big step forward.
- I reran the earlier stages in the pipeline to reflect changes I've decided to make since starting this annotation work -- i.e. scraping date from source alongside title+text; retaining article title + section title for claim. I also made made/released fixes to mwparserfromhtml and mwedittypes (dependencies used for identifying changed citations and extracting associated text) that fixed a few bugs.
- I collected some additional stats to on how often citations are co-located with other citations. Essentially, until now we've been considering a citation+claim pair in isolation. But there are cases where a claim is complex and so editors add multiple citations to cover the different aspects. Out of 261 citations evaluated from the dataset, 71% are the only one in their sentence (convenient) though we'll have to make a decision eventually about how to handle the more complicated cases:
- Cases where it's ecologically valid to evaluate the citation in isolation (not considering other citations in the article):
- 71% of new citations were the only one in that sentence (claim).
- 38% of new citations were the only one in the claim AND the paragraph leading up to that claim. In practice having additional citations on claims earlier in the paragraph doesn't change anything as right now we're still focused on "claim as a single sentence with the citation in it" but I generated the stats out of curiosity.
- Cases where there's just one more citation to consider -- e.g., maybe not too much additional latency/cost to add in a second source:
- An additional 15% of claims.
- An additional 24% of claim + previous text in the paragraph.
- Realistically speaking, it's relatively unbounded after 2. For example, there's a claim in the dataset that has 269 associated citations (likely a table or list with a ton of citations in it so perhaps misleading because we could break that up).
- Cases where it's ecologically valid to evaluate the citation in isolation (not considering other citations in the article):
- I'm working on synthesizing everything but current summary of where things stand (everything up to model evaluation):
- 10-20% of citations could be expected to not be evaluate-able. Breakdown:
- 10% fail fast (no scraping or request made to a model because e.g., citation lacks a URL)
- The other 10% would go through the full pipeline but our ability to handle them depends on if we can fix the extraction to handle more edge-cases. This is mostly from claims that require context from prior in the paragraph to be understood (e.g., usage of pronouns that don't refer to the subject of the article), which I think can be fixed by passing that context to the model in the prompt but we'll see when we run evaluation.
- An additional ~25% of citations would not be scrape-able. Breakdown:
- ~18% would fail somewhat quickly (scrape attempted but no request to model because no text returned from URL)
- The other ~7% return scraped text but it's incomplete. These might be caught by a model as insufficient (sometimes it's obvious) or might return false negatives – i.e. model saying no evidence when there is if you see the full source.
- 10-20% of citations could be expected to not be evaluate-able. Breakdown:
Mar 26 2026
I'm going to decline this for now. If we decide to further develop the tool, we can always reopen but I think not worthwhile for it as long as it stays in a more prototype phase.
Mar 25 2026
New challenge: because the citations are stripped of identifying information, the library thinks they get moved around much more than they do -- i.e. if one is deleted somewhere and another added somewhere else, they will look identical to the library. This is somewhat unavoidable as the citation itself in the HTML really only has numbering information (which does change all the time) and nothing specific to the reference itself. The probably only way to get around this (other than turning off "moves" for citations, which is not desirable either) is to create some fake HTML attribute and attach it to the citation that includes something from the citation text.
Mar 24 2026
From Pablo:
We are fully aligned for the first two thoughts.
Some thoughts I shared:
- Adding <div> as acceptable element alongside <table> seems like an easy win. This seems to cover the case of Italian and Polish Wikipedias. It's also part of the issue with Dutch Wikipedia.
- Adding flexbox as an acceptable class alongside mbox feels good too and addresses the other half of the issue with Dutch Wikipedia. It captures classic messagebox-style elements like the outdated-information box under the Toekomst section in this revision as well as article-for-deletion-type templates as in this revision. I often forget the latter is considered a messagebox but it is on English Wikipedia as well.
- Adding html_tag.has_attr("about") is interesting -- this essentially requires that the messagebox is a top-level transclusion. I think every messagebox will be transcluded but the question is whether we want to allow for nested messageboxes. For example, see this revision that has a multiple issues template that then has two issues listed (lacks in-line citations; lacks focus). Currently the library detects three messageboxes: the umbrella "multiple issues" one and then the two sub-messageboxes. Enforcing the about logic would mean just the multiple-issues messagebox is captured. I can see arguments for either approach.
- bandeau is also tricky. It'd be relevant to French Wikipedia but they don't reliably distinguish between hatnote-style boxes and message boxes so it could also generate a lot of false positives. For example, cnf and voir homonymes are actually hatnote templates though bataille en cours is a classic messagebox. They don't include the role attribute on hatnote templates so that can't be used to easily distinguish. Sometimes the hatnote class is applied but, for instance, in the "for deeper reading" section messages, no hatnote template is applied. This revision is a good example of the various templates for inspection. I've been assuming for a long time that an element can be one thing and one thing only but maybe that's not true? But this points to two decisions that have to be made: 1) how much to relax the criteria on these elements and whether to allow the same element to be retrieved by .get_message_boxes() and .get_notes(), and, 2) which element should come first in _tag_to_element, which is used for assigning element types in mwedittypes and plaintext extraction.
Some examples from Pablo using the function in the description:
