User Details
- User Since
- May 19 2025, 3:26 PM (64 w, 1 d)
- Availability
- Available
- LDAP User
- Guilherme Gonçalves
- MediaWiki User
- GGoncalves-WMF [ Global Accounts ]
Today
Thu, Aug 6
Data Gateway is supported by Data Persistence SRE, so it would be good for them to look into it next.
Wed, Aug 5
Correct, and I think the main factor for moving to private replicas is if the performance regression becomes intolerable.
Tue, Aug 4
Marking this low priority: it would be nice to get to this in Q1, and perhaps we can keep this as a time-boxed spike, but it's perfectly fine if it happens later in the year as we solidify Data Warehouse plans.
Fri, Jul 31
Sounds good to me, and we should assign to a point of contact in Editing to drive this.
Thu, Jul 30
We do need to have the conversations about harmonizing Prep Pantry and SCROLL in the general case. @MLechvien-WMF have chatted briefly about it and should catch up again, now that we have this ticket as a real data point :) I'd say that discussion is out of scope for this ticket.
Fri, Jul 24
I think it's worth taking a step back here when Peter returns, so we look at what the hypothesis is requesting and what's covered by existing SLOs and infrastructure, and what needs to be new.
Wed, Jul 22
I'm marking this as blocked by T425666 which tracks the actual introduction of the new model and webrequest_v2. Note that even though the ticket refers to a numerical bot score, we recommend querying for agent_type = "human" instead of using the numerical score directly.
Having chatted with Halley, what we currently have should be enough for establishing a baseline as a one-off query against webrequest_v2, which now contains scores for non-pageviews. I've also commented to that effect in T430637.
Now that we're updating the bot detection model in the data platform, you should be able to run a query like:
Tue, Jul 21
Boldly moving this one level up, because I think we want to discuss SRE support in earnest only after the controlled experiment has proven that we want to keep this intervention as a long-term investment.
While this is still possible, and I think it will be desirable to measure pageviews with bot detection for non-wikis, this is also more complex than I thought.
Thu, Jul 16
We don't believe this change blocks enabling private tags; until we resolve this ticket, private tags will not be ingested into the Data Lake at all (see Slack discussion).
Tue, Jul 14
Here's the alerting methods I know of, and how I think Superset alerts and reports fit into the picture:
Mon, Jul 13
We will have the need for this if/when we start migrating some of our existing pipelines to dbt. You are correct that it would simplify the proposal though, we can definitely leave this for a separate feature upgrade.
Indeed, KC and I chatted about this afterwards, and that's probably going to be the plan. As I understand it, this ticket is obsolete, so I'll close it now. Feel free to reopen if I'm misunderstanding.
Jul 6 2026
Can you please expand on the "track activity" part? I'd like to understand what we need to measure here so we can find the best table(s) to create and change.
Jul 1 2026
Oh good point! I used "asynchronous" mostly to mean "not in CI/Gitlab", but you're right that this still leaves two options:
Jun 30 2026
That would mean we'd have to regularly import the JSON to be able to exact-match individual UAs against it, rather than extracting information from the UA ourselves like we currently do. Is that right?
Thanks! A couple of questions and comments:
If the goal is to reject events from self-declared bots, it's worth also keeping an eye on T430020, which is attempting to provide one canonical table for those.
Right, I'm totally fine with translating our designs to Wikitech once completed and making that the source of truth. I only prefer Google docs for gathering requirements and discussing finer design details inline.
Jun 29 2026
Looks good to me!
Looks excellent, thank you! I've just made a minor edit to fix a typo and add a link to an issue.
Jun 26 2026
If I did it right, the more sensitive tickets are restricted to the WMF-NDA group. Those are the tickets that list specific bot detection signals and datasets, and how we combine them for final scoring.
Jun 25 2026
Are we going to store only bot-like User-Agents or all of them? Maybe, since we can not extract contact/agent/policy information from them and they are also not useful for KAPOW, we might not want to store them to save storage space.
Jun 23 2026
Jun 22 2026
Having spoken to @KCVelaga_WMF this morning, we can break up the work into a couple of chunks to keep the risk to WE5 manageable:
Jun 19 2026
I don't have a strong opinion on "one row per source", that's sensible if you two agree to it :) For what it's worth, in KAPOW, we'll probably not care about what is the source of a UA, we'd just join to that table based on (unique) UAs.
Jun 17 2026
Thanks, that makes sense!
This looks good to me, thanks for proposing!
Jun 16 2026
I can see user_agent_compliance_hourly is computed from webrequest by joining it with user_agent_compliance_classified_hourly.
Jun 15 2026
Jun 9 2026
May 19 2026
May 15 2026
Hi, thanks for the detailed analysis and I'm asking the team about the data gap in top_pages_per_editor.
May 7 2026
Apr 24 2026
Sounds good to me!
That's the MIME type of the response for a particular transcoding, is that right? So if we serve an image poster for a .ogg video, content_type will say image?
Apr 21 2026
Sorry for the delay here, @Ahoelzl is looking for an assignee.
Apr 13 2026
Thanks! If I understand correctly, we have 3 possible approaches: