hey look at this graph ๐
User Details
- User Since
- Nov 5 2018, 2:54 PM (404 w, 5 d)
- Availability
- Available
- IRC Nick
- cdanis
- LDAP User
- CDanis
- MediaWiki User
- CDanis (WMF) [ Global Accounts ]
Tue, Aug 4
The patch, for posterity: https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1305065
Wed, Jul 29
Mon, Jul 27
It seems to be happening much, much more often since 11:30 UTC today.
Jul 6 2026
Jul 1 2026
00:42:36 <Amir1> okay, the user impact should be gone now
00:43:11 <Amir1> it's removed from writes
Jun 26 2026
Don't have much context here, but I'm very happy to help figure out how LAC/hoarde can best integrate with tracing as deployed in WMF prod.
Jun 17 2026
Today (2026-06-17) was a recurrence of this, with many puppet-compiler jobs stuck in "queued" state for 3+ hours because of issues in other Zuul pipelines.
There's a hidden Option 4 here, which is to declare that urldownloader would be the first Sophroid-only service, only accessible via the service mesh and Envoy, but apparently the little remaining Sophroid work isn't scheduled until Q2.
Jun 16 2026
It might be possible to work around this by creating headless services for these jobs that access the databases -- definitely worth trying.
Jun 12 2026
Jun 11 2026
The MR looks good to me! I'm happy to help with the haproxy patch portion of possibility 1, which is my preference of the two.
Jun 3 2026
This does indeed look promising! Let's watch for another week or two, and then we can probably call it good enough.
May 29 2026
May 28 2026
+1 from me on supporting the format pywikibot already emits for on-wiki usernames.
May 26 2026
Thank you Luca!
I think, as compared to aggressive, middle-ground mostly adds extra work without actually buying much safety or flexibility.
Can confirm this is also affecting https://gitlab.wikimedia.org/repos/sre/CIDERGRINDER
May 20 2026
May 14 2026
May 12 2026
May 7 2026
FWIW I kept the names unchanged on the past-Benthos side
Apr 28 2026
Apologies, SRE put a rule in place as part of responding to an especially-impactful scraping incident, and it was too broadly scoped. It's been reverted since about 09:05 UTC today (thanks @Fabfur!)
Apr 17 2026
BTW -- here are two canned queries for distributed traces of uploads: https://w.wiki/LSLK and https://w.wiki/LSLn
What Jaeger sees only includes the requests from MediaWiki towards ms-fe hosts, of course -- but you can use the x-request-id from the tags to pivot into logs.
I've no idea if Swift can omit OTel, or if it could use a proxy that could.
FYI: as of my Puppet patches above, you can now use an x-request-id value to find all the intra-Swift requests associated with that request.
Apr 15 2026
Apr 13 2026
thanks for the report, sorry for missing this as part of T414486
Live-patched in production, service is restored.
Sounds a lot like T248872: puppet-merge lockout/tagout ?
Apr 1 2026
Mar 31 2026
This should be working now:
deployment.eqiad.wmnet
user fundraising-data-uploader
/var/lib/fundraising-data-uploader (the user's homedir)
Mar 30 2026
Apologies, I'll write a host firewall patch tomorrow morning.
Mar 25 2026
Thanks Joseph <3
Mar 23 2026
Mar 21 2026
Mar 20 2026
+1 from me on making UUIDv4 the only codepath
Mar 19 2026
Mar 16 2026
I haven't played around with https://connectrpc.com/ at all, but it looks like it was designed with these exact concerns in mind, and could allow us to write and ship a fully-compliant gRPC server that also, with zero additional work required, would accept an easy-to-speak dialect of plain HTTP.
Mar 9 2026
Mar 6 2026
Most user JS scripts should be working again.
Mar 5 2026
Apologies for the conflicting information, that's partially my fault.
The app involves fetching a large number of apps to present to the user.
Feb 25 2026
Feb 23 2026
Feb 19 2026
๐root@puppetserver1002.eqiad.wmnet /srv/git/operations/private ๐๐ git reset --hard origin/master HEAD is now at ce722766 (herron) dummy commit to resync logstash collector yaml
tagging High because it is so low-effort
We have used llama-server on our relforge instances (dual socket xeon silver w/ 12 cores each) to provide reranking in our prototype via a custom Qwen3-Reranker-0.6B-Q8_0.gguf but the performance is borderline unacceptable (2-3s for n=10). Q8 was used as it provided a significant improvement to inference latency over the unquantized model.
Feb 18 2026
Sounds good to me.
Feb 11 2026
In particular it'd be good to know if you only care about encryption going over the wire, or if encryption at rest is also necessary.
The latency trend for this past ~week is even worse:
Thanks for filing that @Tgr !
This should be fixed now -- I did some quick checks but further confirmation appreciated :)
Feb 6 2026
Parsoid looks to be consuming basically all the CPU that mw-jobrunners have to offer:
Looks like it lines up with the train deployment of MediaWiki 1.45/wmf.6 ?
It isn't just s2. It's all of them. Since the switchover.
I found the request:
https://logstash.wikimedia.org/goto/78f43ba4e27abb0fa03ad86fcae79dfc
Feb 5 2026
Feb 3 2026
Jan 29 2026
Jan 23 2026
Jan 20 2026
Totally fine from SRE's side, just check in with the current oncallers before you begin the disruptive part of the maintenance.
Jan 16 2026
The appservers RED k8s dashboard makes even heavier queries, and was the trigger of a Thanos outage this week.
Jan 15 2026
Q: How long should I leave this running?
A: Up to one whole workday at a time. Don't set it and totally forget it -- it's directing you towards one of our edge sites, and that will need to change if we need to depool that edge site.
