683 Magic numbers Should we be skeptical of performance guidelines which state that 100 milliseconds feels instantaneous to everyone?
715 Introducing: Autonomous Systems report Today we're publishing our first report of the performance experienced by visitors of Wikimedia websites, focused on the Autonomous Systems visitors are connecting from.
722 Evaluating Element Timing for images In the search for a better user experience metric, we have tried out the upcoming Element Timing for Images API in Chrome.
726 Performance perception: How satisfied are Wikipedia users? We've recently published research on performance perception that we did last year. The micro survey used in this study is still running on multiple Wikipedia languages and gives us insights into perceived performance.
732 Performance perception: The effect of late-loading banners Unlike most websites, Wikipedia and its sister projects are ad-free. This is actually one of the reasons why our performance is so good. We don't have to deal with slow and invasive third-parties.
737 Performance perception: Correlation to RUM metrics When we set out to ask Wikipedia visitors their opinion of page load performance, our main hope was to answer an age-old question: which RUM metric matters the most to users? And more interestingly, which ones matter the most to our users on our content.
741 Tracking down slow event handlers with Event Timing We're taking part in the ongoing Event Timing Chrome origin trial, in order to experiment with that API early and give feedback to its designers.
747 Wikipedia's JavaScript initialisation on a budget This week saw the conclusion of a project that I've been shepherding on and off since September of last year. The goal was for the initialisation of our asynchronous JavaScript pipeline (at the time, 36 kilobytes in size) to fit within a budget of 28 KB – the size of two 14 KB bursts of Internet packets.
752 WikimediaDebug v2 is here! WikimediaDebug is a set of tools for debugging and profiling MediaWiki web requests in a production environment. WikimediaDebug can be used through the accompanying browser extension, or from the command-line. This post highlights changes we made to WikimediaDebug over the past year, and explains more generally how its capabilities work.
782 Celebrating Free and Open Source Software with Google Summer of Code and Outreachy Programs like Google Summer of Code, Google Season of Docs, and Outreachy can provide a platform to accelerate your FOSS journey. Learn about Wikimedia’s outreach programs.
801 From Gerrit to Gitlab: join the discussion Wikimedia Release Engineering is considering a move from Gerrit to Gitlab. Learn more about the reasoning and join the discussion.
833 Diving into Wikipedia’s ocean of errors This blog post documents how we prioritized debugging an issue on mobile; how we went about implementing a solution; and what we learned from the experience.
862 MediaWiki History: the best dataset on Wikimedia content and contributors Learn about using the Mediawiki History Dataset to explore the every day experience of editors on Wikipedia.
917 Bot or Not? Identifying "fake" traffic on Wikipedia We have been working this past year to better identify and tag the "bot spam" traffic so we can produce top pageview lists that (mostly) do not require manual curation.
975 Wikimedia's CDN up to 2018: Varnish and IPSec The 1st of a 3 part series that will describe some of the changes, which included replacing Varnish with Apache Traffic Server (ATS) as the on-disk HTTP cache component of the CDN.
985 The Query Service Tutorial for Wikidata: an interview with Dr. Keren Shatzman from Wikimedia Israel With the Wikidata Query service tutorial, Wikimedia Israel aims to empower new users to query Wikidata, which currently holds almost 90 million items.
1026 Impact of using HTTP connection pooling for PHP applications at scale This post explores the challenges of running PHP applications at a large scale and discusses the effect of using Envoy on MediaWiki applications.
1040 Wikipedia as a castle in the wilderness: modernization in the dynamic world of the internet A developer's journey from software development to systems thinking, and how it impacts the future of Wikipedia.
1052 This bot is judging you: here's how we made it better at it After two years of planning and migration work, the 3rd generation WP 1.0 web tool has launched! Learn more about how it happened and what it means for Wikipedia's editors.
1069 Web performance case study: Wikipedia page previews Preview popups are common and requires careful scripting and styling; they can generate useful learning about performance as a reference for other front-end tasks.
1092 Wikimedia's CDN: the road to ATS The 2nd of a 3 part series that describes some of the changes to Wikimedia’s Content Delivery Network, including the replacement of Varnish with Apache Traffic Server (ATS) as the on-disk HTTP cache component of the CDN.
1164 Wikimedia's CDN: datacenter switchover The 3rd and last article of a series that describes some of the changes to Wikimedia’s Content Delivery Network, including the replacement of Varnish with Apache Traffic Server (ATS) as the on-disk HTTP cache component of the CDN.
1197 Perf Matters at Wikipedia in 2015 Looking back at our ups and downs. Including major performance improvements, deploying WebPageTest, and introducing WANObjectCache for MediaWiki.
1225 Make your technical work on Wikimedia projects more fabulous with these Phabricator video tutorials What is Phabricator? How does it work? Learn about a set of new video tutorials that can help you learn how to use the collaboration tool used by folks working in Wikimedia's technical spaces.
1252 2020 Coolest Tool Awards: Thank you for the tools! For the past two years, the Coolest Tool Awards has celebrated tools created and used by the Wikimedia community. Read about 2020's winners.
1257 Censorship, outages and Internet shutdowns: monitoring Wikipedia’s accessibility around the world This article describes the methodology used by the Wikimedia Foundation to monitor outages on Wikipedia around the world. These events are called anomalies and could be due to various causes, among them censorship.
1285 2020: The Year in Vue A look back at a year of experiences of using Vue within the Structured Data team.
1325 Cinder on Cloud VPS This post explores the new Openstack Cinder feature for Cloud VPS.
1410 Wikimedia SSO: Evaluation In the first of three posts about the implementation of Single Sign On (SSO). This post looks at the original landscape of Wikimedia's web-based services, summarizes requirements for a new SSO provider, looks at existing FLOSS solutions, and explains why Apereo CAS was chosen.
1447 Outreachy round 21: experiences and outcomes Wikimedia's participation in Outreachy Round 21 focused on projects related to data science and engineering. In this post, the interns share the outcomes and experiences of their projects.
1460 Searching for Wikipedia How people use Search to access Wikipedia is a common question by researchers. Until now, however, there has been little data available about this relationship. To help address these questions, the Wikimedia Foundation is releasing a new, faceted dataset on search engine traffic to Wikipedia so you can ask questions like "What is the most common search engine in my country?" or "Which search engine is most-used by Android users?"
1471 Discovering and fixing CVE-2021-33038 in Mailman3 During Wikimedia's Mailman3 migration, we discovered and fixed a security issue that would have disclosed the contents of private list archives during the import process. This post explains the issue, how we discovered it and how it was fixed.
1483 Alone together: Wikimedia Hackathon 2021 Wikimedia's 2021 Hackathon brought participants from all over the globe together online. This post explored the event and its outcomes.
1524 The rollout of single-sign-on (SSO) at the Wikimedia Foundation This is the second part of a three part series on the rollout of single-sign-on (SSO) at the Wikimedia Foundation.
1528 Wikimania 2021 Hackathon: 24 hours to experiment with other Wikimedians The Wikimania 2021 hackathon will take place on the first day of the remote Wikimania 2021, starting August 13th at 5:00 UTC and lasting for 24 hours.
1538 June 2021 Datacenter Switchover In June 2021, most user traffic was switched from our primary Virginia datacenter to our secondary one in Texas. This post covers how the swtichover went and the issues that came up.
1566 Sending messages to Wiki users in their preferred language With the improved MassMessage extension you can send wiki pages or select sections of wiki pages as messages, and these messages will be sent in their target language!
1588 The Wikipedia image/caption matching challenge and a huge release of image data for research! Wikipedia articles are missing images, and Wikipedia images are missing captions. A scientific competition organized by the Research team at the Wikimedia Foundation could help bridge this gap. The WMF is also releasing a large image dataset to help researchers and practitioners build systems for automatic image-text retrieval in the context of Wikipedia.\n
1622 Digging deeper into Quarry Wikimedia Cloud Services is planning improvements to the Quarry service.
1656 Analyzing the Wikipedia clickstream just got easier with WikiNav We have recently developed WikiNav, an interactive tool to analyze and visualize reader navigation, as part of an Outreachy-internship.
1685 Summit on present & future Vue.js user-interface library and the design system at the Wikimedia Foundation The Wikimedia Foundation Design Systems Team provides updates and outcomes from the recent Vue.js Designer Workshop and Developer Summit, which aimed to help us converge on a single design system and UI component library moving forward as we officially adopt Vue.js.
1695 How we deploy code In this post, you'll learn about the complexities of keeping the deployment train running.
1719 Roundup: Google Summer of Code 2021 and Outreachy Round 22 This year's outreach programs were some of the most successful ever. This post highlights projects from Google Summer of Code 2021 and Outreachy Round 22.
1728 Iterating on how we do NFS at Wikimedia Cloud Services NFS is a central piece of infrastructure that is essential to services like Toolforge. Recently, the Cloud Services team at Wikimedia had been reviewing how we do NFS.
1739 The trouble with triples This is part two of a three part series on the Wikidata Query Service.
1741 How we learned to stop worrying and loved the (event) flow This is part one of a three part series on the Wikidata Query Service.
1745 Getting the WDQS Updater to production: a tale of production readiness for Flink on Kubernetes at WMF This is part three of a three part series on the Wikidata Query Service.
1760 How we improved performance of a batch process from two days to five minutes The Growth team recently improved the performance of a script that prepares data for usage in the mentor dashboard. Learn about how they decreased the average runtime of the script from more than 48 hours to less than five minutes.
1806 How we made editing Wikipedia twice as fast Over the past six months we deployed a new technology that sped up Wikipedia’s backend application, reducing the median page-saving time for editors from about 7.5 seconds to 2.5 seconds.
1831 Wikipedia and Apps: A Love Story Wikimedia Foundation’s apps are an essential piece to meet our “this is for everyone” design principle. Wikimedia apps are designed with the philosophy of mobile first in mind.
1868 Pawing around with PAWS: recent updates to Wikimedia Cloud Services' Jupyter notebooks instance Learn more about recent updates to PAWS: A Web Service, Wikimedia Cloud Services' Jupyter notebooks instance.
1947 Propose sessions and projects for Wikimedia Hackathon 2022! It's time to start proposing sessions and projects for hacking at the Wikimedia Hackathon 2022!
1952 What it takes to parse MediaWiki page titles... in Rust MediaWiki page titles are the primary identifiers for all wiki content - learn how they are validated, normalized and parsed and what it took to do so in Rust.
1967 Modernizing our tech stack for serving maps at Wikipedia MediaWiki allows editors to contribute with geospatial content. Learn more about the internals of the map technology we use and the improvements introduced.
1975 New discovery tool for technical documentation The Wikimedia Developer Portal guides technical audiences to key documentation and community resources. By organizing links into thematic sections focused on developer tasks, it helps people more easily find information about Wikimedia technology.
1979 Explore wiki project data faster with mwsql Learn how the mwsql library makes it easier to download and work with SQL dump files in formats like Pandas dataframes or CSV.
1994 What is in an edit? Automated detection of edit types on Wikipedia Learn about a new Python library for automatically detecting and summarizing what content is changed by edits on Wikipedia.
2008 Building DReaMeRS: How and why we opened a datacenter in France Content delivery networks (CDNs) are one of the modern building blocks of the Internet. This blog post highlights a recent addition to Wikimedia's CDN setup.
2072 Introducing Terraform support on Wikimedia Cloud VPS Learn how Wikimedia Cloud VPS was adapted to work with Terraform.
2094 HTTP/2 performance revisited Deploying HTTP/2 support to the Wikimedia CDN significantly changed how browsers negotiate and transfer data during the page load process. We found regressions in performance during the transition and are sharing the lessons we learned.
2147 Web Perf Hero: Valentín Gutierrez Today we celebrate two numbers: 25% lower latency for ATS backend requests at the p75, and up to 1000X reduction of ATS disk read latency at the p999.
2169 How we're building our Kubernetes pipeline in GitLab Our deployment pipeline helps our developers build, test, and release Docker images to our production Kubernetes. Now we're migrating the pipeline to GitLab and seizing the chance to refine our tools.\n
2180 Perf Matters at Wikipedia in 2016 Looking back at our ups and downs. Including HTTP/2 deployment, metric collection improvements, and the begining of our journey to Thumbor.
2239 Around the world: How Wikipedia became a multi-datacenter deployment Learn why we transitioned the MediaWiki platform to serve traffic from multiple data centers, and the challenges we faced along the way.
2355 Flame graphs arrive in WikimediaDebug The new "Excimer UI" option in WikimediaDebug generates flame graphs. What are flame graphs, and when do you need this?
2414 Web Perf Hero: Máté Szabó We recognize the volunteer effort that increased Wikipedia’s backend responses that complete within 50ms by 20%.
2586 Perf Matters at Wikipedia in 2018 Looking back at our ups and downs. Including a literature review, multi-DC prep, and a new data center in Singapore.
2655 Perf Matters at Wikipedia in 2019 Looking back at our ups and downs. Joining the W3C, implementing Paint Timing in upstream WebKit, and researching perceived performance.
2695 Web Perf Hero: Amir Sarabadani Today we recognize Amir's work over the past six months which cut latencies by half!
2749 Perf Matters at Wikipedia in 2020 Looking back at our ups and downs, including FOSDEM and Mobile Device Lab.
2915 Unifying our mobile and desktop domains How we achieved 20% faster mobile response times, improved SEO, and reduced infrastructure load.
2995 Web Perf Hero: Thiemo Kreuz From optimizing the MediaWiki stack on Wikipedia.org, to speeding up CI; Thiemo's work benefits us every day!
198947 Investigating a performance improvement When a performance improvement seems too good to be true, it's time to investigate in depth and find out what happened.
198948 Improving time-to-logo performance with preload links Thanks to a new web standard, we've recently deployed a small performance improvement that highlights some of the unique challenges we encounter on Wikimedia sites.
198949 Looking back: improvements to edit save time Is it faster to save articles on Wikipedia than it was a year ago?
198950 The journey to Thumbor, part 1: rationale Making Thumbor production-ready for Wikimedia is a journey that started a year and a half ago. Let's look at the rationale for this project.
198951 The journey to Thumbor, part 2: thumbnailing architecture To understand why Thumbor is a good fit, it's important to understand where it fits in our overall thumbnailing architecture. A lot of historic constraints come into play, where Thumbor could be adapted to meet those needs.
198952 The journey to Thumbor, part 3: development and deployment strategy Introducing Thumbor replaces an existing service, and as such it's important that it doesn't preform worse than its predecessor. We came up with a strategy to reach feature parity and ensure a launch that would be invisible to end users.
198953 Measuring Wikipedia page load times Here is how we measure and interpret load times on Wikipedia. Let's also look at what real-user metrics are, and how percentiles work.
198954 Thumbor support for private wikis deployed Yesterday we deployed Thumbor support for Wikimedia-hosted private wikis. While 99.9% of our traffic is for public-facing wikis, the Wikimedia Foundation hosts a number of private MediaWiki instances on the same infrastructure.
198955 Mobile web performance: the importance of the device Let's explore our web performance data from an angle we haven't explored before: mobile device type.
198956 Performance testing in a controlled lab environment - the metrics One of the Performance Team responsibilities at Wikimedia is to keep track of Wikipedias performance. Why is performance important for us? In our case it is easy: We have so many users and if we have a performance regression, we are really affecting people's lives.
198957 Best friends forever We use both synthetic and RUM testing for Wikipedia. These two ways of testing performance are best friends and help us verify regressions. Today, we will look at two regressions where it helped us to get metrics both ways.
198958 Machine learning: how to undersample the wrong way Machine learning is a powerful tool, but it's easy to use it incorrectly and draw biased conclusions, as we'll show in this real world example.
198959 Why performance matters Performance is important in many ways, some of which matter particularly for the Wikimedia Foundation.
198961 Debugging production with X-Wikimedia-Debug Here are ways we debug issues in production when users report bugs.
198962 Computational knowledge: Wikidata, Wikidata query Service, and women who are mayors! One of the main aims of Wikidata is to represent knowledge in a way that is computable—that is, amenable to automatic processing. Wikipedia already contains a lot of information; much of it is reasonably easy for a human to understand—though some of the more esoteric bits are decidedly not—but it’s not at all readily crunchable by a computer.
198963 Parsoid in PHP, or there and back again In December 2019, we replaced the original version of Parsoid, written in JavaScript, with a version written in PHP, the primary programming language of MediaWiki. This new version, called Parsoid/PHP, is roughly twice as fast as the original JavaScript version. Parsoid/PHP brings us one step closer to integrating Parsoid and other MediaWiki wikitext-handling code into a single system.
198964 Google Code-In 2019: The next generation of technologists contribute to Wikimedia’s code Each year the Google Code-in contest brings students and mentors from around the world together to improve their technical skills and to make contributions to the Wikimedia movement.
198965 Saying no to proprietary code in production is hard work: the GPU chapter Maintaining and improving one of the largest websites in the world using Open Source software requires a continuous commitment. The site is always evolving, so for every new component we want (or need!) to deploy, we need to evaluate the Open Source solutions available.
198966 Fixing npm security issues immediately in MediaWiki projects LibUp writes a commit message by mostly analyzing the diff, fixes up some changes, and pushes the commit to Gerrit to pass through CI and be merged. If npm is aware of the CVE ID for the security update, that will be mentioned in the commit message. Each package upgrade is tagged, so if you want to e.g. look for all commits that bumped MediaWiki Codesniffer to v26, it's a quick search away.
198967 Organizing and running a developer room at FOSDEM For FOSDEM 2020, the Wikimedia Performance Team organized a Web Performance devroom. In this post, they share their experience.
198968 What's trending? New report lets editors know when Wikipedia articles go viral Posts on social media can make an otherwise obscure Wikipedia article go viral. A new traffic report gives English Wikipedians new insight into which ones are being read and shared most on four major social media platforms.
198969 Using Zulip, an Open Source tool, for engaging participants in Wikimedia's technical outreach programs Zulip enables organizers and mentors to provide support and guidance to participants in each phase of technical outreach programs.
198970 Measuring the performance of Wikipedia visitors' devices We have been collecting microbenchmark scores for over a year. This lets us see the long-term evolution of our audience as a whole. The information gives us an idea of how fast device/operating system/browser environments improve on their own.
198971 What does it mean to “keep community in the loop” when building algorithms for Wikipedia? We argue that if you follow three steps, you can “keep community in the loop” as you build algorithmic systems, making you more likely to avoid catastrophic and community-damaging consequences.
198972 A better Toolforge, part 1: upgrading the Kubernetes cluster This article focuses on how we made a better Toolforge by integrating a newer version of Kubernetes and, along with it, some more modern workflows.
198973 Permuting Khmer: restructuring Khmer syllables for search Khmer is written left-to-right in syllable groups, without spaces between words—though that’s not its most complicated feature!
198974 Celebrating 600,000 commits for Wikimedia A reflection on the developer services we offer and our community of developers.
198975 A better Toolforge, part 2: a technical deep dive In this follow up post, we dive deeper into the technical details of the recent Kubernetes upgrade for Toolforge.
198976 First-ever Coolest Tool Award: Thank you for the tools! The Coolest Tool Award acknowledges projects and with that, all the people who have been contributing to those projects — primarily the developers and maintainers, and the people who help with documentation, translation, feedback, technical advice, design or communication.
198977 How we contributed Paint Timing API to WebKit The story of how we decided to commission the implementation of Paint Timing API, a feature that lets us observe web performance from an end-user perspective. This web browser feature tells us at what point in time content started to appear on the screen for a visitor.
198978 An Orca screen reader tutorial This tutorial covers how to install, set up and use Orca on Linux systems, with sighted developers, product managers, and user experience designers in mind. Orca is a screen reader available on Linux which is being continuously developed as part of the GNOME project. It is probably the best choice to test screen reader conformance when developing on Linux.
198979 All code is built. Build what you can before you ship. Every page request in the ResourceLoader pipeline benefits from compilation, minification, and bundling build steps but not every build step fits in the runtime space.
198980 Designing technical workshops for the Indic community This blog post summarizes planning that went into designing a technical workshop series for the Indic community, key outcomes, success stories, lessons learned, and some next steps! It targets potential organizers who might be interested in conducting similar training in their wiki community.
198981 Wikimedia’s Event Data Platform, or JSON is ok too The Wikimeda Foundation has been working with event data since 2012. Over time, our event collection systems have transitioned from being used only to collect analytics data to being used to build important user facing features. This 3 part series will focus on how Wikimedia has adapted these ideas for our own unique technical environment.
198982 Wikimedia’s Event Data Platform - JSON & Event Schemas In the previous post, we talked about why Wikimedia chose JSONSchema instead of Avro for our Event Data Platform. This post will discuss the conventions we adopted and the tooling we built to support an Event Data Platform using JSON and JSONSchema.
198983 Wikimedia’s Event Data Platform - Event Intake Part 3 of 3 posts on Wikimedia's event data platform.
198984 Ceph distributed VM storage coming to Cloud Services Part 1 of 2 posts exploring Ceph and Wikimedia Cloud Services.
198985 Ceph at WMCS, the numbers and the details Part 2 of 2 posts exploring Ceph and Wikimedia Cloud Services.
198986 Lalitha's story: an Outreachy intern shares her experience In this post, Lalitha, an Outreachy intern from 2020, shares her experience with the program.
198987 A journey of a single step begins with a thousand miles Sometimes a simple bugfix can take longer than expected, but seeing it through is worth the journey.
198989 Wikimedia and Google Season of Docs 2020 For the last two years, Wikimedia has participated in Google Season of the docs. In this post, Gbadebo Bello shares his experiences with the program.
198990 In search of the perfect search for Wikipedia In a recent project, the Android app team decided to improve one of their core experiences: searching for articles on Wikipedia. Our goal was to make the discovery process for readers more intelligent, personalized and efficient to present the right result at the right time.
198991 Profiling PHP in production at scale We built an efficient sampling profiler for PHP. It runs continually in production on live requests, and generates trace logs and flame graphs.
198992 Upgrading Hadoop in just one day The Wikimedia Analytics Engineering team manages multiple systems, all gravitating around a big (for our standards) Hadoop cluster. This post describes our path to changing our Hadoop distribution in a single day, together with the lessons learned while doing it.
198993 Introducing Database as a Service on Cloud VPS This post introduces a new service, Database as a service, on Cloud VPS.
198994 Call for smaller language wikis to participate in the technical capacity-building program! Learn more about how to participate in the small wiki toolkits initiative.
198996 Toolforge GridEngine Debian 10 Buster migration In accordance with our operating system upgrade policy, we should migrate our servers to Debian Buster.
198997 Wikimedia Hackathon 2022: save the date and apply for grants and scholarships now! The Wikimedia Hackathon 2022 will take place as a hybrid event on May 20-22, 2022. The Hackathon will be held online and there are grants available to support local in-person meetups around the world.
198998 Join the Wikimania Hackathon, August 12-14 2022! The Wikimania 2022 Hackathon is a free, online event where anyone interested in Wikimedia technology can work together on projects, learn new skills, and meet other technical contributors. Don’t miss this opportunity to come together with other Wikimedia technical community members!
199000 From hell to HTML: releasing a Python package to easily work with Wikimedia HTML dumps Announcing mwparserfromhtml, a new library that makes it easy to parse the HTML content of Wikipedia articles