Page MenuHomePhabricator

hashar (Antoine Musso)
LogisticsAdministrator

Today

  • No visible events.

Tomorrow

  • No visible events.

Friday

  • No visible events.

User Details

User Since
Oct 3 2014, 2:31 PM (619 w, 4 d)
Roles
Administrator
Availability
Available
IRC Nick
hashar
LDAP User
Hashar
MediaWiki User
Unknown

https://www.mediawiki.org/wiki/User:Hashar

I am based in France CET/CEST (UTC+1, UTC+2). I have been a volunteer since ~ 2002 and employed at the Wikimedia Foundation since 2011.

My team is Release-Engineering-Team in which I notably maintain Jenkins Zuul Gerrit Continuous-Integration-Infrastructure Continuous-Integration-Config and various other things such as running the weekly MediaWiki deployment.

The preferred ways to reach me are:

IRC Libera.Chat

  • #wikimedia-releng
  • #wikimedia-operations
  • Direct message /query hashar

File a task in Phabricator and subscribe me to it (@hashar).

Email, Slack etc are read on an inconsistent best effort basis

Recent Activity

Yesterday

hashar claimed T435186: Addition of zuul1004 broke zuul-eqiad zookeeper.
Tue, Aug 18, 1:13 PM · Patch-For-Review, Collaboration-Services, Continuous-Integration-Infrastructure (Zuul upgrade)
hashar created T435186: Addition of zuul1004 broke zuul-eqiad zookeeper.
Tue, Aug 18, 10:56 AM · Patch-For-Review, Collaboration-Services, Continuous-Integration-Infrastructure (Zuul upgrade)
hashar closed T435112: Coverage builds fail with HTTP 429 error as Resolved.

I am assuming that one was related to yesterday GitHub outage.

Tue, Aug 18, 7:32 AM · Release-Engineering-Team, Continuous-Integration-Infrastructure

Mon, Aug 17

hashar added a project to T434749: Offload queue wait time from the feedback loop time: Jenkins.
Mon, Aug 17, 10:18 AM · Continuous-Integration-Infrastructure, Jenkins, Spike, Test Platform, Castor
hashar added a comment to T434749: Offload queue wait time from the feedback loop time.

I had a look at the plugin dev instance, created a new job and looked for:

  • build steps, it has Conditional Step (single), Conditional Step (multiple).
  • post build action (publishers), does not have the conditional step.
Mon, Aug 17, 10:18 AM · Continuous-Integration-Infrastructure, Jenkins, Spike, Test Platform, Castor
hashar added a comment to T434749: Offload queue wait time from the feedback loop time.

I don't remember why I went to use an unconditional execution of the castor-save-workspace-cache job (that was a decade ago T112560).

Mon, Aug 17, 10:13 AM · Continuous-Integration-Infrastructure, Jenkins, Spike, Test Platform, Castor

Sat, Aug 15

hashar reopened T429547: Migrate sonar analysis from deprecated sonar.login to sonar.token as "Open".

Reopening since there are still open changes to transition from sonar.login to sonar.token: https://gerrit.wikimedia.org/r/q/bug:T429547 Some can probably be raised to the #wmf-java channel on Slack :)

Sat, Aug 15, 9:07 AM · Test Platform (Aktau 28), Patch-For-Review, Continuous-Integration-Config, Quality-and-Test-Engineering-Team (SonarCloud Admin)

Fri, Aug 14

hashar added a comment to T256168: Move beta cluster automatic deployment to a dedicated infrastructure.

I have deleted the 3 Jenkins jobs (beta-code-update-eqiad, beta-scap-sync-world, beta-update-databases-eqiad) and the resulting view that became empty: https://integration.wikimedia.org/ci/view/Beta/

Fri, Aug 14, 10:06 AM · User-bd808, Release-Engineering-Team (Doing 😎), Continuous-Integration-Infrastructure, Quality-and-Test-Engineering-Team (Test Infrastructure), Jenkins, Continuous-Integration-Config, Beta-Cluster-Infrastructure

Thu, Aug 13

hashar added a comment to T434024: CI build slow and timing out sometimes.

@Lars yes that is my suspicion and I will keep using that task for now since it has all the debugging context ;)

Thu, Aug 13, 10:05 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I'd like to find how the build is so much slower (15 vs 24 minutes when CPU speeds are not that different: 2.6GHz vs 2.1GHz). Maybe I can try to reproduce the CiviCRM on a fast vs a slow host and see whether I can find something, but maybe it is due the CPU power management.

Thu, Aug 13, 4:20 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I will try to reproduce the CiviCRM on a fast vs a slow host and see whether I can find something. I'd like to find how the build is so much slower (15 vs 24 minutes when CPU speeds are 2.9GHz vs 2.6GHz).

Thu, Aug 13, 4:08 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T433980: PCC: no nodes found for class: Class/Role::Jenkins.

As a workaround, instead of using the roles:

Hosts: O:ci
Hosts: O:jenkins

I went to use the hosts I am interesting for:

Hosts: contint1002.wikimedia.org
Hosts: contint1003.wikimedia.org
Hosts: contint2003.wikimedia.org
Thu, Aug 13, 1:12 PM · Infrastructure-Foundations, Puppet CI
hashar updated the task description for T434766: Normalize URI host in Turnilo's webrequest_sampled_live.
Thu, Aug 13, 10:29 AM · Data-Engineering-Radar, Traffic, Data-Engineering
hashar closed T434712: Error: Class "Wikibase\Repo\WikibaseRepo" not found as Resolved.

The hotfix has been deployed. I have promoted wmf.15 to group 1 wikis and the error is no more happening 🎉

Thu, Aug 13, 9:32 AM · MW-1.47-notes (1.47.0-wmf.15; 2026-08-11), Wikibase Reuse Team, User-ItamarWMDE, User-brennen, Wikidata, Wikidata Lexicographical data, Wikimedia-production-error
hashar closed T434712: Error: Class "Wikibase\Repo\WikibaseRepo" not found, a subtask of T430834: 1.47.0-wmf.15 deployment blockers, as Resolved.
Thu, Aug 13, 9:32 AM · User-brennen, Essential-Work, Release-Engineering-Team (Priority Backlog 📥), Release, Train Deployments

Wed, Aug 12

hashar closed T427922: Speedup the mwext-phpunit-coverage-publish castor job as Resolved.

The npm cache directory has not reappeared which solves this issue:

$ sudo du -s -h /srv/castor/*/*/mwext-phpunit-coverage*/
314M	/srv/castor/castor-mw-ext-and-skins/master/mwext-phpunit-coverage-publish/
Wed, Aug 12, 7:54 PM · Patch-For-Review, Quibble, Test Platform (Nairobi 29), Castor
hashar closed T427922: Speedup the mwext-phpunit-coverage-publish castor job, a subtask of T427450: Find the root cause for the long Castor wait times, as Resolved.
Wed, Aug 12, 7:54 PM · Patch-For-Review, Castor, Test Platform (Basel 26)
hashar closed T427922: Speedup the mwext-phpunit-coverage-publish castor job, a subtask of T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting, as Resolved.
Wed, Aug 12, 7:54 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar closed T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting, a subtask of T427471: Speedup mwext-codehealth-master-non-voting Castor job, as Resolved.
Wed, Aug 12, 7:52 PM · Patch-For-Review, Continuous-Integration-Infrastructure, Castor
hashar closed T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting as Resolved.

The npm cache directory has not reappeared which solves this issue.

du: cannot access '/srv/castor/*/*/*codehealth*/npm': No such file or directory
Wed, Aug 12, 7:52 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar added a comment to T434024: CI build slow and timing out sometimes.

The list of CPU we for later reference:

Wed, Aug 12, 2:41 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

Looking again at https://integration.wikimedia.org/ci/job/wikimedia-fundraising-civicrm-bookworm/buildTimeTrend , there are fast builds happening on other hosts than integration-agent-docker-1093:

Wed, Aug 12, 1:41 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I went to query all 25 instances for the reported MHz using integration-cumin.integration.eqiad1.wikimedia.cloud:

sudo cumin --force -p 0 'name:docker' 'grep -m 1 MHz /proc/cpuinfo'
Wed, Aug 12, 1:19 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T427922: Speedup the mwext-phpunit-coverage-publish castor job.

I have updated the 3 mwext-phpunit-coverage* jobs and nuked the /cache/npm directory. I am letting this open for a little while so we can verify the cache does not get populated back. That can be done using:

ssh integration-castor06.integration.eqiad1.wikimedia.cloud \
  sudo du -s -h /srv/castor/*/*/mwext-phpunit-coverage*/npm

Which was 11GB!!

Wed, Aug 12, 12:34 PM · Patch-For-Review, Quibble, Test Platform (Nairobi 29), Castor
hashar reassigned T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting from hashar to Peter.

Assigning to @Peter who found the root cause.

Wed, Aug 12, 12:09 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar added a comment to T434334: cloudcephosd1043 drives are very very busy.

After bdev_enable_discard got rolled back ( https://gerrit.wikimedia.org/r/c/operations/puppet/+/1322878 ), the disk went well again on cloudcephosd1043 and the latency / disk business metrics I reported on T434024#12196170 look all fine now. Thank you!

Wed, Aug 12, 7:33 AM · Cloud-VPS, tools-infrastructure-team
hashar added a comment to T434566: deeptest Jenkins jobs became slower than they used to.

@Xqt thank you for the detailed reply!

Wed, Aug 12, 7:27 AM · ci-test-error, Pywikibot
hashar merged task T434515: CI on gerrit can't run git - citoid-pipeline-test into T434572: inference-services-pipeline-pre-commit-check failing with JDK error after Java upgrade.
Wed, Aug 12, 7:18 AM · Continuous-Integration-Config, ci-test-error
hashar merged T434515: CI on gerrit can't run git - citoid-pipeline-test into T434572: inference-services-pipeline-pre-commit-check failing with JDK error after Java upgrade.
Wed, Aug 12, 7:18 AM · Release-Engineering-Team, Jenkins, Continuous-Integration-Infrastructure

Tue, Aug 11

hashar added a comment to T434024: CI build slow and timing out sometimes.

Back in Spring 2019, I have discovered some hardware hosts were quite slower than others, some fancy Xeon servers were slower than my laptop (the one I still have, but maybe I bought an over powered one and that was a smart choice since I had it for 7/8 years now) or than previous machine, an Intel NUC that was even older than that. Anyway, I digress.

Tue, Aug 11, 9:00 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I then looked at the underlying hosts for each instances, their metrics can be found on https://grafana.wikimedia.org/d/000000377/host-overview

Tue, Aug 11, 8:42 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I pick 6893 and 6894 because they both started at 7:00 am UTC which should be low traffic and they ran on different hosts. We can look at how busy each hosts was while the builds were running using https://grafana.wmcloud.org/d/0g9N-7pVz/cloud-vps-project-board There is a panel showing the Pressure Stall Information (Kernel.org doc) which in short shows process waiting for CPU/IO/Memory.

Tue, Aug 11, 8:22 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

That table is amazing. That made me remember Jenkins provides the duration for the last 50 builds which can be found at https://integration.wikimedia.org/ci/job/wikimedia-fundraising-civicrm-bookworm/buildTimeTrend with a graph:

civicrm_builds.png (495×398 px, 121 KB)

So yeah things have improved a bit. We can see the build variances ranging between 15 and 26 minutes as is reflected in your table.

Tue, Aug 11, 8:02 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

I think we need a few builds to check the behavior, at least it is no more timing out so I guess things have improved. May you build a table of how long each steps take now so it can be compared with the Reference (fastest) column in this task description?

Tue, Aug 11, 4:23 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T177826: Upgrade CI Jenkins ssh key to ecdsa .

I am not working on it, last time I acted on that front was a couple years ago at T371930 in order to add a new ssh key and is a subtask hence why this one is stallen. The other task has fallen under the radar.

Tue, Aug 11, 4:19 PM · Jenkins, Collaboration-Services, Release-Engineering-Team (Seen), Continuous-Integration-Infrastructure, SRE
hashar closed T434572: inference-services-pipeline-pre-commit-check failing with JDK error after Java upgrade as Resolved.
Tue, Aug 11, 3:24 PM · Release-Engineering-Team, Jenkins, Continuous-Integration-Infrastructure
hashar added a project to T434572: inference-services-pipeline-pre-commit-check failing with JDK error after Java upgrade: Release-Engineering-Team.
java.io.IOException: Failed to exec spawn helper: pid: 251512, exit code: 1, error: 0 (none) 
Possible reasons:
  - Spawn helper ran into JDK version mismatch
  - Spawn helper ran into unexpected internal error
  - Spawn helper was terminated by another process
Possible solutions:
  - Restart JVM, especially after in-place JDK updates
  - Check system logs for JDK-related errors
  - Re-install JDK to fix permission/versioning problems
  - Switch to legacy launch mechanism with -Djdk.lang.Process.launchMechanism=VFORK
Tue, Aug 11, 3:23 PM · Release-Engineering-Team, Jenkins, Continuous-Integration-Infrastructure
hashar created T434566: deeptest Jenkins jobs became slower than they used to.
Tue, Aug 11, 2:58 PM · ci-test-error, Pywikibot
hashar claimed T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting.

And I think the problem is that mwext-codehealth-master-non-voting runs different containers which have those different cache configurations:

Very good finding @Peter ! I have enhanced the task description to give the whole context. The quick solution is to use quibble --skip-npm-install which was crafted by @Mhurd with T427922 :]

Tue, Aug 11, 2:14 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar added a parent task for T427922: Speedup the mwext-phpunit-coverage-publish castor job: T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting.
Tue, Aug 11, 1:43 PM · Patch-For-Review, Quibble, Test Platform (Nairobi 29), Castor
hashar added a subtask for T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting: T427922: Speedup the mwext-phpunit-coverage-publish castor job.
Tue, Aug 11, 1:43 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar updated the task description for T427822: Investigate multiple npm cache folders for mwext-codehealth-master-non-voting.
Tue, Aug 11, 1:42 PM · Test Platform, Continuous-Integration-Infrastructure, Castor
hashar added a comment to T434470: quibble-with-gated-extensions-vendor-mysql-php83 performance regression 2026-07-24/25.

@Peter you did a Quibble run on your local machine which reproduced the large delay. May you retry it with https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1324303 applied? (quibble --change 1324303 ...) should download and apply it).

Tue, Aug 11, 1:23 PM · MW-1.47-notes (1.47.0-wmf.17; 2026-08-25), Test Platform
hashar added a comment to T427922: Speedup the mwext-phpunit-coverage-publish castor job.

I have rolled Quibble 1.19.0 which now supports --skip-npm-install. @Mhurd can we sync up on deploying the change https://gerrit.wikimedia.org/r/c/integration/config/+/1321626 which updates the mwext-phpunit-coverage* jobs? Due to the timezone difference, maybe it is better done pairing with someone else than me (Release Engineering or Vaugh which I recently trained on the updating Jenkins jobs).

Tue, Aug 11, 1:19 PM · Patch-For-Review, Quibble, Test Platform (Nairobi 29), Castor
hashar added a parent task for T427471: Speedup mwext-codehealth-master-non-voting Castor job: T427450: Find the root cause for the long Castor wait times.
Tue, Aug 11, 1:14 PM · Patch-For-Review, Continuous-Integration-Infrastructure, Castor
hashar added a subtask for T427450: Find the root cause for the long Castor wait times: T427471: Speedup mwext-codehealth-master-non-voting Castor job.
Tue, Aug 11, 1:14 PM · Patch-For-Review, Castor, Test Platform (Basel 26)
hashar closed T433331: Experiment disabling castor cache saves for mwext-codehealth-master-non-voting for one week to measure castor queue impact as Declined.

An implementation was proposed as https://gerrit.wikimedia.org/r/c/integration/config/+/1294992 and abandoned. Rather than skipping the cache saving, we should fix the root causes of the slowness. That is T427450 / T427471.

Tue, Aug 11, 1:08 PM · Test Platform (Nairobi 29), Spike
hashar closed T433331: Experiment disabling castor cache saves for mwext-codehealth-master-non-voting for one week to measure castor queue impact, a subtask of T432684: Castor improvements, as Declined.
Tue, Aug 11, 1:08 PM · Castor, Test Platform (Epics), Epic
hashar added a project to T432684: Castor improvements: Castor.
Tue, Aug 11, 1:03 PM · Castor, Test Platform (Epics), Epic
hashar added a comment to T432671: investigate reducing the ~15GB castor cache of the codehealth jobs.

I am cleaning up the task related to mwext-codehealth-master-non-voting. I did the investigation in T427471, see T427471#11964274 and the following comments showing the issue are dupe npm cache and the large cypress files.

Tue, Aug 11, 1:03 PM · Test Platform (Nairobi 29), Spike
hashar merged T432671: investigate reducing the ~15GB castor cache of the codehealth jobs into T427471: Speedup mwext-codehealth-master-non-voting Castor job.
Tue, Aug 11, 1:01 PM · Patch-For-Review, Continuous-Integration-Infrastructure, Castor
hashar merged task T432671: investigate reducing the ~15GB castor cache of the codehealth jobs into T427471: Speedup mwext-codehealth-master-non-voting Castor job.
Tue, Aug 11, 1:01 PM · Test Platform (Nairobi 29), Spike
hashar added a comment to T434470: quibble-with-gated-extensions-vendor-mysql-php83 performance regression 2026-07-24/25.

T434024#12196170 is an issue with the Ceph filesystem which certainly caused a slowdown. That one started on July 24/25 and has tentatively been fixed on Friday August 7th 20:00 UTC by https://gerrit.wikimedia.org/r/c/operations/puppet/+/1322878 . I haven't looked at the performance of the civicrm job but that change at least got rid of timeouts though it is still slower than it used to be ( T434024#12196806 ). But I'll follow up on that other task. It is entirely possible there is another issue in the underlying infrastructure.

Tue, Aug 11, 7:48 AM · MW-1.47-notes (1.47.0-wmf.17; 2026-08-25), Test Platform

Mon, Aug 10

hashar added a comment to T434357: Add python-is-python3 to releng/node images.

I intentionally did not do this, as I wanted to flush out the remaining Python 2 uses and fix them individually. :-(

Mon, Aug 10, 4:10 PM · Release-Engineering-Team, ci-test-error, 3D
hashar added a comment to T434378: CI broken for 3d2png.git.

I have added the PipelineLib configuration to CI ( https://gerrit.wikimedia.org/r/c/integration/config/+/1323849 ) and did a recheck of the refreshed Blubber config https://gerrit.wikimedia.org/r/c/3d2png/+/516709 which missed a .pipeline/config.yaml file. It now has a single test pipeline. I guess the dev variant should be build as well, and maybe published to the Docker registry?

Mon, Aug 10, 7:37 AM · Continuous-Integration-Config, MediaWiki-Media-Platform-Team, 3D
hashar closed T434357: Add python-is-python3 to releng/node images, a subtask of T418403: Create dev enviroment for 3d2png, as Resolved.
Mon, Aug 10, 6:56 AM · Community-Tech (Sea Lion Squad), 3D
hashar closed T434357: Add python-is-python3 to releng/node images as Resolved.

Thank you for having split the case of 3d2png to a standalone task (T434378)

Mon, Aug 10, 6:56 AM · Release-Engineering-Team, ci-test-error, 3D

Sat, Aug 8

hashar added a comment to T434357: Add python-is-python3 to releng/node images.

I have build the images and update the generic-node24 job via https://gerrit.wikimedia.org/r/c/integration/config/+/1323667 . I did not update the other jobs that will be for Monday, hence I am leaving this task open for now.

Sat, Aug 8, 8:53 PM · Release-Engineering-Team, ci-test-error, 3D

Fri, Aug 7

hashar added a comment to T434024: CI build slow and timing out sometimes.

I went to look at wmcs dashboards, there are several of them for Ceph/OSD, given I know nothing about about the infra, I am pasting below my findings.

Fri, Aug 7, 5:37 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar updated subscribers of T434024: CI build slow and timing out sometimes.
Fri, Aug 7, 3:35 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

And from a Quibble / MediaWiki job, the daily average rose on July 25th (view over last 30 days and civicrm view over last 30 days:

Fri, Aug 7, 3:30 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar added a comment to T434024: CI build slow and timing out sometimes.

And the changes we made to the docker-registry.wikimedia.org/releng/civicrm image:

civicrm (0.9-s4) wikimedia; urgency=high
Fri, Aug 7, 3:22 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog
hashar updated subscribers of T434024: CI build slow and timing out sometimes.

i think @Mhurd and @Peter noticed the MediaWiki job became slower fairly recently.

Fri, Aug 7, 3:11 PM · Continuous-Integration-Infrastructure, Release-Engineering-Team, Fundraising Tech - Chaos Crew, Wikimedia-Fundraising-CiviCRM, Fundraising-Backlog

Thu, Aug 6

hashar added a comment to T434187: CI runs for repos with codesniffer 51 fail due to PKSA-rdkp-vv9z-mjkg.

Hello Given this affects PHP CodeSniffer which is the styling tool we use across a thousand of MediaWiki repositories × a few supported branches (releases, wmf, fundraising) that is going to be a few thousands of changes to upload to Gerrit and pass through CI.

Thu, Aug 6, 1:47 PM · MW-1.47-notes (1.47.0-wmf.14; 2026-08-04), MW-1.46-release, MW-1.43-release, MW-1.45-release, Patch-For-Review, MediaWiki-Codesniffer, Composer, ci-test-error
hashar edited projects for T427922: Speedup the mwext-phpunit-coverage-publish castor job, added: Quibble; removed Patch-For-Review.
Thu, Aug 6, 1:27 PM · Patch-For-Review, Quibble, Test Platform (Nairobi 29), Castor
hashar added a watcher for Castor: hashar.
Thu, Aug 6, 1:26 PM
hashar added a comment to T305571: Set content model JSON for Web2Cit configuration pages in metawiki's "Main" namespace.

Well done, I have approved the change which has been merged. It should deploy automatically on the beta cluster infrastructure over the next 10 minutes :]

Thu, Aug 6, 1:14 PM · Patch-For-Review, JsonConfig, Wikimedia-Site-requests, Web2Cit

Tue, Aug 4

hashar created T433980: PCC: no nodes found for class: Class/Role::Jenkins.
Tue, Aug 4, 1:51 PM · Infrastructure-Foundations, Puppet CI

Mon, Aug 3

hashar awarded Continuous Integrator to recipient: vaughnwalters.
Mon, Aug 3, 4:44 PM

Thu, Jul 30

hashar created T433615: Grant Access to ciadmin for Vaughn Walters.
Thu, Jul 30, 2:51 PM · SRE, LDAP-Access-Requests

Wed, Jul 29

hashar closed T432791: [Request] Add me to trusted contributor in integration/config as Resolved.

Hello @Les4353, I have added you to the Trusted Contributors group in Gerrit, it brings you additional rights such as editing hashtags, topics, rebase or add a patch to another person change (use those two carefully!).

Wed, Jul 29, 12:44 PM · Continuous-Integration-Config

Mon, Jul 27

hashar added a comment to T383047: Could not send confirmation email: Unknown error in PHP's mail() function..

Thank you! I confirm I have managed to send an email via the wikis. The page took a while to process the send (which I guess correspond to the 4 seconds delay)

Mon, Jul 27, 8:37 PM · MW-1.44-notes, MW-1.46-notes, MW-1.45-notes, MW-1.47-notes (1.47.0-wmf.9; 2026-06-30), MW-1.43-notes, Patch-For-Review, Product Safety and Integrity, MediaWiki-extensions-EmailAuth, MediaWiki-Email, Mail, Infrastructure-Foundations, MediaWiki-Core-Platform-Team, MediaWiki-User-login-and-signup, Wikimedia-production-error
hashar added a comment to T429901: The email verification message from gerrit does not contain a clickable link.

https://gerrit.wikimedia.org/r/1316125 got deployed and the new template should have been taken in consideration automatically. The template will be carried as we upgrade Gerrit. We will just have to remember to remove it from our Puppet repository once we upgrade to a version of Gerrit that has it, but it is not a big issue.

Mon, Jul 27, 9:48 AM · RoadToWiki, Gerrit

Mon, Jul 20

hashar added a comment to T305571: Set content model JSON for Web2Cit configuration pages in metawiki's "Main" namespace.

I gave some links to @diegodlh over IRC:

Mon, Jul 20, 3:00 PM · Patch-For-Review, JsonConfig, Wikimedia-Site-requests, Web2Cit
hashar updated subscribers of T406169: Setup a daily job on jenkins to run ReadingLists tests.

TLDR: given the job running core+extensions together is now way faster (thanks @Peter & all), I think we can decline this and instead add ReadingLists to the gated extensions which is T403560 .

Mon, Jul 20, 10:18 AM · OKR-Work, Reader Experience Team (REx Sprint 2 [Q1 FY26-27 July 15-28]), FY252627 Reading Lists (Graduating from Beta), MediaWiki-extensions-ReadingLists

Jul 14 2026

hashar added a comment to T431663: integration.wikimedia.org should have a FOSS license.

The code we wrote is now placed under GPL2.0 or later

Jul 14 2026, 8:41 AM · doc.wikimedia.org, Software-Licensing, Continuous-Integration-Infrastructure

Jul 11 2026

hashar added a comment to T372404: Gerrit's syntax highlighting for PHP code breaks when encountering an apostrophe in a // comment in a function call.

@Paladox has let me know (thank you) that upstream has released 11.11.2 in June 2026 which includes my fix.

Jul 11 2026, 2:51 PM · Upstream, Gerrit

Jul 9 2026

hashar added a comment to T430775: Validate Gerrit backup recoverability after the fileset exclusions.

The backup is being conducted while the service is actively writing to LFS object files (rarely happens), git packfiles and objects. We also have git garbage collection triggering on a weekly basis.

Jul 9 2026, 3:07 PM · Collaboration-Services, Gerrit
hashar claimed T411583: Gerrit backups are growing.

From my list of action on T411583#12074321

Jul 9 2026, 3:00 PM · Collaboration-Services, Gerrit
hashar closed T430774: Compact Gerrit H2 databases to reclaim disk as Resolved.

This happens automatically when Gerrit is being stopped (more exactly when the H2 DB Driver is about to disconnect from the db). We had multiple controlled restart of Gerrit over the last few days and the files got shrunk:

-rw-r--r-- 1 gerrit gerrit 762M Jul  9 14:50 git_file_diff.h2.db
-rw-r--r-- 1 gerrit gerrit 767M Jul  9 14:50 gerrit_file_diff.h2.db
-rw-r--r-- 1 gerrit gerrit 774M Jul  9 14:50 conflicts.h2.db
-rw-r--r-- 1 gerrit gerrit 785M Jul  9 14:48 diff_intraline.h2.db
-rw-r--r-- 1 gerrit gerrit 832M Jul  9 14:49 change_kind.h2.db
-rw-r--r-- 1 gerrit gerrit 845M Jul  9 14:50 mergeability.h2.db
-rw-r--r-- 1 gerrit gerrit 1.9G Jul  9 14:49 diff_summary.h2.db
-rw-r--r-- 1 gerrit gerrit 2.2G Jul  9 14:49 comment_context.h2.db
Jul 9 2026, 2:53 PM · Release-Engineering-Team, Gerrit, Collaboration-Services
hashar closed T430774: Compact Gerrit H2 databases to reclaim disk, a subtask of T411583: Gerrit backups are growing, as Resolved.
Jul 9 2026, 2:53 PM · Collaboration-Services, Gerrit
hashar closed T430774: Compact Gerrit H2 databases to reclaim disk, a subtask of T425667: Investigate Gerrit root disk usage and logging, as Resolved.
Jul 9 2026, 2:53 PM · Gerrit, Collaboration-Services
hashar closed T257744: Decide if Gerrit's indices should get backed up as Resolved.

We have historically not backed up the search indices. We had them backed up by mistake as part of some path change and they take a sizeable amount of disk space, they are also ever growing (T411583)

Jul 9 2026, 2:48 PM · Gerrit
hashar added a project to T431663: integration.wikimedia.org should have a FOSS license: doc.wikimedia.org.

+ doc.wikimedia.org since the repo is shared with https://doc.wikimedia.org/

Jul 9 2026, 2:36 PM · doc.wikimedia.org, Software-Licensing, Continuous-Integration-Infrastructure
hashar closed T430829: 1.47.0-wmf.10 deployment blockers as Resolved.

I am claiming 1.47.0-wmf.10 to have successfully rolled out.

Jul 9 2026, 2:26 PM · Essential-Work, Release-Engineering-Team (Priority Backlog 📥), Release, Train Deployments
hashar added a comment to T352319: castor-save-workspace-cache aborted during postbuild.

Self note review the two cases in T352319#10288944

Jul 9 2026, 1:29 PM · Upstream, Castor, Jenkins, Continuous-Integration-Infrastructure
hashar merged task T419488: PostBuild changing the status of successful builds to failure for no apparent reason into T352319: castor-save-workspace-cache aborted during postbuild.
Jul 9 2026, 1:27 PM · ci-test-error (WMF-deployed Build Failure), Continuous-Integration-Infrastructure, Castor, Continuous-Integration-Config
hashar merged T419488: PostBuild changing the status of successful builds to failure for no apparent reason into T352319: castor-save-workspace-cache aborted during postbuild.
Jul 9 2026, 1:27 PM · Upstream, Castor, Jenkins, Continuous-Integration-Infrastructure
hashar added a comment to T419488: PostBuild changing the status of successful builds to failure for no apparent reason.

I am pretty sure that is the same as T352319 and is an issue within the Jenkins Parameterized Build plugin

Jul 9 2026, 1:27 PM · ci-test-error (WMF-deployed Build Failure), Continuous-Integration-Infrastructure, Castor, Continuous-Integration-Config
hashar added a comment to T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.

🎉 Thank you for having confirmed the resolution. Very much appreciated!

Jul 9 2026, 12:34 PM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar renamed T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail from contint1003 is missing docker-buildx, causing blubber-based CI pipelines to fail to contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.
Jul 9 2026, 11:52 AM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar added a comment to T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.

If I try to push an existing image on the old host:

contint1002$ docker --config /etc/docker-pusher push docker-registry.wikimedia.org/releng/release-notes:0.1.0-s1
The push refers to repository [docker-registry.wikimedia.org/releng/release-notes]
...
357ab391beed: Layer already exists 
unauthorized: authentication required
Jul 09 11:17:52 contint1002 dockerd[1160]: time="2026-07-09T11:17:52.996083076Z" level=info msg="Attempting next endpoint for push after error: unauthorized: authentication required"
Jul 9 2026, 11:21 AM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar added a comment to T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.

I have confirmed the credential file has the same auth information on the old host (contint1002) and the new host (contint1003). It is provisioned by Puppet.

Jul 9 2026, 11:10 AM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar added a comment to T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.

The docker daemon has:

Jul 09 09:41:01 contint1003 dockerd[1293]:
time="2026-07-09T09:41:01.925689187Z"
level=error
msg="Upload failed: unauthorized: <html>\r\n<head><title>401 Authorization Required</title></head>\r\n<body>\r\n<center><h1>401 Authorization Required</h1></center>\r\n<hr><center>nginx/1.22.1</center>\r\n</body>\r\n</html>\r\n"
Jul 9 2026, 10:55 AM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar added a comment to T431582: contint1003 is missing docker-buildx / auth to registry fails, causing blubber-based CI pipelines to fail.

Thanks @kevinbazira! The build ran on contint1003 which is a production agent and as such should have the credentials to push to the registry (to be verified). From the console log, the build does:

+ sudo /usr/local/bin/docker-pusher docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services-tts:2026-07-09-094058-publish
The push refers to repository [docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services-tts]
...
0c28ce2f5fb5: Waiting
unauthorized: <html>
<head><title>401 Authorization Required</title></head>
...
Jul 9 2026, 10:50 AM · Release-Engineering-Team (Doing 😎), Collaboration-Services, Jenkins, Continuous-Integration-Infrastructure
hashar reopened T431675: OpenSearch mediawiki new errors filters out some errors as "Open".

TLDR must_not without Exclude results ends up being an OR filter.

Jul 9 2026, 10:02 AM · Release-Engineering-Team (Doing 😎), Deployments
hashar added a comment to T431675: OpenSearch mediawiki new errors filters out some errors.

must_not

Logical not operator. All matches are excluded from the results. If must_not has multiple clauses, only documents that do not match any of those clauses are returned.
For example, "must_not":[{clause_A}, {clause_B}] is equivalent to NOT(A OR B).

Jul 9 2026, 9:59 AM · Release-Engineering-Team (Doing 😎), Deployments
hashar added a comment to T431675: OpenSearch mediawiki new errors filters out some errors.

Good catch, and for the record the filter had the custom label NOT T394850 - JsonConfig: Undefined array key with the OpenSearch Query DSL being:

 {
  "query": {
    "bool": {
      "must_not": [
        {
          "match_phrase": {
            "error.stack_trace.text": "extensions/JsonConfig/includes/GlobalJsonLinksQuery.php(274)"
          }
        },
        {
          "match_phrase": {
            "error.message": "PHP Warning: Undefined array key"
          }
        }
      ]
    }
  }
}
Jul 9 2026, 9:46 AM · Release-Engineering-Team (Doing 😎), Deployments
hashar updated the task description for T431675: OpenSearch mediawiki new errors filters out some errors.
Jul 9 2026, 9:25 AM · Release-Engineering-Team (Doing 😎), Deployments
hashar created T431675: OpenSearch mediawiki new errors filters out some errors.
Jul 9 2026, 9:23 AM · Release-Engineering-Team (Doing 😎), Deployments