Description
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Open | None | T353925 gate-and-submit backlogged due to waiting for castor-save-workspace-cache | |||
| Open | Peter | T427450 Find the root cause for the long Castor wait times | |||
| Open | None | T427471 Speedup mwext-codehealth-master-non-voting Castor job | |||
| Resolved | Peter | T427822 Investigate multiple npm cache folders for mwext-codehealth-master-non-voting | |||
| Resolved | Mhurd | T427922 Speedup the mwext-phpunit-coverage-publish castor job | |||
| Open | Mhurd | T427752 Investigate why the job median time regressed in May for quibble-with-gated-extensions-vendor-mysql-php83 |
Event Timeline
Happening again in similar circumstances...
Do we really only have one host doing one of these jobs at a time?
This is still happening, and often a major bottleneck when there are many changes being submitted, with waits in the order of minutes.
Mentioned in SAL (#wikimedia-releng) [2025-05-06T16:16:16Z] <hashar> restarting CI Jenkins due to a deadlock affecting castor-save-workspace which ends up blocking jobs # T353925
@Peter rediscovered it as one job has a very large cache. The investigation has been conducted at T427450
@Mhurd was investigating CI jobs slowness and found the IO rate limiting on the WMCS instance to be problematic: T427752#12037841
Lets reuse T353925 as the main parent tracking task. I am adding the two other tasks as sub tasks.



