Page MenuHomePhabricator

gate-and-submit backlogged due to waiting for castor-save-workspace-cache
Open, Needs TriagePublic

Description

Screenshot 2023-12-22 at 02.38.18.png (347×881 px, 107 KB)

Screenshot 2023-12-22 at 02.38.50.png (428×755 px, 73 KB)

Event Timeline

Happening again in similar circumstances...

Do we really only have one host doing one of these jobs at a time?

Screenshot 2024-01-05 at 19.52.07.png (178×1,173 px, 97 KB)

Screenshot 2024-01-05 at 19.52.27.png (369×980 px, 82 KB)

This is still happening, and often a major bottleneck when there are many changes being submitted, with waits in the order of minutes.

Mentioned in SAL (#wikimedia-releng) [2025-05-06T16:16:16Z] <hashar> restarting CI Jenkins due to a deadlock affecting castor-save-workspace which ends up blocking jobs # T353925

@Peter rediscovered it as one job has a very large cache. The investigation has been conducted at T427450

@Mhurd was investigating CI jobs slowness and found the IO rate limiting on the WMCS instance to be problematic: T427752#12037841

Lets reuse T353925 as the main parent tracking task. I am adding the two other tasks as sub tasks.