Page MenuHomePhabricator

Reduce the amount of messages sent through channel:Memcached during failures
Open, Needs TriagePublic

Assigned To
None
Authored By
jijiki
Mar 31 2025, 10:54 AM
Referenced Files
F58951543: image.png
Mar 31 2025, 10:54 AM
F58951539: image.png
Mar 31 2025, 10:54 AM
F58951536: image.png
Mar 31 2025, 10:54 AM
F58951492: image.png
Mar 31 2025, 10:54 AM
F58951554: image.png
Mar 31 2025, 10:54 AM

Description

During this incident, 2025-03-12 ExternalStorage Database Cluster Overload, MediaWiki emitted millions of events to channel:Memcached, for thousands of requests.

In this incident, mw-{api-int,parsoid, jobrunner} were trying to connect to an invalid memcached server address, no to avail. Thus, they sent out millions of messages per 5 mins.

image.png (1,294×212 px, 56 KB)

image.png (2,276×1,452 px, 496 KB)

During this particular incident, this behaviour overwhelmed logstash, causing delays and alerts not being fired.

Below we can see envoy's POV. While I have not done the math yet, but my imprerssion is that we are logging quite aggressively.

image.png (2,996×850 px, 330 KB)

image.png (2,972×860 px, 356 KB)

image.png (2,966×878 px, 305 KB)