Cloudelastic seems to complain about a high fix rate, we should take a look to understand the cause of this.
Description
Event Timeline
silenced the alert for 30 days: https://alerts.wikimedia.org/?q=%40silenced_by%3Dec1062b5-c0b3-41e4-b4c2-d27be3c0bce2
We suspect the issue is when Sanitizer fails to connect to Elasticsearch. We know this is sometime an issue. Let's fix this error case first and remove that noise from the data.
Maybe I already fixed the mentioned issue with connection errors and simply forgot. I tried reproducing the error's we had in the past but could not. On review on the related code I found https://gerrit.wikimedia.org/r/c/mediawiki/extensions/CirrusSearch/+/980955 which fixes the issue i was thinking about. That means there is some unique problem going on here.
Separately, the graphs for cloudelastic are now back in-line with the other clusters. I was way up for a few weeks, but then came back down. Maybe it was a real problem that got cleaned up? But what was the source?
Indeed, https://grafana-rw.wikimedia.org/d/2DIjJ6_nk/cirrussearch-saneitizer-historical-fix-rate?orgId=1 shows a bump mid October and is now back to "normal", cause is unknown and might be hard to investigate now, boldly closing.