Query killer didn't kill slow queries during the s4 incident (T398448) which prolonged the incident. We should make sure it's properly installed in all of s4 replicas (or better, check it against all core sections)
Description
Description
Related Objects
Related Objects
Event Timeline
Comment Actions
Was a case of query killer not being installed or the "normal" issue we experienced before: when the host is super overloaded the query killer cannot be fast enough to kill them?
Comment Actions
To my understanding, the former. Or at least it was installed but didn't kick in. I quickly killed all of them manually.
Comment Actions
I guess someone fixed it? I just checked and the query killer is present everywhere. Per the monitoring, we've already had this: https://phabricator.wikimedia.org/T254738