Page MenuHomePhabricator

Incident: 2022-12-09 api appserver worker starvation
Closed, ResolvedPublic

Event Timeline

Clement_Goubert renamed this task from Incident: 2022-12-12 api appserver worker starvation to Incident: 2022-12-09 api appserver worker starvation.Dec 12 2022, 5:16 PM

Tagging other responders to help fill out the incident report.

Change 867597 had a related patch set uploaded (by Clément Goubert; author: Clément Goubert):

[operations/deployment-charts@master] Revert "eventgate-analytics: bump replicas from 20 to 30"

https://gerrit.wikimedia.org/r/867597

Since this incident was caused by a temporary raise in logging volume, and our response was to scale up the bottleneck, this task will only track reverting to the state before the incident.

reverting to the state before the incident

Hm, do we need to revert? I don't mind either way, but +10 replicas (in each DC), probably is okay and will allow us to handle more load if something like this happens again?

reverting to the state before the incident

Hm, do we need to revert? I don't mind either way, but +10 replicas (in each DC), probably is okay and will allow us to handle more load if something like this happens again?

I thought we had talked about reverting when the train had passed reverting the root cause, but checking back in my logs, I misremembered.

We have the resources to keep it at 30 replicas at the moment in my opinion. @JMeybohm what do you think?

Hi, The flood of logs is still incoming, the revert of logspam has not been deployed yet so I advise against reducing pods now.

Yes, the commit message of the above changelog makes it very clear it is not to be deployed before the train has ran on Thursday. We might end up not scaling back anyways.

Change 867597 abandoned by Clément Goubert:

[operations/deployment-charts@master] Revert "eventgate-analytics: bump replicas from 20 to 30"

Reason:

https://gerrit.wikimedia.org/r/867597

We have the resources to keep it at 30 replicas at the moment in my opinion. @JMeybohm what do you think?

Sorry, totally missed that. Yes, resource wise we are absolutely fine keeping the 30 replicas.

No worries, I took a look at the resources and it seemed fine to leave it like that. We can always revisit when/if contention happens with other services.

Joe subscribed.

Removing the sustainability tag as it doesn't seem like there is any related actionable here. @Clement_Goubert if you think that's unfair please undo my change.

No more action needed on this incident.