Page MenuHomePhabricator

Investigate why mobileapps and proton required a manual restart
Open, Needs TriagePublic

Description

During incident, mobileapps and proton required a manual restart, while it was not expected.

Additional context is available in the incident document, but a failure of mw-api-int caused mobileapps and proton to continue to serve a level of errors (around 20 RPS) despite the service having recovered. This could be queued requests with invalid responses stored, cached IP addresses or other stale state maintained by the service. The fact it has happened to both services might indicate an issue with a shared library in service-runner. This persisted until a restart

Action item extracted from incident 2025-09-11 mwapi error rate/DB failure

Incident document: https://docs.google.com/document/d/1gKZR-5zWTl22EcNkab0suOuYkHufzGWsVW9pr9_ZvEo/edit?tab=t.0#heading=h.nz4dlhpgbsjm

Event Timeline

@hnowlan for this and T417059 , you were Incident Coordinator, do you recall more context?

hnowlan updated the task description. (Show Details)
hnowlan subscribed.