Page MenuHomePhabricator

Investigate Quickstatements Outage
Open, Needs TriagePublic

Description

We've had two alerts fire (and resolve).

I see the pods have also restarted 19 times with OOM (since the pod was created on Monday).

Looking at the google logs it appears all recent traffic is to https://memory-prime.wikibase.cloud. I suspect, perhaps, they are trying to import more in a single request than we can handle or similar.

Event Timeline

User reported trying to create large items with >700 statements. They attempted to stop creating items of such size. We had at least one more restart today with big but slightly smaller items that snuck through the filter of around 658 statements. We'll keep this under review until the restarts have stopped for at least 24hrs.

We should also follow up with a ticket to make it impossible for users to cause this issue; we should gracefully handle these larger entities.

The type of batch causing the issue can be found here:
https://memory-prime.wikibase.cloud/wiki/Item_talk:Q179

Feel free to run it for testing.