Page MenuHomePhabricator

Investigate significant changeprop backlogs after PCS migration
Closed, ResolvedPublic

Description

After migrating most page requests out of restbase and directly into PCS we are seeing significant backlogs of messages developing in changeprop. These grow to very significant levels and then get zeroed (possibly due to events expiring?).

Backlogs:

image.png (1,919×998 px, 216 KB)

Details

Related Changes in Gerrit:
SubjectAuthorRepoBranchLines +/-
Jgiannelosmediawiki/extensions/EventBusmaster+57 -1
Jgiannelosoperations/mediawiki-configmaster+4 -0
Jgiannelosoperations/deployment-chartsmaster+2 -1
Jgiannelosmediawiki/services/mobileappsmaster+17 -0
Jgiannelosmediawiki/coremaster+139 -23
Jgiannelosmediawiki/services/mobileappsmaster+6 -8
Hnowlanoperations/deployment-chartsmaster+2 -2
Jgiannelosoperations/deployment-chartsmaster+3 -3
Hnowlanoperations/deployment-chartsmaster+2 -2
Hnowlanoperations/deployment-chartsmaster+9 -4
Hnowlanoperations/deployment-chartsmaster+9 -2
Hnowlanmediawiki/services/change-propagationmaster+1 -0
Hnowlanoperations/deployment-chartsmaster+14 -3
Hnowlanoperations/deployment-chartsmaster+8 -0
Jgiannelosoperations/deployment-chartsmaster+6 -1
Hnowlanoperations/deployment-chartsmaster+1 -1
Hnowlanoperations/deployment-chartsmaster+6 -5
Hnowlanoperations/deployment-chartsmaster+3 -3
Hnowlanoperations/deployment-chartsmaster+4 -0
Hnowlanoperations/deployment-chartsmaster+2 -1
Show related patches Customize query in gerrit

Event Timeline

Since merging the changes to disable restbase parsoid/page rules in changeprop yesterday, topic: eqiad.change-prop.transcludes.resource-change group: changeprop-pcs_rerender_mobile_html_native_transcludes is seeing runaway growth in all partitions. This topic is possibly not being processed or is being populated unnecessarily

image.png (2,115×1,224 px, 192 KB)

Change #1160045 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: increase concurrency on pcs_rerender_mobile_html_native

https://gerrit.wikimedia.org/r/1160045

Change #1160045 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: increase concurrency on pcs_rerender_mobile_html_native

https://gerrit.wikimedia.org/r/1160045

Change #1160052 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: increase memory limit

https://gerrit.wikimedia.org/r/1160052

Change #1160052 abandoned by Hnowlan:

[operations/deployment-charts@master] changeprop: increase memory, cpu limit

Reason:

values-production has significantly higher values already

https://gerrit.wikimedia.org/r/1160052

Change #1160089 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop, mobileapps: bump pcs job concurrency, mobileapps replicas

https://gerrit.wikimedia.org/r/1160089

Change #1160089 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop, mobileapps: bump pcs job concurrency, mobileapps replicas

https://gerrit.wikimedia.org/r/1160089

Change #1160704 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: bump concurrency for pcs_rerender_native_on_null

https://gerrit.wikimedia.org/r/1160704

Change #1160704 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: bump concurrency for pcs_rerender_native_on_null

https://gerrit.wikimedia.org/r/1160704

topic: eqiad.change-prop.transcludes.resource-change group: changeprop-pcs_rerender_mobile_html_native_transcludes still has a significant backlog but is slowly going down due to increased concurrency and resources allocated. The same is true of pcs_rerender_native_on_null although the backlog is a lot smaller and it will recover much quicker.

The following topics have a significant backlog that will expire but never be processed as these events are not being consumed any longer after the retirement of restbase-related resource_change events:
changeprop-mw_purge
changeprop-summary_definition_rerender
changeprop-mobile-html_rerender
changeprop-media-list_rerender

Change #1160727 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: bump replicas

https://gerrit.wikimedia.org/r/1160727

Change #1160727 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: bump replicas

https://gerrit.wikimedia.org/r/1160727

Change #1160753 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[operations/deployment-charts@master] changeprop: Add header with event timestamp for PCS requests

https://gerrit.wikimedia.org/r/1160753

Change #1160751 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[mediawiki/services/mobileapps@master] RB sunset: Reduce changeprop backlog by rejecting old event

https://gerrit.wikimedia.org/r/1160751

Change #1160897 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[operations/deployment-charts@master] RB sunset: Abandon event processing for PCS events older than cache TTL

https://gerrit.wikimedia.org/r/1160897

Change #1160897 merged by jenkins-bot:

[operations/deployment-charts@master] RB sunset: Configure claim TTL for PCS related endpoints

https://gerrit.wikimedia.org/r/1160897

Change #1161467 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] admin_ng: increase changeprop resource quotas

https://gerrit.wikimedia.org/r/1161467

Change #1161467 merged by jenkins-bot:

[operations/deployment-charts@master] admin_ng: increase changeprop resource quotas

https://gerrit.wikimedia.org/r/1161467

Change #1161559 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: pcs concurrency config at values level, bump native transcludes

https://gerrit.wikimedia.org/r/1161559

Change #1161565 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[mediawiki/services/change-propagation@master] base_executor: emit abandoned metric

https://gerrit.wikimedia.org/r/1161565

Change #1161559 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: pcs concurrency config at values level, bump native transcludes

https://gerrit.wikimedia.org/r/1161559

Change #1161565 merged by jenkins-bot:

[mediawiki/services/change-propagation@master] base_executor: emit abandoned metric

https://gerrit.wikimedia.org/r/1161565

Change #1161893 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: emit abandoned events metric

https://gerrit.wikimedia.org/r/1161893

Change #1161957 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: implement batch_size parameter for pcs job

https://gerrit.wikimedia.org/r/1161957

Change #1161893 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: emit abandoned events metric

https://gerrit.wikimedia.org/r/1161893

Change #1161957 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: implement batch_size parameter for pcs job

https://gerrit.wikimedia.org/r/1161957

Change #1163345 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] mobileapps: increase memory limit, drop replicas

https://gerrit.wikimedia.org/r/1163345

Change #1163345 merged by jenkins-bot:

[operations/deployment-charts@master] mobileapps: increase memory limit, drop replicas

https://gerrit.wikimedia.org/r/1163345

Do we know whether we are generating the correct number of events? It appears that a sustained period of ~3k resource-change events over 2.5 hours results in close to 4 million events in a single partition

image.png (1,053×786 px, 105 KB)

T398243 is probably related to this issue

hnowlan triaged this task as High priority.Jul 7 2025, 10:24 AM

Change #1168041 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[operations/deployment-charts@master] changeprop: Ignore more namespace on pcs transclusion rules

https://gerrit.wikimedia.org/r/1168041

Change #1168041 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: Ignore more namespaces on pcs transclusion rules

https://gerrit.wikimedia.org/r/1168041

Change #1168144 had a related patch set uploaded (by Hnowlan; author: Hnowlan):

[operations/deployment-charts@master] changeprop: correct amended regex

https://gerrit.wikimedia.org/r/1168144

Change #1168144 merged by jenkins-bot:

[operations/deployment-charts@master] changeprop: correct amended regex

https://gerrit.wikimedia.org/r/1168144

Some context around a possible root cause:

Before when parsoid was served via RB on each parsoid pregen request:

  • rb fetched content from cassandra
  • rb fetched content from parsoid
  • compared the html
  • if the html was unchanged it didn't continue the chain of events that would eventually pregenerate pcs

now we don't have that and instead we use mediawiki events directly (eg. mediawiki.revision_create).
But transclusions (on a stateless environment like changeprop) cannot be optimized for unchanged content (edited)

I think the need for a proper parser output content changed event is more pressing at this point (edited)
Some more ideas on temporary workarounds:

  • Avoid caching big categories of unused content
    • We already started with File: namespace and we saw a huge increase in performance
      • Backlog was practically halved
  • Increase concurrency of the pcs on transclusion matcher on changeprop
  • Increase pods on k8s
  • Increase concurrency on pcs service runner
    • In theory we are way more efficient with purges instead of pregen because we don't do any IO for MW apis
    • We only do a fire and forget delete in cassandra and send eventgate events
  • Make rules more specific to the domain we target
    • eg. commons changeprop events are only affecting summary endpoint but instead we send pregen/purge requests for all endpoints and then PCS would raise a 404: Domain not supported.
    • In the scale we operate, matching the rules, occupying concurrency resources, and sending thousand of requests to fail is a legit overhead

Ultimately I believe that we need a specific single event for parser output content changed:

  • Purge/pregen only what is actually changed
  • Simplify spaghetti rules on changeprop
  • Optimize based on namespace

Here are the amount of requests that would avoid unneeded pregen traffic from Parsoid-on-RB days:
https://grafana.wikimedia.org/goto/ar50jqyNg?orgId=1
This x3 are the events overhead because on PCS we have {mobile-html, media-list, summary}

Change #1169036 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[mediawiki/core@master] Run hook on parser cache output change

https://gerrit.wikimedia.org/r/1169036

Change #1169665 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[mediawiki/extensions/EventBus@master] Send resource change event on parseroutput content change

https://gerrit.wikimedia.org/r/1169665

Change #1170174 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[operations/mediawiki-config@master] Configure stream for parser cache change events

https://gerrit.wikimedia.org/r/1170174

Change #1170269 had a related patch set uploaded (by Jgiannelos; author: Jgiannelos):

[mediawiki/services/mobileapps@master] caching: Stop sending resource change events

https://gerrit.wikimedia.org/r/1170269

Change #1170269 merged by jenkins-bot:

[mediawiki/services/mobileapps@master] caching: Stop sending resource change events

https://gerrit.wikimedia.org/r/1170269

Jgiannelos closed this task as Resolved.EditedAug 7 2025, 9:51 AM
Jgiannelos claimed this task.

Backlog has been consistency manageable for the past few weeks. The reasons behind this are:

  • High traffic topics are using cache invalidation instead of pregeneration
    • We only invalidate caches on request and we don't continue with the req/res cycle with MW api requests, parsing, rendering etc
  • We don't send pcs resource change events
    • Before we were flooding the resource change event topic with resource changes from PCS too
    • That caused increase in message traffic and on top of that processing in order to discard them from processing
  • We exclude some namespaces from invalidation
    • Traffic is pretty low for the excluded namespaces
    • Eventually cached content will be evicted because of TTL so in the range of ~7 days content will be fresh

The proper solution is to have a specific event from MW that signals a parser output content change. But closing this ticket for now.

Change #1169036 abandoned by Jgiannelos:

[mediawiki/core@master] Run hook on parser cache output change

Reason:

Closing this one, its preferred to use events instead

https://gerrit.wikimedia.org/r/1169036

The proper solution is to have a specific event from MW that signals a parser output content change. But closing this ticket for now.

Chiming in to say that this would be beneficial for more than just PCS! A single mediawiki.page_rendering_change or something (modeled on mediawiki/state event schema fragments) would be useful for things like T360794: Event stream with latest revision HTML & parent revision HTML diff and other potential 'derived data' usage outside of MediaWiki.

Change #1160751 abandoned by Jgiannelos:

[mediawiki/services/mobileapps@master] RB sunset: Reduce changeprop backlog by rejecting old event

https://gerrit.wikimedia.org/r/1160751

Change #1160753 abandoned by Jgiannelos:

[operations/deployment-charts@master] changeprop: Add header with event timestamp for PCS requests

https://gerrit.wikimedia.org/r/1160753

Change #1170174 abandoned by Jgiannelos:

[operations/mediawiki-config@master] Configure stream for parser cache change events

https://gerrit.wikimedia.org/r/1170174

Change #1169665 abandoned by Jgiannelos:

[mediawiki/extensions/EventBus@master] Send events on parser cache change

https://gerrit.wikimedia.org/r/1169665