Page MenuHomePhabricator

Reduce secondary caching of fragments under the WikiLambdaFunctionCall key to the minimum necessary
Closed, ResolvedPublic

Description

Description

Orchestrator responses stored under the WikiLambdaFunctionCall key are large objects averaging ~3400 bytes, compared to ~100 bytes for the fragment itself.
These entries are keyed with today's date, so they become orphaned after one day and are never retrieved again.

The scale of the problem becomes clear when projected across realistic topic and locale counts. For example, assuming a reasonable number of 50 fragments per page, and following phase projections for Abstract Wikipedia integration, this would be the space needed to host integrated fragments:

PhaseTopicsLocalesFragmentsWikiLambdaAbstractFragment sizeWikiLambdaFunctionCall size
Phase 11,0003150,000~116 MB~14 GB
Phase 210,0002512,500,000~9 GB~1.2 TB
Phase 310,000,000300150,000,000,000~110 TB~14 PB

For every 116 MB of necessary fragment data, we are storing ~14 GB of redundant orchestrator responses. At Phase 1 scale alone, WikiLambdaFunctionCall occupies roughly 120x the storage of the actual fragment cache, for data that is unreachable within 24 hours. See more details of this analysis.

We are not fully removing the key, as it provides some safeguard against (D)DOS attacks on the orchestrator. A TTL of one minute is sufficient to absorb request bursts while clearing the redundant data early enough to have no meaningful impact on Memcached capacity.

Desired behavior/Acceptance criteria

  • WikiLambdaFunctionCall entries in Memcached expire after TTL_MINUTE
  • No functional regression in fragment serving

Completion checklist