Many hours, at the top of the hour, we're seeing load spikes of reads to all s8 replicas.
Every once in a while, this causes some impact, up to and including user-visible errors and paging of SREs.
The first time was 2020-10-02 22:06:48 where it caused restbase issues, and made the LVS healthchecks for wikifeeds and termbox fail:
2020-10-02 22:06:48 <+icinga-wm> PROBLEM - LVS wikifeeds codfw port 4101/tcp - A node webservice supporting featured wiki content feeds. termbox.svc.eqiad.wmnet IPv4 #page on wikifeeds.svc.codfw.wmnet is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/LVS%23Diagnosing_problems
The second time was today:
2020-10-06 22:03:11 <+icinga-wm> PROBLEM - Not enough idle PHP-FPM workers for Mediawiki api_appserver at codfw #page on alert1001 is CRITICAL: 0.2271 lt 0.3 https://bit.ly/wmf-fpmsat https://grafana.wikimedia.org/d/RIA1lzDZk/application-servers-red-dashboard?panelId=54&fullscreen&orgId=1&from=now-3h&to=now&var-datasource=codfw+prometheus/ops&var-cluster=api_appserver
The issue shows most clearly as read_key traffic, although read_next also shows the pattern as well:
grafana/explore link, be logged in on grafana first: https://w.wiki/fTq
Looking at the past week, whatever the load spike is seems especially heavy at 22:00UTC.
What is this workload? Is it possible to get some query logs? And then to spread it out?


