Page MenuHomePhabricator

Investigate DispatchChanges Normal job backlog time (mean avg, 15min) alert post datacenter switch
Open, Needs TriagePublic

Description

The alert DispatchChanges Normal job backlog time (mean avg, 15min) became Firing at 14:40 UTC, but Resolved at 15:00 UTC (according to my email inbox). However, on Grafana it was still Alerting until I changed the data source (following a recommendation by @fgiunchedi). It’s currently unclear why it seemed to disappear from alerts.w.o when grafana.w.o still had it.

Event Timeline

The two emails I received (I assume they’re fine to share):


Quoting the first one for convenience:

[1] Firing
Labels
alertname = DispatchChanges Normal job backlog time (mean avg, 15min) alert
__alert_rule_uid__ = MF0FSjJ4z
__contacts__ = "AlertManager","cxserver"
grafana_folder = Wikidata
rule_uid = MF0FSjJ4z
severity = critical
team = wikidata
Annotations
_alertId__ = 309
__dashboardUid__ = TUJ0V-0Zk
__orgId__ = 1
__panelId__ = 28
__value_string__ = [ var='B0' metric='NoData' labels={} value=null ]
grafana_state_reason = NoData
message = DispatchChanges job backlog is over 10 minutes! Normal values are between 0.5s and 1s

Source

However, on Grafana it was still Alerting until I changed the data source (following a recommendation by @fgiunchedi).

FTR, I just applied the same fix to the corresponding panel on the Wikidata Alerts dashboard.