Disclaimer: This info is pieced together from the scrollback in Wikimedia-Search . If you have better info, feel free to correct this.
Beginning at 0752 UTC today, all wdqs graph split hosts in EQIAD alerted:
[07:52:25] <jinxer-wm> FIRING: SystemdUnitFailed: systemd-timedated.service on wdqs1023:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed
Per earlier conversation with @dcausse , the working theory is that load-categories-daily.timer ran before we could reload the wdqs-categories database. When this happens, the daily job tries to write to a namespace that doesn't exist , which fails and generates enough log spam to lock up the host. The log spam floods the console as well, making it unresponsive.
The actual impact was minimal:
- this service is quite new and does not serve a lot of traffic
- it was apparently still serving queries via query-main.wikidata.org, so no actual user impact