An alert for the reference-need-predictor service was triggered on 16/07/2025. The initial alert information was:
reference-need-predictor revision-models istio-system k8s-mlserve critical eqiad prometheus
I looked at the reference-need-predictor logs in logstash: https://logstash.wikimedia.org/goto/c558820a21decb143949f8319d3e6fbb and there were spikes in the minute of 10:34 UTC. They showed that the service was failing due to:
concurrent.futures.process.BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending.
see logs here: https://phabricator.wikimedia.org/P79250
@elukey suggested looking at the Istio dashboard since all the traffic goes through Istio before reaching the isvcs. After drilling into to 10 to 11 UTC: https://logstash.wikimedia.org/goto/efdaccbb1f9112eeb474fdb85ef502f8, we saw that the user agent is: WME/2.0 (https://enterprise.wikimedia.com/; wme_mgmt@wikimedia.org)
This alert resolved itself at ~13 UTC but we are going to investigate whether the model inputs sent by the WME client were unusually large, whether they require more resources, or if it was a different issue.
