Over the last ~3 days fourohfour has been failing its probes consistently, the first being at 2026-03-01T08:59:51:
metricsinfra-alertmanager-2:~# journalctl -u alertmanager-webhook-logger.service --since -1w --grep http_this_tool_does_not_exist_toolforge_org | wc -l 1183
Scaling up the deployment didn't help, eventually the tool fails its healthcheck and then gets restarted by k8s in a loop, eventually the pods end up in crashloop
