Today ATSBackendErrorsHigh alert pages on an absolute number of errors (3/s) although I think the severity/SLO of the alert depends on the traffic levels of the backend service itself. In other words 3 errors/s on a service doing 3k requests/s is far different in terms of impact on users and SRE oncall than on a service doing, say, 10x or 100x less traffic.
Thus I'd like to propose switching ATSBackendErrorsHigh to page based on service availability i.e. failures / all requests and decide said number. Please also note that we can consider a blend of services in this case, namely keep some services on fix thresholds and some others based on availability. Having a single paging policy however will likely be easier to understand and troubleshoot.



