Description:
Rollouts in the revertrisk namespace stall on exceeded quota: the new Knative revision comes up before the old one goes away, and their combined requests don't fit. Pods sit Pending and the namespace is left mixed-version until the old revision is deleted by hand.
Steady-state usage is ~98 CPU / ~148Gi against a 140 CPU / 200Gi quota — not enough headroom for one service to double up during a rollout. Raised the revertrisk quota to 160 CPU / 230Gi in CL 1326234 (ml-serve-eqiad, ml-serve-codfw, ml-staging-codfw).
Scoped to revertrisk. Quotas in the other ml-serve namespaces and GPU rollout headroom (T420883) are not covered here.