Some users reported on Slack that they can't reach images on upload.w.o, located in US.
NEL reports spikes in tcp.timed_out
- Depooled ulsfo with sre.dns.admin cookbook to investigate about the TCP timeouts (https://logstash.wikimedia.org/goto/acb7d9be4e06cfed7c9b6b92c58bb002)
- Recent ULSFO DC reimage (and cache hosts IP changes) affected liberica load balancer on a lvs host (lvs4008). This host was still using the old IPs and hence failed to repool some cache hosts (mainly upload hosts).
- This resulted in unavailability and degradation of services like users ability to visualize images on wikipedia.
- Affected users (ones that were geographically close to the ulsfo DC) weren’t able to visualize content from upload.w.o (mainly images)
- A restart of liberica services on lvs4008 correctly updated the loadbalancer configuration and pooled back the impacted cache hosts.
- Other lvs hosts in ulsfo are not affected by this because were reimaged recently and already pick up the new IPs.
- Command used: sudo cookbook sre.loadbalancer.upgrade --query 'P{lvs4009*}' --reason "config reload" restart
- New page at 15:59 due to excessive usage of ulsfo <-> codfw link. This is probably due to cold cache in ulsfo for reimaging combined with “partial repool” due to liberica not actually pooling cache hosts. Decided to move CA traffic from ulsfo to codfw to lower the pressure on ulsfo while the cache warms up: https://gerrit.wikimedia.org/r/c/operations/dns/+/1284699
- Rebalanced traffic between codfw and ulsfo, an error for user holding the lock on Homer appeared but promptly resolved by XioNoX
{F80145426}
