* Toolforge tools were not responding to http requests (Tools-proxy-9 was returning an error page)
* We found that Ceph had intermittent issues since last night, after some hosts were upgraded to Bookworm.
* Bookworm hosts were: cloudcephosd100[6-8], cloudcephosd1003[5-7].
* This caused intermittent issues to both Toolforge and Cloud VPS
* We are downgrading the hosts that were upgraded to Bookworm, which is in itself proving challenging
Incident doc: https://docs.google.com/document/d/1CLY_iZyXDTyJEl4fKYeU1aRSNsheO9-TZcjyW9wFyEk/edit?tab=t.0#heading=h.nz4dlhpgbsjm
### Timeline (UTC)
(see incident doc above for updated timeline)
01:59 PROBLEM - SSH on cloudcephosd1035 is CRITICAL
02:08 RECOVERY - SSH on cloudcephosd1035 is OK
04:20 SSH on cloudcephosd1036 is CRITICAL
05:02 RECOVERY - SSH on cloudcephosd1036 is OK
05:28 PROBLEM - SSH on cloudcephosd1037 is CRITICAL
05:31 RECOVERY - SSH on cloudcephosd1037 is OK
07:13 many cloud-vps hosts reported down
07:14 PROBLEM - SSH on cloudcephosd1037 is CRITICAL
07:30 RECOVERY - SSH on cloudcephosd1037 is OK
07:16 FIRING: CephSlowOps: Ceph cluster in eqiad has 1451 slow ops
07:19 FIRING: WidespreadInstanceDown: Widespread instances down in project cloudinfra
07:20 cloud-vps back to normal (no hosts reported down)
07:21 RESOLVED: CephSlowOps: Ceph cluster in eqiad has 779 slow ops
07:24 RESOLVED: WidespreadInstanceDown: Widespread instances down in project cloudinfra
07:27 FIRING: CephSlowOps: Ceph cluster in eqiad has 908 slow ops
07:32 RESOLVED: CephSlowOps: Ceph cluster in eqiad has 908 slow ops
08:08 many cloud-vps hosts again reported down
08:10 FIRING: CephSlowOps: Ceph cluster in eqiad has 1678 slow ops
08:13 PROBLEM - SSH on cloudcephosd1036 is CRITICAL
08:18 wmcs-dnsleaks fails on cloudcontrol1007 (possibly unrelated)
08:18 FIRING: WidespreadInstanceDown: Widespread instances down in project cloudinfra
08:19 cloud-vps back to normal (no hosts reported down)
08:20 Manuel reports switchmaster.toolforge.org is down
08:23 RESOLVED: WidespreadInstanceDown: Widespread instances down in project cloudinfra
08:23 FIRING: [2x] ProbeDown: Service tools-k8s-haproxy-5:30000 has failed probes (http_admin_toolforge_org_ip4)
08:27 lucas.werkmeister@wikimedia.de reports all tools are returning an error from tools-proxy-9
08:39 <lucaswerkmeister> I can SSH into tools-proxy-9, the only failed systemd unit is logrotate which judging by the journal has been broken for a long time, probably not related
08:44 <lucaswerkmeister> I think tools-proxy-9 times out trying to reach k8s.tools.eqiad1.wikimedia.cloud in turn
08:44 <lucaswerkmeister> I can SSH into that one too, no high load there either
08:52 Incident opened. Francesco Negri becomes IC.
08:55 Toolforge is working again. No action was taken.
08:58 RESOLVED: [2x] ProbeDown: Service tools-k8s-haproxy-5:30000 has failed probes (http_admin_toolforge_org_ip4)
09:01 Incident is resolved.
09:25 Francesco Negri starts wmcs.toolforge.k8s.reboot for tools-k8s-worker-nfs-77, tools-k8s-worker-nfs-68, tools-k8s-worker-nfs-37, as they were alerting with “many processes in D state”
09:28 SSH on cloudcephosd1036 is OK
09:39 PROBLEM - SSH on cloudcephosd1035 is CRITICAL
09:42 RECOVERY - SSH on cloudcephosd1035 is OK
10:20 PROBLEM - SSH on cloudcephosd1036 is CRITICAL
10:21 FIRING: CephSlowOps: Ceph cluster in eqiad has 847 slow ops
10:24 RECOVERY - SSH on cloudcephosd1036 is OK
10:26 RESOLVED: CephSlowOps: Ceph cluster in eqiad has 1386 slow ops
10:28 FIRING: CephSlowOps: Ceph cluster in eqiad has 5134 slow ops
10:33 RESOLVED: CephSlowOps: Ceph cluster in eqiad has 1272 slow ops
11:35 SSH on cloudcephosd1008 is CRITICAL
11:41 SSH on cloudcephosd1008 is OK
11:41 FIRING: WidespreadInstanceDown
11:46 RESOLVED: WidespreadInstanceDown
12:13 SSH on cloudcephosd1035 is CRITICAL
12:18 FIRING: WidespreadInstanceDown
12:20 SSH on cloudcephosd1035 is OK
12:23 RESOLVED: WidespreadInstanceDown
14:12 FIRING: WidespreadInstanceDown
14:24 Reopening the incident
14:29 <dhinus> things seem to get worse after 14:12 UTC
14:30 <andrewbogott> 1007 is frozen right now. So we /do/ have two down at once, which could maybe explain current bad behavior.
14:41 <dhinus> we have now 9 OSDs down (compared to 16 before)
15:30 Most (all?) VMs have recovered. Ceph health still shows flapping of various OSDs.
16:05 Cloudcephosd1037 is now back to running bullseye, and is repooled. Cloudcephosd1036 is in the process of reimaging.