Page MenuHomePhabricator

Migrate servers in codfw rack C1 from asw-c1-codfw to lsw1-c1-codfw
Closed, ResolvedPublic

Description

Currently scheduled for Wed Sept 4th 2024 16:00 UTC

As part of the scheduled refresh of switch equipment in codfw rows C and D we need to move the network connections for servers in rack C1 from the old to new switch.

Hosts in this rack are managed by the following teams:

Core Platform
Data Persistence
Infrastructure Foundations
Search Platform
ServiceOps

A full list of the specific hosts can be found below. We will use the sheet to plan the moves and co-ordinate with other SRE teams on actions required to ensure things go smoothly:

https://docs.google.com/spreadsheets/d/16xoZuDeC_-o6s70uEMnvdgn4BlT1f8__WPYprRuduIA#gid=856344840

Server links will be moved one-by-one from old to the new switch. So no two hosts will be offline at once.

Based on previous experience each host is likely to only lose comms for ~10 seconds. It is inevitable that a small number of the new cables do not work, however, or there is some minor glitch in the move. So it is possible in an edge case that a host will be offline for 2-3 minutes. On previous occasions this happened with about 1 out of 20 hosts.

Event Timeline

cmooney triaged this task as Medium priority.
Restricted Application added a subscriber: Aklapper. · View Herald Transcript
ABran-WMF subscribed.
db2125s2
db2138s2
db2149s3
db2190s3
db2206s4
db2207s2 candidate master
es2031es2 standalone
es2032es1 standalone
es2036es6
pc2013pc3

those hosts will be depooled before operations

Noting that es2032 is a "perceived master" (since dbctl requires a master) so you can't just depool it. You need to switch it over but since it's RO, it's just the dbctl command (and DNS mayyyybe but I'm not sure).

Noting that es2032 is a "perceived master" (since dbctl requires a master) so you can't just depool it. You need to switch it over but since it's RO, it's just the dbctl command (and DNS mayyyybe but I'm not sure).

my intent was to dbctl set-master xxxx and depool the host; and I'll double check the DNS indeed great catch!

Mentioned in SAL (#wikimedia-operations) [2024-09-04T07:09:14Z] <akosiaris> T373095 depool kubernetes2011, kubernetes2012, kubernetes2036, kubernetes2037, wikikube-worker2037, wikikube-worker2038, mw2436, mw2437

All wikikube hosts have been depooled. RESTBase and mc-wf should be good to do at anytime per comments in the sheet.

All wikikube hosts have been depooled. RESTBase and mc-wf should be good to do at anytime per comments in the sheet.

Thanks!

Mentioned in SAL (#wikimedia-operations) [2024-09-04T11:07:00Z] <topranks> migrating VMs off ganeti2009 in advance of network maintence codfw rack C1 - T373095

[...] I'll double check the DNS indeed great catch!

no DNS modification needed, I'll run sudo dbctl --scope codfw section es1 set-master es2030

$ rg 'es1-master'
templates/wmnet
57:es1-master      5M  IN CNAME    es1027.eqiad.wmnet.

Mentioned in SAL (#wikimedia-operations) [2024-09-04T14:39:30Z] <arnaudb@cumin1002> dbctl commit (dc=all): 'swap masters for es1 - T373095', diff saved to https://phabricator.wikimedia.org/P68648 and previous config saved to /var/cache/conftool/dbconfig/20240904-143928-arnaudb.json

Mentioned in SAL (#wikimedia-operations) [2024-09-04T15:43:23Z] <topranks> configure lsw1-c1-codfw interfaces for servers in advance of move T373095

Icinga downtime and Alertmanager silence (ID=6ba6c00e-f364-45da-8be3-ee80785b36c0) set by cmooney@cumin1002 for 0:30:00 on 27 host(s) and their services with reason: Move server uplinks codfw racks C1

cumin2002.codfw.wmnet,db[2125,2138,2149,2190,2206-2207].codfw.wmnet,es[2031-2032,2036].codfw.wmnet,ganeti[2009-2010].codfw.wmnet,kubernetes[2011-2012,2036-2037].codfw.wmnet,mc-wf2001.codfw.wmnet,mw[2436-2437].codfw.wmnet,pc2013.codfw.wmnet,restbase[2022,2031-2032].codfw.wmnet,sessionstore2005.codfw.wmnet,wcqs2002.codfw.wmnet,wikikube-worker[2037-2038].codfw.wmnet

Mentioned in SAL (#wikimedia-operations) [2024-09-04T16:06:50Z] <topranks> migrating servers in codfw rack C1 from asw-c-codfw to lsw1-c1-codfw T373095

Link moves completed, all servers now responding to ping again so looks ok. Unsure of exact times for each but looking at cumin2002 it was less than 5 seconds :)

[2024-09-04T16:08:00.964101] 2: eno1: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc mq state DOWN group default
[2024-09-04T16:08:05.065653] 2: eno1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default

Mentioned in SAL (#wikimedia-operations) [2024-09-04T16:29:29Z] <claime> T373095 repool kubernetes2011, kubernetes2012, kubernetes2036, kubernetes2037, wikikube-worker2037, wikikube-worker2038, mw2436, mw2437

cmooney claimed this task.