Now that {T364095} is effectively completed we can begin the process of moving existing hosts from the old ASW switches in codfw rows C & D to the new Leaf switches there.
**Hosts****Schedule**
Hosts on the following vlans need to be moved (or see [[ https://phabricator.wikimedia.org/P66915 | full list ]]):The full schedule of planned maintenance windows is as listed below. On most days we will try to move the servers in two full racks, which seems a reasonable compromise between not disrupting too many hosts on a given day, and completing the moves in good time. It should allow us to get the work done before the datacentre switchover to codfw on September 25th.
|Vlan|Hosts||Racks|Date|Task|
|-----|----|------|-----|
|private1-c-codfw|P66913||C1|Wed Sept 4th 16:00 UTC|T373095|
|public1-c-codfw|P66911||C2 & C3|Thu Sept 5th 16:30 UTC|T373096|
|private1-d-codfw|P66914||C4 & C5|Tue Sep 10th 16:00 UTC|T373097|
|public1-d-codfw|P66912|
**Order of moves**C6 & C7|Wed Sep 11th 16:00 UTC|T373101|
|D1 & D2|Thu Sep 12th 16:00 UTC|T373102|
|D3 & D4|Tue Sep 17th 16:00 UTC|T373103|
|D5 & D6|Wed Sep 18th 16:00 UTC|T373104|
|D7 & D8|Thu Sep 19th 16:00 UTC|T373105|
We probably need to approach this in several phases to ensure it goes smoothly and minimise the disruption to other teams and live services during the work. One lesson from T355544,Server links will be moved one-by-one from old to the new switch. when we moved the servers from the old switches to new in rows A & B, is that for certain clusters of hosts it is a lot more difficult to schedule than others.So **no two hosts will be offline at once.**
In brief, for some clusters of hosts providing a particular serviceBased on previous experience **each host is likely to only lose comms for ~10 seconds**. It is inevitable that a small number of the new cables do not work, they are all basically "equal"however, and depooling any set of N hosts requires the same effort and analysis as any other random set of N hosts. In those cases moving the hosts on a rack-by-rack basis works just fine,or there is some minor glitch in the move. So it is possible in an edge case that a host will be offline for 2-3 minutes. it doesn't make a difference to the service running on them what order things are depooledOn previous occasions this happened with about 1 out of 20 hosts.
Other services, however, are not like that. For instance if specific hosts in the cluster are performing certain duties (master nodes etc), then moving the hosts in an effectively random order (rack-by-rack) can be difficult to orchestrate for the service owners. The database hosts maintained by our Data Platform were particularly tricky to get moved the last time.**Planning**
TWe are using the best way to proclow gsheet to plan the moves and detail actions that need would therefore seem to be:to be taken to ensure things go smoothly:
# Move LVS hosts as they need to connect to the new per-rack vlans before we start
# Move hosts for services that may require special attention on a host-by-host basis (i.e. db's)https://docs.google.com/spreadsheets/d/16xoZuDeC_-o6s70uEMnvdgn4BlT1f8__WPYprRuduIA
SREs can help by:
* Listing the actions to be taken for a given type of host - if they can all be treated the same - on the first tab
# Move remaining hosts on a rack-by-rack basis* Listing the actions to be taken for a particular host - if we need to detail the plans for specific hosts - on the tab for each day
**Planning**The [[ https://docs.google.com/spreadsheets/d/16xoZuDeC_-o6s70uEMnvdgn4BlT1f8__WPYprRuduIA#gid=610489305 | Schedule by Team & Host Type ]] tab lists all hosts, grouped by team and host type, with hosts of the same type to be moved on the same day highlighted.
The [[ https://docs.google.com/spreadsheets/d/16xoZuDeC_-o6s70uEMnvdgn4BlT1f8__WPYprRuduIA#gid=397380523 | Schedule by Team & Date ]] tab lists all the hosts, grouped by team and date, with all hosts to be moved on a given date highlighted.
Hopefully this will assist in planning for the moves.
**Exceptions**
We are planning the moves on a rack-by-rack schedule as this is the easiest way to group them and perform the actual work. However hosts can theoretically be moved in any order. So if there is a particular blocker for a given maintenance, for instance if two hosts are scheduled for the same day but they cannot both be depooled at the same time, we can try to move one or other ahead of the overall rack move, so they aren't both affected on the same day.
First step is to plan the moves on a rack-by-rack basis. When the tasks and schedule for that is clear we can discuss with service owners and see if we need to address any specific hosts ahead of the schedule, for instance if they need to be downtimed and it is not possible to downtime two particular hosts at the one time we can move one prior to the scheduled dateAny such exceptions can be noted here or please discuss with topranks on irc.