Page MenuHomePhabricator

Rerack mc-gp2006 and mc2046
Closed, ResolvedPublic

Description

Those servers currently co-exist in the same racks with their stand-by servers, which is not great.

Please re-rack the following machines:

#1 mc-gp2006 : from D4 → D3

#2 mc2046: from B7 → B8

A heads up in -sre-private would be nice. You may do this work any time as long as those those two servers are not offline simultaneously.

Event Timeline

I can re-rack mc-gp2006 in U-26 of D3 on Wednesday.

rack B5 is not a good location to rack mc2046 because of 10G port availability on that switch. B8 or B3 would work better.
I can do that on Thursday just to be extra sure down time doesn't overlap.

@jijiki i got the server moved. the mgmt pings.

lmk if you have a preference for mc2046

FYI, there was this outstanding diff:

Change for lsw1-d4-codfw.mgmt.codfw.wmnet:

[edit interfaces xe-0/0/43]
-   description mc-gp2006;

It's because the cable got removed, but the old interface was still enabled.

I disabled the interface (https://netbox.wikimedia.org/extras/changelog/288031/) then ran Homer on the switch.

Change #1314701 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] mcrouter_wancache: temporary remove mc-gp2006

https://gerrit.wikimedia.org/r/1314701

jijiki changed the task status from Open to Stalled.EditedThu, Jul 23, 8:27 AM

I can re-rack mc-gp2006 in U-26 of D3 on Wednesday.

rack B5 is not a good location to rack mc2046 because of 10G port availability on that switch. B8 or B3 would work better.
I can do that on Thursday just to be extra sure down time doesn't overlap.

Thank you Jenn! We need to re-IP mc-gp2006 as it turns out, which means that moving mc2046 can happen after that. Do you have time this week help me with this?

B8 for mc2046 will be great, I have updated the description

Change #1314701 merged by Effie Mouzeli:

[operations/puppet@production] mcrouter_wancache: temporary remove mc-gp2006

https://gerrit.wikimedia.org/r/1314701

jijiki changed the task status from Stalled to In Progress.Thu, Jul 23, 9:58 AM
jijiki triaged this task as Medium priority.

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm executed with errors:

  • mc-gp2006 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console mc-gp2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm executed with errors:

  • mc-gp2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console mc-gp2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc-gp2006.codfw.wmnet with OS bookworm completed:

  • mc-gp2006 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607241513_jiji_2383655_mc-gp2006.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

@Jhancock.wm mc-gp2006 is alive an well (tx @cmooney for sorting out the networking bits). Is it alright if we re-rack mc2046 sometime next week?

Change #1315946 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] mcrouter_wancache: re-add mc-gp2006 to the gutter-pool

https://gerrit.wikimedia.org/r/1315946

Change #1315946 merged by Effie Mouzeli:

[operations/puppet@production] mcrouter_wancache: re-add mc-gp2006 to the gutter-pool

https://gerrit.wikimedia.org/r/1315946

@jijiki i can get this tomorrow some time. I will ping when i can start it.

@jijiki i can get this tomorrow some time. I will ping when i can start it.

Grand, thanks!

@jijiki it's moved. but unless you want to reimage it, you'll need to redo whatever magic cathal did.

Networking is no magic :)

The server is still connected to the B7 rack though, what is its new switch port ?

For the re-image, the "hack" is to set the host's switch port vlan to one of the LEGACY_VLANS https://github.com/wikimedia/operations-cookbooks/blob/master/cookbooks/sre/hosts/__init__.py#L38 (like it's currently configured, but on the wrong old B7 switch port)
Then run the re-image cookbook with --move-vlan

it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro was on it.
@jijiki lmk which distro or you can take over.

it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro was on it.
@jijiki lmk which distro or you can take over.

I made the change now. @Jhancock.wm if you need to do this again the trick Arzhel is suggesting is when you connect the new switch port to the server in Netbox, set the "untagged vlan" on the new port to one of the row-wide ones. In this case I used "private1-b-codfw".

When the reimage cookbook is run with --move-vlan, and it sees the port is on one of those, it will change the port vlan to a rack-specific one, say private1-b8-codfw. But it will also assign new IPs for the host and fix up the dns (the bit I did manually last time). So it can save a few steps when we do server moves like this.

No hassle we are always around if needed, but that's the quickest way to do it.

it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro was on it.
@jijiki lmk which distro or you can take over.

bookworm please, but tell us if you don't have the bandwidth and we'll set it up.

i got it done. ty for given me the chance to learn a new thing!

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie completed:

  • mc2046 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608051157_jiji_871101_mc2046.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie

Change #1321564 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] mcrouter_wancache: replace mc2046's IP

https://gerrit.wikimedia.org/r/1321564

Change #1321564 merged by Effie Mouzeli:

[operations/puppet@production] mcrouter_wancache: replace mc2046's IP

https://gerrit.wikimedia.org/r/1321564

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie completed:

  • mc2046 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608051438_jiji_989551_mc2046.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
jijiki claimed this task.

Thank you @Jhancock.wm and @cmooney for doing all the magic stuff, server is serving :)