Page MenuHomePhabricator

Re-IP Swift hosts to per-rack subnets in codfw rows A-D
Closed, ResolvedPublic

Description

As part of the move from per-row to per-rack redundancy model hosts in codfw rows A-D need to be configured / moved to new per-rack vlans/subnets. This work can be tackled once we have completed the physical move of all hosts in those rows from old 'asw' switch devices to new 'lsw' ones.

In discussion on irc we touched on some of the challenges for these hosts, which as I understand may use IP addresses as identifiers. We also need to consider how clusters function with hosts on different subnets that were previous layer-2 adjacent.

Having tested the migration process on ms-be2075, we know that it works thus:

  1. Drain node
  2. Remove node from rings
  3. Reimage node (with --move-vlan)
  4. Make sure the swift ring manager knows about the relevant per-rack subnet
  5. Add node back to rings

Newer nodes automatically get added to new-style subnets (ms-be2081 and later); so start with the newest node with old-style networking and move backwards, meaning that the oldest nodes get done last (and might have been aged out in the mean time).

  • ms-be2080
  • ms-be2079
  • ms-be2078
  • ms-be2077
  • ms-be2076
  • ms-be2075
  • ms-be2074
  • ms-be2073
  • ms-be2072
  • ms-be2071
  • ms-be2070

Below this point, nodes are using old-style storage, so we might want to fix that at the same time

  • ms-be2069
  • ms-be2068
  • ms-be2067
  • ms-be2066
  • ms-be2065
  • ms-be2064
  • ms-be2063
  • ms-be2062
  • ms-be2061 [this and below decommissioned]
  • ms-be2060
  • ms-be2059
  • ms-be2058
  • ms-be2057

There are three places where puppet has to be changed when moving a system to the new storage layout (with the sre.swift.convert-disks cookbook):

  1. hieradata/regex.yaml
  2. modules/install_server/files/autoinstall/scripts/partman_early_command.sh
  3. modules/profile/data/profile/installserver/preseed.yaml

Details

Related Changes in Gerrit:
SubjectAuthorRepoBranchLines +/-
MVernonoperations/puppetproduction+2 -0
MVernonoperations/puppetproduction+4 -4
MVernonoperations/puppetproduction+0 -5
MVernonoperations/puppetproduction+6 -2
MVernonoperations/puppetproduction+5 -5
MVernonoperations/puppetproduction+0 -4
MVernonoperations/puppetproduction+6 -2
MVernonoperations/puppetproduction+4 -8
MVernonoperations/puppetproduction+7 -2
MVernonoperations/puppetproduction+2 -2
MVernonoperations/puppetproduction+1 -1
MVernonoperations/puppetproduction+1 -1
MVernonoperations/puppetproduction+0 -6
MVernonoperations/puppetproduction+6 -3
MVernonoperations/puppetproduction+7 -0
MVernonoperations/puppetproduction+0 -6
MVernonoperations/puppetproduction+6 -3
MVernonoperations/puppetproduction+3 -0
MVernonoperations/puppetproduction+0 -6
MVernonoperations/puppetproduction+10 -3
MVernonoperations/puppetproduction+0 -6
MVernonoperations/puppetproduction+9 -3
MVernonoperations/puppetproduction+0 -6
MVernonoperations/puppetproduction+19 -3
MVernonoperations/puppetproduction+1 -0
MVernonoperations/puppetproduction+0 -2
MVernonoperations/puppetproduction+2 -1
MVernonoperations/puppetproduction+1 -0
MVernonoperations/puppetproduction+0 -2
Show related patches Customize query in gerrit
Related Changes in GitLab:
TitleReferenceAuthorSource BranchDest Branch
Add codfw racks B2, C4repos/data_persistence/swift-ring!16mvernoncodfw_racksmain
Add Codfw rack D2repos/data_persistence/swift-ring!15mvernoncodfw_d2main
Add codfw rack A4repos/data_persistence/swift-ring!11mvernoncodfw_a4main
Customize query in GitLab

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2072.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2073.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2072.codfw.wmnet with OS bullseye completed:

  • ms-be2072 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603111154_mvernon_1466038_ms-be2072.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2073.codfw.wmnet with OS bullseye completed:

  • ms-be2073 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603111205_mvernon_1469346_ms-be2073.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1250591 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: add 5 codfw backends

https://gerrit.wikimedia.org/r/1250591

Change #1250591 merged by MVernon:

[operations/puppet@production] swift: add 5 codfw backends

https://gerrit.wikimedia.org/r/1250591

Change #1264368 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: drain 3 codfw nodes for reimage

https://gerrit.wikimedia.org/r/1264368

Change #1264368 merged by MVernon:

[operations/puppet@production] swift: drain 3 codfw nodes for reimage

https://gerrit.wikimedia.org/r/1264368

Change #1270382 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] codfw: remove 3 drained ms be nodes for reimage

https://gerrit.wikimedia.org/r/1270382

Change #1270382 merged by MVernon:

[operations/puppet@production] codfw: remove 3 drained ms be nodes for reimage

https://gerrit.wikimedia.org/r/1270382

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2070.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2070.codfw.wmnet with OS bullseye completed:

  • ms-be2070 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604131414_mvernon_2220287_ms-be2070.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2069.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2069.codfw.wmnet with OS bullseye completed:

  • ms-be2069 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604131543_mvernon_2300741_ms-be2069.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye

Change #1270860 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] preseed: move ms-be206[8,9] to new-style storage

https://gerrit.wikimedia.org/r/1270860

Change #1270860 merged by MVernon:

[operations/puppet@production] preseed: move ms-be206[8,9] to new-style storage

https://gerrit.wikimedia.org/r/1270860

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye completed:

  • ms-be2068 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604140805_mvernon_3003200_ms-be2068.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1270899 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] partman: also add ms-be206[8-9] to partman_early_command

https://gerrit.wikimedia.org/r/1270899

Change #1270899 merged by MVernon:

[operations/puppet@production] partman: also add ms-be206[8-9] to partman_early_command

https://gerrit.wikimedia.org/r/1270899

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye executed with errors:

  • ms-be2068 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604141335_mvernon_3231627_ms-be2068.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console ms-be2068.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye executed with errors:

  • ms-be2068 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604141459_mvernon_3287667_ms-be2068.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console ms-be2068.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2068.codfw.wmnet with OS bullseye executed with errors:

  • ms-be2068 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604141615_mvernon_3337436_ms-be2068.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console ms-be2068.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Change #1271676 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] hiera: ms-be206[8-9] need new-style storage

https://gerrit.wikimedia.org/r/1271676

Change #1271676 merged by MVernon:

[operations/puppet@production] hiera: ms-be206[8-9] need new-style storage

https://gerrit.wikimedia.org/r/1271676

Change #1271726 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: restore 3 nodes to rings, drain 2 more for reimage

https://gerrit.wikimedia.org/r/1271726

Change #1271726 merged by MVernon:

[operations/puppet@production] swift: restore 3 nodes to rings, drain 2 more for reimage

https://gerrit.wikimedia.org/r/1271726

Change #1279372 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: remove 2 drained nodes from rings for reimage

https://gerrit.wikimedia.org/r/1279372

Change #1279372 merged by MVernon:

[operations/puppet@production] swift: remove 2 drained nodes from rings, set for new-style storage

https://gerrit.wikimedia.org/r/1279372

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2066.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2067.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2066.codfw.wmnet with OS bullseye completed:

  • ms-be2066 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604291644_mvernon_496259_ms-be2066.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2067.codfw.wmnet with OS bullseye completed:

  • ms-be2067 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604291708_mvernon_511799_ms-be2067.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1280083 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: restore 2 nodes to rings, drain 2 more for reimage

https://gerrit.wikimedia.org/r/1280083

Change #1280083 merged by MVernon:

[operations/puppet@production] swift: restore 2 nodes to rings, drain 2 more for reimage

https://gerrit.wikimedia.org/r/1280083

Mentioned in SAL (#wikimedia-operations) [2026-05-14T09:20:55Z] <Emperor> rebalance codfw swift rings T354872

Change #1287365 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: remove 2 drained nodes for reimage

https://gerrit.wikimedia.org/r/1287365

Change #1287365 merged by MVernon:

[operations/puppet@production] swift: remove 2 drained nodes for reimage

https://gerrit.wikimedia.org/r/1287365

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2064.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2064.codfw.wmnet with OS bullseye executed with errors:

  • ms-be2064 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console ms-be2064.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Change #1287826 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: set ms-be206[4,5] to be new-style storage

https://gerrit.wikimedia.org/r/1287826

Change #1287826 merged by MVernon:

[operations/puppet@production] swift: set ms-be206[4,5] to be new-style storage

https://gerrit.wikimedia.org/r/1287826

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2064.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2065.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2064.codfw.wmnet with OS bullseye completed:

  • ms-be2064 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202605151024_mvernon_2916451_ms-be2064.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202605151034_mvernon_2916451_ms-be2064.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2065.codfw.wmnet with OS bullseye completed:

  • ms-be2065 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202605151104_mvernon_2939402_ms-be2065.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1288441 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: restore 2 nodes to rings, drain final 2 for reimage

https://gerrit.wikimedia.org/r/1288441

Change #1288441 merged by MVernon:

[operations/puppet@production] swift: restore 2 nodes to rings, drain final 2 for reimage

https://gerrit.wikimedia.org/r/1288441

Change #1298665 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: remove 2 drained nodes for reimage

https://gerrit.wikimedia.org/r/1298665

Change #1298666 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: move ms-be206[2,3] to new-style storage

https://gerrit.wikimedia.org/r/1298666

Change #1298665 merged by MVernon:

[operations/puppet@production] swift: remove 2 drained nodes for reimage

https://gerrit.wikimedia.org/r/1298665

Change #1298666 merged by MVernon:

[operations/puppet@production] swift: move ms-be206[2,3] to new-style storage

https://gerrit.wikimedia.org/r/1298666

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2062.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2002 for host ms-be2063.codfw.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2062.codfw.wmnet with OS bullseye completed:

  • ms-be2062 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606081213_mvernon_1093975_ms-be2062.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2002 for host ms-be2063.codfw.wmnet with OS bullseye completed:

  • ms-be2063 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606081219_mvernon_1095115_ms-be2063.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1298773 had a related patch set uploaded (by MVernon; author: MVernon):

[operations/puppet@production] swift: restore 2 nodes to rings

https://gerrit.wikimedia.org/r/1298773

Change #1298773 merged by MVernon:

[operations/puppet@production] swift: restore 2 nodes to rings

https://gerrit.wikimedia.org/r/1298773

MatthewVernon claimed this task.
MatthewVernon updated the task description. (Show Details)

All done! And all codfw backends moved to new-style storage.