Page MenuHomePhabricator

Migrate stat hosts to Bookworm or later
Closed, ResolvedPublic

Description

Per this Slack conversation with @BTullis , we need to get these hosts on Bookworm or later before the end of September.

These hosts need a bit of special handling:

  • Reimages should take place in a maintenance window to minimize user impact.
  • We should use the reuse-parts partman recipes so we can preserve user data.

Creating this ticket to:

  • Update partman recipes
  • Schedule maintenance windows/contact stat host users
  • Reimage the hosts
  • Verify operation
  • Update docs/procedures as needed

Event Timeline

bking changed the task status from Open to In Progress.Tue, Aug 11, 9:14 PM
bking claimed this task.
bking triaged this task as Medium priority.

Each stat host has a different disk configuration, I've posted the output of lsblk on each host here .

The next step is to write (or find/use) a partman recipe for each host that preserves the contents of /srv.

stat10[09-10].eqiad.wmnet: 2 drives, probably modules/install_server/files/autoinstall/partman/custom/reuse-analytics-raid1-2dev.cfg
stat1011.eqiad.wmnet: 4 drives, probably modules/install_server/files/autoinstall/partman/custom/reuse-analytics-stat-4dev.cfg.

I'll get a patch up shortly.

Change #1325921 had a related patch set uploaded (by Bking; author: Bking):

[operations/puppet@production] stat hosts: update partitioning to retain /srv data

https://gerrit.wikimedia.org/r/1325921

Change #1325921 merged by Bking:

[operations/puppet@production] stat hosts: update partitioning to retain /srv data

https://gerrit.wikimedia.org/r/1325921

The proposed schedule is as follows:

stat1008 - Wednesday 19th @ 10:00 UTC
stat1009 - Wednesday 19th @ 16:00 UTC
stat1010 - Thursday 20th @ 10:00 UTC
stat1011 - Thursday 20th @ 16:00 UTC

BTullis updated the task description. (Show Details)

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1008.eqiad.wmnet with OS bookworm

Change #1327088 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] install_server: Install grub on all RAID members in reuse recipes

https://gerrit.wikimedia.org/r/1327088

Change #1327088 merged by Btullis:

[operations/puppet@production] install_server: Install grub on all RAID members in reuse recipes

https://gerrit.wikimedia.org/r/1327088

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1008.eqiad.wmnet with OS bookworm executed with errors:

  • stat1008 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console stat1008.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1008.eqiad.wmnet with OS bookworm

First puppet run on stat1008 has failed.

image.png (1,154×451 px, 63 KB)

Working on a puppet patch now.

Change #1327105 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Install amd_rocm version 6.1 on stat1008 after reimage to bookworm

https://gerrit.wikimedia.org/r/1327105

Change #1327105 merged by Btullis:

[operations/puppet@production] Install amd_rocm version 6.1 on stat1008 after reimage to bookworm

https://gerrit.wikimedia.org/r/1327105

I fixed the amd_rocm version number, but during the first puppet run I saw this on the console of stat1008.

Debian GNU/Linux 12 stat1008 ttyS1

stat1008 login: root
Password: 
[ 3589.672355] ---[ end Kernel panic - not syncing: GHES: Fatal hardware error ]---

Investigating.

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1008.eqiad.wmnet with OS bookworm completed:

  • stat1008 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202608191148_btullis_1371291_stat1008.out, asking the operator what to do
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202608191207_btullis_1371291_stat1008.out, asking the operator what to do
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202608191233_btullis_1371291_stat1008.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608191306_btullis_1371291_stat1008.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm

Change #1327505 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Install amd_rocm version 6.1 on stat1010 after reimage to bookworm

https://gerrit.wikimedia.org/r/1327505

Change #1327505 merged by Btullis:

[operations/puppet@production] Install amd_rocm version 6.1 on stat1010 after reimage to bookworm

https://gerrit.wikimedia.org/r/1327505

Change #1327516 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Add a new partman reuse recipe for hardware RAID two drive hosts

https://gerrit.wikimedia.org/r/1327516

Change #1327516 merged by Btullis:

[operations/puppet@production] Add a new partman reuse recipe for hardware RAID two drive hosts

https://gerrit.wikimedia.org/r/1327516

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm executed with errors:

  • stat1010 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console stat1010.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm

Change #1327520 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] install_server: Generate the H750 stat host reuse recipe at install time

https://gerrit.wikimedia.org/r/1327520

Change #1327520 merged by Btullis:

[operations/puppet@production] install_server: Generate the H750 stat host reuse recipe at install time

https://gerrit.wikimedia.org/r/1327520

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm executed with errors:

  • stat1010 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console stat1010.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host stat1009.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1010.eqiad.wmnet with OS bookworm completed:

  • stat1010 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608201224_btullis_1894602_stat1010.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host stat1009.eqiad.wmnet with OS bookworm completed:

  • stat1009 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608201232_btullis_1894737_stat1009.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host stat1011.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host stat1011.eqiad.wmnet with OS bookworm completed:

  • stat1011 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608201643_bking_3391430_stat1011.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
bking updated the task description. (Show Details)

stat1011 successfully reimaged, which means all stat hosts have been successfully migrated to Bookworm. Closing...