Page MenuHomePhabricator

Migrate Mediawiki memcached to Debian Trixie
Open, MediumPublic

Description

Update all memcached servers across both datacentres to Debian Trixie.

How?
We will reimage on of the hosts we are retiring from eqiad (mc1054) to trixie, add it to the pool for a while (2 days, maybe more), and if there is nothing odd, we can proceed with the upgrade.

eqiad

  • mc1055.eqiad.wmnet
  • mc1056.eqiad.wmnet
  • mc1057.eqiad.wmnet
  • mc1058.eqiad.wmnet
  • mc1059.eqiad.wmnet
  • mc1060.eqiad.wmnet
  • mc1061.eqiad.wmnet
  • mc1062.eqiad.wmnet
  • mc1063.eqiad.wmnet
  • mc1064.eqiad.wmnet
  • mc1065.eqiad.wmnet
  • mc1066.eqiad.wmnet
  • mc1067.eqiad.wmnet
  • mc1068.eqiad.wmnet
  • mc1069.eqiad.wmnet
  • mc1070.eqiad.wmnet
  • mc1071.eqiad.wmnet
  • mc1072.eqiad.wmnet
  • mc-wf1001.eqiad.wmnet
  • mc-wf1002.eqiad.wmnet

codfw

  • mc2038.codfw.wmnet
  • mc2039.codfw.wmnet
  • mc2040.codfw.wmnet
  • mc2041.codfw.wmnet
  • mc2042.codfw.wmnet
  • mc2043.codfw.wmnet
  • mc2044.codfw.wmnet
  • mc2045.codfw.wmnet
  • mc2046.codfw.wmnet
  • mc2047.codfw.wmnet
  • mc2048.codfw.wmnet
  • mc2049.codfw.wmnet
  • mc2050.codfw.wmnet
  • mc2051.codfw.wmnet
  • mc2052.codfw.wmnet
  • mc2053.codfw.wmnet
  • mc2054.codfw.wmnet
  • mc2055.codfw.wmnet
  • mc-gp2004.codfw.wmnet
  • mc-gp2005.codfw.wmnet
  • mc-gp2006.codfw.wmnet
  • mc-wf2001.codfw.wmnet
  • mc-wf2002.codfw.wmnet

Notes

  • codfw can wait for Q2 when we will also refresh the memcached cluster
  • @RLazarus we have two mc-wf servers per DC, what should we do in order to not disrupt the application using them?
  • For mc-gp servers needing to be re-IPed, the process will be to
    1. Degrade the gutter pool to 1 server (ie remove the two hosts from profile::mediawiki::mcrouter_wancache::shards
    2. run puppet on deployment servers and re-deploy mw-mcrouter in both DCs
    3. Reimage/re-IP (--move-vlan)
    4. rinse

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes
Nemoralis renamed this task from MIgrate Mediawiki memcached to Debian Trixie to Migrate Mediawiki memcached to Debian Trixie.May 12 2026, 12:12 PM
jijiki updated the task description. (Show Details)
JMeybohm triaged this task as Medium priority.May 21 2026, 9:12 AM
JMeybohm moved this task from Inbox to Scheduled (this Q) on the ServiceOps new board.

Change #1293656 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] site.pp: re-add mc1054 for trixie testing

https://gerrit.wikimedia.org/r/1293656

Change #1293656 merged by Blake:

[operations/puppet@production] site.pp: re-add mc1054 for trixie testing

https://gerrit.wikimedia.org/r/1293656

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1054.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1054.eqiad.wmnet with OS trixie completed:

  • mc1054 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202605261414_blake_759281_mc1054.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1294216 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] mcrouter_wancache: swap mc1055 for mc1054 for trixie testing

https://gerrit.wikimedia.org/r/1294216

i've verified that mc1054 is running the versions of memkeys and prometheus-memcached-exporter we expect, so i believe we can now add it to the pool.

Change #1294216 merged by Blake:

[operations/puppet@production] mcrouter_wancache: swap mc1055 for mc1054 for trixie testing

https://gerrit.wikimedia.org/r/1294216

mc1054 has been added to the pool, and things look good (also, memkeys was tested, and that worked!). i'll reimage mc1055 to trixie on monday, and then swap these servers back so we can proceed with the decom of 1054.

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1055.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1055.eqiad.wmnet with OS trixie completed:

  • mc1055 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606020914_blake_4006665_mc1055.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1296513 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] mcrouter_wancache: swap mc1054 for mc1055 to enable decom

https://gerrit.wikimedia.org/r/1296513

Change #1296513 merged by Blake:

[operations/puppet@production] mcrouter_wancache: swap mc1054 for mc1055 to enable decom

https://gerrit.wikimedia.org/r/1296513

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1056.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1057.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1056.eqiad.wmnet with OS trixie completed:

  • mc1056 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021129_blake_4105790_mc1056.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1058.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1057.eqiad.wmnet with OS trixie completed:

  • mc1057 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021133_blake_4106022_mc1057.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1059.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1058.eqiad.wmnet with OS trixie completed:

  • mc1058 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021204_blake_4137421_mc1058.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1060.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1059.eqiad.wmnet with OS trixie completed:

  • mc1059 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021207_blake_4139522_mc1059.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1061.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1060.eqiad.wmnet with OS trixie executed with errors:

  • mc1060 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021235_blake_4168021_mc1060.out
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console mc1060.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1062.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1061.eqiad.wmnet with OS trixie executed with errors:

  • mc1061 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021242_blake_4175315_mc1061.out
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console mc1061.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1063.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1062.eqiad.wmnet with OS trixie completed:

  • mc1062 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021303_blake_4191962_mc1062.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1064.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1063.eqiad.wmnet with OS trixie completed:

  • mc1063 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021323_blake_15586_mc1063.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1065.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1064.eqiad.wmnet with OS trixie completed:

  • mc1064 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021342_blake_27735_mc1064.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1066.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1065.eqiad.wmnet with OS trixie completed:

  • mc1065 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021400_blake_47007_mc1065.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1067.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1066.eqiad.wmnet with OS trixie completed:

  • mc1066 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021420_blake_69523_mc1066.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1068.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1067.eqiad.wmnet with OS trixie completed:

  • mc1067 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021434_blake_82109_mc1067.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1069.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1068.eqiad.wmnet with OS trixie completed:

  • mc1068 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021454_blake_99001_mc1068.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1070.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1069.eqiad.wmnet with OS trixie completed:

  • mc1069 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021520_blake_119509_mc1069.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1071.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1070.eqiad.wmnet with OS trixie completed:

  • mc1070 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021535_blake_132357_mc1070.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc1072.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1071.eqiad.wmnet with OS trixie completed:

  • mc1071 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021554_blake_148381_mc1071.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc1072.eqiad.wmnet with OS trixie completed:

  • mc1072 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021614_blake_162361_mc1072.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

All of the memcached hosts in the main pool in eqiad have been reimaged to Trixie.

Word from the Abstract Wikipedia folks is that you should go ahead and reimage the mc-wf hosts without any prework -- just, one at a time please.

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-wf2001.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-wf2001.codfw.wmnet with OS trixie completed:

  • mc-wf2001 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606091458_blake_2455955_mc-wf2001.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-wf2002.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-wf2002.codfw.wmnet with OS trixie completed:

  • mc-wf2002 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606091549_blake_2475963_mc-wf2002.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-wf1001.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-wf1001.eqiad.wmnet with OS trixie completed:

  • mc-wf1001 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606091630_blake_2508357_mc-wf1001.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-wf1002.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-wf1002.eqiad.wmnet with OS trixie completed:

  • mc-wf1002 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606091714_blake_2523467_mc-wf1002.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1303350 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] mcrouter_wancache: Remove 2 gutterpool servers for maintenance.

https://gerrit.wikimedia.org/r/1303350

Change #1303350 merged by Blake:

[operations/puppet@production] mcrouter_wancache: Remove 2 gutterpool servers for maintenance.

https://gerrit.wikimedia.org/r/1303350

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-gp1005.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-gp1006.eqiad.wmnet with OS trixie

Change #1303419 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] mcrouter_wancache: Swap gutterpool servers under maintenance.

https://gerrit.wikimedia.org/r/1303419

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-gp1005.eqiad.wmnet with OS trixie completed:

  • mc-gp1005 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606171310_blake_463902_mc-gp1005.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-gp1006.eqiad.wmnet with OS trixie completed:

  • mc-gp1006 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606171314_blake_464088_mc-gp1006.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1303419 merged by Blake:

[operations/puppet@production] mcrouter_wancache: Swap gutterpool servers under maintenance.

https://gerrit.wikimedia.org/r/1303419

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host mc-gp1004.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host mc-gp1004.eqiad.wmnet with OS trixie completed:

  • mc-gp1004 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606171444_blake_482916_mc-gp1004.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1303469 had a related patch set uploaded (by Blake; author: Blake):

[operations/puppet@production] mcrouter_wancache: Bring mc-gp1004 back into use.

https://gerrit.wikimedia.org/r/1303469

Change #1303469 merged by Blake:

[operations/puppet@production] mcrouter_wancache: Bring mc-gp1004 back into use.

https://gerrit.wikimedia.org/r/1303469