Page MenuHomePhabricator

ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets
Open, LowPublic

Description

eqiad rows C and D have been migrated to new Nokia switches, and more importantly to the new network design.

You can find all the information on https://wikitech.wikimedia.org/wiki/Vlan_migration

All new servers are by default getting in those new vlans, but to not have to wait for a full 5+ years server refresh cycle, we're now asking service owners to re-image their existing baremetal servers using the --move-vlan cookbook parameter, at their own pace/convenience. This will change the server's IP.
There will of course be some special cases (like Ganeti, or DBs) and that's ok to not covert 100% of the servers, but the most we can get, the better.

Please contact netops for any help.

cumin1003:~$ sudo cumin 'A:owner-serviceops and P{P:netbox::host%location ~ "[C|D].*eqiad"} and P{F:fqdn ~ ".wmnet$"} and not A:vms and not P{F:netmask = "255.255.255.0"}'
115 hosts will be targeted:
conf1009.eqiad.wmnet,kafka-main[1008-1009].eqiad.wmnet,kubestage1004.eqiad.wmnet,mc[1045-1054].eqiad.wmnet,mc-gp[1005-1006].eqiad.wmnet,mc-wf1002.eqiad.wmnet,rdb[1012,1014].eqiad.wmnet,wikikube-ctrl1003.eqiad.wmnet,wikikube-worker[1004,1016,1019-1020,1034,1036-1037,1051-1055,1062-1063,1067-1071,1083,1096-1097,1107-1110,1135-1141,1154-1157,1159-1165,1167-1168,1260-1275,1305-1306,1313,1328-1346,1348-1349,1362-1370].eqiad.wmnet

Better put:

  • conf1009.eqiad.wmnet
  • kafka-main[1008-1009].eqiad.wmnet
  • kubestage1004.eqiad.wmnet (to be refreshed 2026)
  • mc[1045-1054].eqiad.wmnet (to be refreshed 2026)
  • mc-gp[1005-1006].eqiad.wmnet re-IPed as part of T426044: Migrate Mediawiki memcached to Debian Trixie
  • mc-wf1002.eqiad.wmnet
  • rdb1012.eqiad.wmnet host is EOL
  • rdb1014.eqiad.wmnet
  • wikikube-ctrl1003.eqiad.wmnet (to be refreshed 2026)
  • wikikube-worker[1004,1016,1019-1020,1034,1062-1063,1083,1096-1097,1107-1110,1167-1168].eqiad.wmnet (to be refreshed 2026)
  • wikikube-worker[1036-1037,1051-1055,1067-1071,1135-1141,1154-1157,1159-1165,1260-1275,1305-1306,1313,1335-1346,1348-1349].eqiad.wmnet
  • wikikube-worker[1328-1334,1362-1370].eqiad.wmnet

=> wikikube-worker tracking spreadsheet

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1363.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1363.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1363 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603311404_ayounsi_395777_wikikube-worker1363.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1364.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1364.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1364 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603311449_ayounsi_455107_wikikube-worker1364.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1365.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1365.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1365 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603311629_ayounsi_521928_wikikube-worker1365.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1366.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1366.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1366 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604010654_ayounsi_699330_wikikube-worker1366.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1367.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1367.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1367 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604010759_ayounsi_709776_wikikube-worker1367.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1368.eqiad.wmnet with OS trixie

JMeybohm updated the task description. (Show Details)

We should remove servers that are do be decommissioned this calendar year from the list. Phasing them out organically should be good enough.

I've marked everything that is to be refreshed in 2026 and those that could be done as part of OS upgrades. Since at least some of those are planned for this Q, I'm scheduling this as well.

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1368.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1368 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604010909_ayounsi_744305_wikikube-worker1368.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1369.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1369.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1369 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604010950_ayounsi_811124_wikikube-worker1369.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host wikikube-worker1370.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1003 for host wikikube-worker1370.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1370 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604011040_ayounsi_863741_wikikube-worker1370.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by cgoubert@cumin1003 Renumbering for host wikikube-worker1273.eqiad.wmnet

Cookbook cookbooks.sre.k8s.renumber-node started by cgoubert@cumin1003 Renumbering for host wikikube-worker1273.eqiad.wmnet completed:

  • wikikube-worker1273.eqiad.wmnet (FAIL)
    • Failed to reimage node wikikube-worker1273.eqiad.wmnet, sre.hosts.reimage returned 2

Cookbook cookbooks.sre.k8s.renumber-node was started by cgoubert@cumin1003 Renumbering for host wikikube-worker1273.eqiad.wmnet

Cookbook cookbooks.sre.k8s.renumber-node started by cgoubert@cumin1003 Renumbering for host wikikube-worker1273.eqiad.wmnet completed:

  • wikikube-worker1273.eqiad.wmnet (FAIL)
    • Failed to cordon node wikikube-worker1273.eqiad.wmnet, sre.k8s.pool-depool-node returned 2

We should create a spreadsheet for this otherwise we'll lose track of which servers are renumbered.
For wikikube workers, I have a patch up to fix sre.k8s.renumber-node https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1268568

Change #1268568 had a related patch set uploaded (by Clément Goubert; author: Clément Goubert):

[operations/cookbooks@master] sre.k8s.renumber-node: Fix pool-depool coobook call

https://gerrit.wikimedia.org/r/1268568

Change #1268568 merged by jenkins-bot:

[operations/cookbooks@master] sre.k8s.renumber-node: Fix pool-depool coobook call

https://gerrit.wikimedia.org/r/1268568

Change #1299455 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] site.pp: reimage rdb1015 and rdb1016 as redis servers

https://gerrit.wikimedia.org/r/1299455

Change #1299455 merged by Effie Mouzeli:

[operations/puppet@production] site.pp: reimage rdb1015 and rdb1016 as redis servers

https://gerrit.wikimedia.org/r/1299455

Change #1300114 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/deployment-charts@master] mediawiki_common: update IP for rdb1014

https://gerrit.wikimedia.org/r/1300114

Change #1300114 merged by jenkins-bot:

[operations/deployment-charts@master] mediawiki_common: update IP for rdb1014

https://gerrit.wikimedia.org/r/1300114

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-worker1160.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-worker1160.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1160 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606251709_jasmine_2026974_wikikube-worker1160.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-worker1161.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-worker1161.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1161 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606251820_jasmine_2040345_wikikube-worker1161.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-worker1162.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-worker1162.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1162 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606251929_jasmine_2058708_wikikube-worker1162.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1036.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1036.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1036.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1036.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1036.eqiad.wmnet
  • wikikube-worker1036.eqiad.wmnet (PASS)
    • Host wikikube-worker1036.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1036 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071258_blake_2192371_wikikube-worker1036.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1036.eqiad.wmnet completed:

  • wikikube-worker1036.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1036.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1036.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1036.eqiad.wmnet
  • wikikube-worker1036.eqiad.wmnet (PASS)
    • Host wikikube-worker1036.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1036.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1036 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071258_blake_2192371_wikikube-worker1036.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1036.eqiad.wmnet completed:

  • wikikube-worker1036.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1036.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1036.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1036.eqiad.wmnet
  • wikikube-worker1036.eqiad.wmnet (PASS)
    • Host wikikube-worker1036.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1036.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1036 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071258_blake_2192371_wikikube-worker1036.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1037.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1037.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1037.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1037.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1037.eqiad.wmnet
  • wikikube-worker1037.eqiad.wmnet (PASS)
    • Host wikikube-worker1037.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1037 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071454_blake_2218232_wikikube-worker1037.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1037.eqiad.wmnet completed:

  • wikikube-worker1037.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1037.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1037.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1037.eqiad.wmnet
  • wikikube-worker1037.eqiad.wmnet (PASS)
    • Host wikikube-worker1037.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1037.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1037 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071454_blake_2218232_wikikube-worker1037.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1037.eqiad.wmnet completed:

  • wikikube-worker1037.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1037.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1037.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1037.eqiad.wmnet
  • wikikube-worker1037.eqiad.wmnet (PASS)
    • Host wikikube-worker1037.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1037.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1037 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607071454_blake_2218232_wikikube-worker1037.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1051.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1051.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1051.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1051.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1051.eqiad.wmnet
  • wikikube-worker1051.eqiad.wmnet (PASS)
    • Host wikikube-worker1051.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1051 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607081305_blake_2605050_wikikube-worker1051.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1051.eqiad.wmnet completed:

  • wikikube-worker1051.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1051.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1051.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1051.eqiad.wmnet
  • wikikube-worker1051.eqiad.wmnet (PASS)
    • Host wikikube-worker1051.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1051.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1051 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607081305_blake_2605050_wikikube-worker1051.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1051.eqiad.wmnet completed:

  • wikikube-worker1051.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1051.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1051.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1051.eqiad.wmnet
  • wikikube-worker1051.eqiad.wmnet (PASS)
    • Host wikikube-worker1051.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1051.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1051 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607081305_blake_2605050_wikikube-worker1051.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1052.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1052.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1052.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1052.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1052.eqiad.wmnet
  • wikikube-worker1052.eqiad.wmnet (PASS)
    • Host wikikube-worker1052.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1052 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607091359_blake_3243349_wikikube-worker1052.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1052.eqiad.wmnet completed:

  • wikikube-worker1052.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1052.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1052.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1052.eqiad.wmnet
  • wikikube-worker1052.eqiad.wmnet (PASS)
    • Host wikikube-worker1052.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1052.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1052 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607091359_blake_3243349_wikikube-worker1052.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1052.eqiad.wmnet completed:

  • wikikube-worker1052.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1052.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1052.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1052.eqiad.wmnet
  • wikikube-worker1052.eqiad.wmnet (PASS)
    • Host wikikube-worker1052.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1052.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1052 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607091359_blake_3243349_wikikube-worker1052.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1053.eqiad.wmnet

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1053.eqiad.wmnet completed:

  • wikikube-worker1053.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1053.eqiad.wmnet
    • Failed to reimage node wikikube-worker1053.eqiad.wmnet, sre.hosts.reimage returned 94
  • wikikube-worker1053.eqiad.wmnet (PASS)
    • Host wikikube-worker1053.eqiad.wmnet depooled from wikikube-eqiad

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1053.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1053.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1053.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1053.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1053.eqiad.wmnet
  • wikikube-worker1053.eqiad.wmnet (PASS)
    • Host wikikube-worker1053.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1053 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202607100941_blake_3370018_wikikube-worker1053.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101045_blake_3370018_wikikube-worker1053.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1053.eqiad.wmnet completed:

  • wikikube-worker1053.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1053.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1053.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1053.eqiad.wmnet
  • wikikube-worker1053.eqiad.wmnet (PASS)
    • Host wikikube-worker1053.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1053.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1053 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202607100941_blake_3370018_wikikube-worker1053.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101045_blake_3370018_wikikube-worker1053.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1053.eqiad.wmnet completed:

  • wikikube-worker1053.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1053.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1053.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1053.eqiad.wmnet
  • wikikube-worker1053.eqiad.wmnet (PASS)
    • Host wikikube-worker1053.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1053.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1053 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202607100941_blake_3370018_wikikube-worker1053.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101045_blake_3370018_wikikube-worker1053.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node was started by blake@cumin1003 Renumbering for host wikikube-worker1054.eqiad.wmnet

Cookbook cookbooks.sre.hosts.reimage was started by blake@cumin1003 for host wikikube-worker1054.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by blake@cumin1003 for host wikikube-worker1054.eqiad.wmnet with OS trixie completed:

  • wikikube-worker1054.eqiad.wmnet (PASS)
    • Successfully cordoned node wikikube-worker1054.eqiad.wmnet
  • wikikube-worker1054.eqiad.wmnet (PASS)
    • Host wikikube-worker1054.eqiad.wmnet depooled from wikikube-eqiad
  • wikikube-worker1054 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101333_blake_3397387_wikikube-worker1054.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1054.eqiad.wmnet completed:

  • wikikube-worker1054.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1054.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1054.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on registry servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1054.eqiad.wmnet
  • wikikube-worker1054.eqiad.wmnet (PASS)
    • Host wikikube-worker1054.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1054.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1054 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101333_blake_3397387_wikikube-worker1054.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.k8s.renumber-node started by blake@cumin1003 Renumbering for host wikikube-worker1054.eqiad.wmnet completed:

  • wikikube-worker1054.eqiad.wmnet (FAIL)
    • Successfully cordoned node wikikube-worker1054.eqiad.wmnet
    • Successfully reimaged node wikikube-worker1054.eqiad.wmnet
    • Failed to run puppet agent on deployment servers
    • Failed to run puppet agent on registry servers
    • Successfully ran puppet agent on registry servers
    • Pooled and uncordoned node wikikube-worker1054.eqiad.wmnet
  • wikikube-worker1054.eqiad.wmnet (PASS)
    • Host wikikube-worker1054.eqiad.wmnet depooled from wikikube-eqiad
    • Host wikikube-worker1054.eqiad.wmnet pooled in wikikube-eqiad
  • wikikube-worker1054 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607101333_blake_3397387_wikikube-worker1054.out
    • Unable to run puppet on config-master2001.codfw.wmnet,config-master1001.eqiad.wmnet to update configmaster.wikimedia.org with the new host SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB