Page MenuHomePhabricator

Decommission or recommission all snapshot and dumpsdata servers
Closed, ResolvedPublic

Description

Now that the current generation of dumps is running from Airflow and Kubernetes, we can re-purpose the previous bare-metal servers.

The following two servers are EoL and should be decommissioned:

  • dumpsdata1003
  • snapshot1010

The remaining snapshot servers (https://wikitech.wikimedia.org/wiki/Dumps/Snapshot_hosts) will likely make good dse-k8s-worker nodes, since they have relatively powerful CPUs, although only 64 GB of RAM each.
The snapshot servers have been renamed and re-commissioned like this.

  • snapshot1011 -> dse-k8s-worker1015
  • snapshot1012 -> dse-k8s-worker1016
  • snapshot1013 -> dse-k8s-worker1017
  • snapshot1015 -> dse-k8s-worker1018
  • snapshot1016 -> dse-k8s-worker1019

The dumpsdata servers are to be renamed and recommissioned like this:

  • dumpsdata1004 -> an-worker1233
  • dumpsdata1005 -> an-worker1234
  • dumpsdata1006 -> an-worker1235
  • dumpsdata1007 -> an-worker1236

The dumpsdata servers (https://wikitech.wikimedia.org/wiki/Dumps/Dumpsdata_hosts) have lower-spec CPU and are 2U servers with 12 x 4TB disks each.
These are a little under-specced as either Hadoop or dse-k8s workers, but we might be able to find a role for them, or we could simply mark them as spare for DC-Ops to allocate accordingly.

Details

Other Assignee
brouberol
Related Changes in Gerrit:
SubjectAuthorRepoBranchLines +/-
Btullisoperations/puppetproduction+1 -5
Btullisoperations/puppetproduction+29 -6
Btullisoperations/puppetproduction+7 -3
Btullisoperations/puppetproduction+5 -1
Btullisoperations/puppetproduction+0 -147
Btullisoperations/puppetproduction+1 -1
Btullisoperations/puppetproduction+2 -2
Brouberoloperations/puppetproduction+1 -1
Brouberoloperations/puppetproduction+5 -3
Brouberoloperations/puppetproduction+4 -2
Brouberoloperations/puppetproduction+4 -2
Brouberoloperations/puppetproduction+4 -2
Brouberoloperations/puppetproduction+6 -2
Btullisoperations/puppetproduction+0 -9
Btullisoperations/dnsmaster+0 -4
Btullisoperations/puppetproduction+2 -2
Btullisoperations/puppetproduction+2 -7 K
Btullisoperations/puppetproduction+0 -59
Btullisoperations/puppetproduction+0 -1
Btullisoperations/puppetproduction+2 -66
Btullisoperations/puppetproduction+7 -38
Btullisoperations/puppetproduction+1 -6
Btullisoperations/puppetproduction+19 -19
Show related patches Customize query in gerrit

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes

Reopening because we still haven't re-commisisoned the four remaining dumpsdata servers.
I've asked in T401299: Investigate whether we can add RAM to dumpsdata100[4-7] from any decommissioned hosts whether we can find some RAM to allocate so that they would be useful as Hadoop worker nodes.

Change #1184068 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Add four hadoop workers from repurposed dumpsdata hosts

https://gerrit.wikimedia.org/r/1184068

Change #1184070 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Remove references to dumpsdata servers

https://gerrit.wikimedia.org/r/1184070

Change #1184068 merged by Btullis:

[operations/puppet@production] Add four hadoop workers from repurposed dumpsdata hosts

https://gerrit.wikimedia.org/r/1184068

Cookbook cookbooks.sre.hosts.rename started by btullis@cumin1003 from dumpsdata1004 to an-worker1233 completed:

  • dumpsdata1004 (PASS)
    • ✔️ Downtimed host on Icinga/Alertmanager
    • ✔️ Disabled puppet and its timer
    • ✔️ Disabled debmonitor-client timer
    • ✔️ Netbox updated
    • ✔️ BMC Hostname updated
    • ✔️ DNS updated
    • ✔️ Switch description updated
    • ✔️ Removed from DebMonitor
    • ✔️ Removed from Puppet master and PuppetDB
    • Rename completed 👍 - now please run the re-image cookbook on the new name with --new

Cookbook cookbooks.sre.hosts.rename started by btullis@cumin1003 from dumpsdata1005 to an-worker1234 completed:

  • dumpsdata1005 (PASS)
    • ✔️ Downtimed host on Icinga/Alertmanager
    • ✔️ Disabled puppet and its timer
    • ✔️ Disabled debmonitor-client timer
    • ✔️ Netbox updated
    • ✔️ BMC Hostname updated
    • ✔️ DNS updated
    • ✔️ Switch description updated
    • ✔️ Removed from DebMonitor
    • ✔️ Removed from Puppet master and PuppetDB
    • Rename completed 👍 - now please run the re-image cookbook on the new name with --new

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye

Change #1184489 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Fix the partman config for the newly renamed an-worker123[3-6]

https://gerrit.wikimedia.org/r/1184489

Change #1184489 merged by Btullis:

[operations/puppet@production] Fix the partman config for the newly renamed an-worker123[3-6]

https://gerrit.wikimedia.org/r/1184489

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1234 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1234.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye completed:

  • an-worker1234 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509031129_btullis_825303_an-worker1234.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.rename started by btullis@cumin1003 from dumpsdata1006 to an-worker1235 completed:

  • dumpsdata1006 (PASS)
    • ✔️ Downtimed host on Icinga/Alertmanager
    • ✔️ Disabled puppet and its timer
    • ✔️ Disabled debmonitor-client timer
    • ✔️ Netbox updated
    • ✔️ BMC Hostname updated
    • ✔️ DNS updated
    • ✔️ Switch description updated
    • ✔️ Removed from DebMonitor
    • ✔️ Removed from Puppet master and PuppetDB
    • Rename completed 👍 - now please run the re-image cookbook on the new name with --new

Cookbook cookbooks.sre.hosts.rename started by btullis@cumin1003 from dumpsdata1007 to an-worker1236 completed:

  • dumpsdata1007 (PASS)
    • ✔️ Downtimed host on Icinga/Alertmanager
    • ✔️ Disabled puppet and its timer
    • ✔️ Disabled debmonitor-client timer
    • ✔️ Netbox updated
    • ✔️ BMC Hostname updated
    • ✔️ DNS updated
    • ✔️ Switch description updated
    • ✔️ Removed from DebMonitor
    • ✔️ Removed from Puppet master and PuppetDB
    • Rename completed 👍 - now please run the re-image cookbook on the new name with --new

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1236 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1236.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye completed:

  • an-worker1235 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509050923_btullis_1116200_an-worker1235.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye completed:

  • an-worker1236 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509050927_btullis_1116245_an-worker1236.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1233 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1233.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye

Change #1184070 merged by Btullis:

[operations/puppet@production] Remove references to dumpsdata servers

https://gerrit.wikimedia.org/r/1184070

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1233 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1233.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1233 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1233.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye completed:

  • an-worker1233 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509081413_btullis_1622199_an-worker1233.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Mentioned in SAL (#wikimedia-releng) [2025-09-09T16:24:12Z] <Krinkle> Remove unused deployment-snapshot- prefix Hiera file. https://gerrit.wikimedia.org/r/plugins/gitiles/cloud/instance-puppet/+/0c6a99c4e12e9d1a55ce353e302d7dc891ae1177 - Follows-up decom of deployment-snapshot05 host. ref T398438, T360995

I have created the kerberos credentials for the four new Hadoop workers.

btullis@krb1002:~$ sudo generate_keytabs.py --realm WIKIMEDIA hadoop_creds_T398438.txt 
HTTP/an-worker1233.eqiad.wmnet@WIKIMEDIA
hdfs/an-worker1233.eqiad.wmnet@WIKIMEDIA
yarn/an-worker1233.eqiad.wmnet@WIKIMEDIA
Entry for principal HTTP/an-worker1233.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1233.eqiad.wmnet/hadoop/HTTP.keytab.
Entry for principal hdfs/an-worker1233.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1233.eqiad.wmnet/hadoop/hdfs.keytab.
Entry for principal yarn/an-worker1233.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1233.eqiad.wmnet/hadoop/yarn.keytab.
HTTP/an-worker1234.eqiad.wmnet@WIKIMEDIA
hdfs/an-worker1234.eqiad.wmnet@WIKIMEDIA
yarn/an-worker1234.eqiad.wmnet@WIKIMEDIA
Entry for principal HTTP/an-worker1234.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1234.eqiad.wmnet/hadoop/HTTP.keytab.
Entry for principal hdfs/an-worker1234.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1234.eqiad.wmnet/hadoop/hdfs.keytab.
Entry for principal yarn/an-worker1234.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1234.eqiad.wmnet/hadoop/yarn.keytab.
HTTP/an-worker1235.eqiad.wmnet@WIKIMEDIA
hdfs/an-worker1235.eqiad.wmnet@WIKIMEDIA
yarn/an-worker1235.eqiad.wmnet@WIKIMEDIA
Entry for principal HTTP/an-worker1235.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1235.eqiad.wmnet/hadoop/HTTP.keytab.
Entry for principal hdfs/an-worker1235.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1235.eqiad.wmnet/hadoop/hdfs.keytab.
Entry for principal yarn/an-worker1235.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1235.eqiad.wmnet/hadoop/yarn.keytab.
HTTP/an-worker1236.eqiad.wmnet@WIKIMEDIA
hdfs/an-worker1236.eqiad.wmnet@WIKIMEDIA
yarn/an-worker1236.eqiad.wmnet@WIKIMEDIA
Entry for principal HTTP/an-worker1236.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1236.eqiad.wmnet/hadoop/HTTP.keytab.
Entry for principal hdfs/an-worker1236.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1236.eqiad.wmnet/hadoop/hdfs.keytab.
Entry for principal yarn/an-worker1236.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-worker1236.eqiad.wmnet/hadoop/yarn.keytab.
btullis@krb1002:~$

I created the new RAID0 logical volumes.

On the two older hosts, an-worker123[34], I used the older megacli command:

sudo  megacli -CfgEachDskRaid0 WB RA Direct CachedBadBBU -a0

On the two newer hosts, an-worker123[56],I used the newer perccli64 command:

sudo perccli64 /c0 add vd each r0 wb ra

Now I am able to initialize these with the sre.hadoop.init-hadoop-workers cookbook.

Change #1186998 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Add four new (renamed) an-worker nodes to the Hadoop cluster

https://gerrit.wikimedia.org/r/1186998

Change #1186998 merged by Btullis:

[operations/puppet@production] Add four new (renamed) an-worker nodes to the Hadoop cluster

https://gerrit.wikimedia.org/r/1186998

Change #1187029 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Temporarily exlude the 4 new hadoop workers to facilitate vlan change

https://gerrit.wikimedia.org/r/1187029

Change #1187029 merged by Btullis:

[operations/puppet@production] Temporarily exlude the 4 new hadoop workers to facilitate vlan change

https://gerrit.wikimedia.org/r/1187029

cookbooks.sre.hosts.decommission executed by btullis@cumin1003 for hosts: an-worker[1233-1236].eqiad.wmnet

  • an-worker1233.eqiad.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found physical host
    • Downtimed management interface on Alertmanager
    • Wiped all swraid, partition-table and filesystem signatures
    • Powered off
    • [Netbox] Set status to Decommissioning, deleted all non-mgmt IPs, updated switch interfaces (disabled, removed vlans, etc)
    • Configured the linked switch interface(s)
    • Removed from DebMonitor
    • Removed from Puppet master and PuppetDB
  • an-worker1234.eqiad.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found physical host
    • Downtimed management interface on Alertmanager
    • Wiped all swraid, partition-table and filesystem signatures
    • Powered off
    • [Netbox] Set status to Decommissioning, deleted all non-mgmt IPs, updated switch interfaces (disabled, removed vlans, etc)
    • Configured the linked switch interface(s)
    • Removed from DebMonitor
    • Removed from Puppet master and PuppetDB
  • an-worker1235.eqiad.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found physical host
    • Downtimed management interface on Alertmanager
    • Wiped all swraid, partition-table and filesystem signatures
    • Powered off
    • [Netbox] Set status to Decommissioning, deleted all non-mgmt IPs, updated switch interfaces (disabled, removed vlans, etc)
    • Configured the linked switch interface(s)
    • Removed from DebMonitor
    • Removed from Puppet master and PuppetDB
  • an-worker1236.eqiad.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found physical host
    • Downtimed management interface on Alertmanager
    • Wiped all swraid, partition-table and filesystem signatures
    • Powered off
    • [Netbox] Set status to Decommissioning, deleted all non-mgmt IPs, updated switch interfaces (disabled, removed vlans, etc)
    • Configured the linked switch interface(s)
    • Removed from DebMonitor
    • Removed from Puppet master and PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1233.eqiad.wmnet with OS bullseye completed:

  • an-worker1233 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run failed and logged in /var/log/spicerack/sre/hosts/reimage/202509111544_btullis_2117294_an-worker1233.out, asking the operator what to do
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509111639_btullis_2117294_an-worker1233.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • Failed to run the sre.puppet.sync-netbox-hiera cookbook, run it manually

Mentioned in SAL (#wikimedia-operations) [2025-09-12T08:43:43Z] <btullis@cumin1003> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "T398438 - btullis@cumin1003"

Mentioned in SAL (#wikimedia-operations) [2025-09-12T08:43:48Z] <btullis@cumin1003> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "T398438 - btullis@cumin1003"

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1234.eqiad.wmnet with OS bullseye completed:

  • an-worker1234 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509120914_btullis_2225247_an-worker1234.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1236 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1236.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1235 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1235.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1236 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1236.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1235.eqiad.wmnet with OS bullseye completed:

  • an-worker1235 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509291259_btullis_494219_an-worker1235.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is not optimal, downtime not removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1236 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1236.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1236.eqiad.wmnet with OS bullseye completed:

  • an-worker1236 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202509291424_btullis_505620_an-worker1236.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is not optimal, downtime not removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Change #1192239 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Add 28 new hadoop workers to the analytics_hadoop cluster

https://gerrit.wikimedia.org/r/1192239

Change #1192239 merged by Btullis:

[operations/puppet@production] Add 28 new hadoop workers to the analytics_hadoop cluster

https://gerrit.wikimedia.org/r/1192239

BTullis updated the task description. (Show Details)

All four of the remaining dumpsdata servers have now been re-commissioned as Hadoop workers, so this ticket can now be resolved.

Change #1195154 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Re-enable YARN and HDFS on an-worker123[3-6]

https://gerrit.wikimedia.org/r/1195154

Change #1195154 merged by Btullis:

[operations/puppet@production] Re-enable YARN and HDFS on an-worker123[3-6]

https://gerrit.wikimedia.org/r/1195154