Page MenuHomePhabricator

[Post kafka-main 3.7 upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers
Closed, ResolvedPublic

Description

Splitting off the kafka-main Debian Trixie upgrade work from T419216: Upgrade kafka-main to Kafka 3.7 for better visibility
and consistency with T426835: Upgrade kafka-jumbo to trixie/JDK21 and T417001: Upgrade Observability Kafka-logging hosts to trixie

Having upgraded kafka-main to kafka 3.7, we're implementing the final part of the upgrade plan enacted by the Kafka upgrade WG, which is to upgrade the brokers to Debian Trixie / jdk21.

Additional migrations/upgrades:

codfw: T428191: ServiceOps: Re-IP codfw private baremetal hosts to new per-rack vlans/subnets
eqiad: T421711: ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets

The VLAN migration practically means running the reimage cookbook with the --move-vlan flag. This will change the IP of the host which needs additional attention:

  • Update kafka_brokers_main in hieradata/common.yaml, run puppet on conf100* and the other brokers in eqiad
  • Update external-services in k8s.
    • Run puppet on deployment server(s)
    • Deploy the external service release to update network policies to include the new broker IP (and remove the old)
    • Roll restart eventgate-main out of precaution since it's pretty picky with loosing brokers

Post-upgrade cleanup:

  • Clean up per-host hiera configs, move to role level

Cluster states:

kafka-main codfw:

Kafka BrokerConfluent distribution 77Inter-broker protocolHost-level override patchDebian Trixie
kafka-main2006.codfw.wmnetUpgraded ✅3.7 ✅1288917
kafka-main2007.codfw.wmnetUpgraded ✅3.7 ✅1288918
kafka-main2008.codfw.wmnetUpgraded ✅3.7 ✅1288919
kafka-main2009.codfw.wmnetUpgraded ✅3.7 ✅1288920
kafka-main2010.codfw.wmnetUpgraded ✅3.7 ✅1288921

kafka-main eqiad:

Kafka BrokerConfluent distribution 77Inter-broker protocolHost-level override patchDebian Trixie
kafka-main1006.eqiad.wmnetUpgraded ✅3.7 ✅1285474
kafka-main1007.eqiad.wmnetUpgraded ✅3.7 ✅1285475
kafka-main1008.eqiad.wmnetUpgraded ✅3.7 ✅1285476
kafka-main1009.eqiad.wmnetUpgraded ✅3.7 ✅1285477
kafka-main1010.eqiad.wmnetUpgraded ✅3.7 ✅1285478

Details

Related Changes in Gerrit:
SubjectAuthorRepoBranchLines +/-
Jasmineoperations/puppetproduction+0 -16
Jasmineoperations/puppetproduction+4 -20
Jasmineoperations/puppetproduction+2 -2
Jasmineoperations/puppetproduction+4 -0
Jasmineoperations/puppetproduction+2 -2
Jasmineoperations/puppetproduction+2 -2
Jasmineoperations/puppetproduction+2 -2
Jasmineoperations/puppetproduction+4 -0
Jasmineoperations/puppetproduction+2 -2
Jasmineoperations/puppetproduction+4 -0
Jasmineoperations/puppetproduction+4 -0
JMeybohmoperations/puppetproduction+8 -5
JMeybohmoperations/puppetproduction+2 -1
JMeybohmoperations/puppetproduction+3 -0
JMeybohmoperations/puppetproduction+1 -1
Jasmineoperations/puppetproduction+4 -0
Show related patches Customize query in gerrit

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes
jasmine_ renamed this task from Upgrade kafka-main brokers to Trixie/JDK21 to [Post kafka 3.7 upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers.May 22 2026, 5:48 PM
jasmine_ renamed this task from [Post kafka 3.7 upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers to [Post kafka 3.7 on kafka-main upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers.
jasmine_ updated the task description. (Show Details)
jasmine_ renamed this task from [Post kafka 3.7 on kafka-main upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers to [Post kafka-main 3.7 upgrade work] Reimage brokers to trixie/JDK21 & vlan migrations on select brokers.May 22 2026, 5:58 PM
jasmine_ updated the task description. (Show Details)

Change #1288917 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main2006: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288917

Change #1288918 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main2007: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288918

Change #1288919 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main2008: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288919

Change #1288920 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main2009: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288920

Change #1288921 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main2010: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288921

Change #1288917 merged by Jasmine:

[operations/puppet@production] kafka-main2006: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288917

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie executed with errors:

  • kafka-main2006 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console kafka-main2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie executed with errors:

The debian installer is hanging in partman, asking if we're sure to set the host up without swap. Maybe there is a problem with the partman receipt since the md has been assembled correctly and the logical volumes have been scanned as well:

~ # cat /proc/mdstat 
Personalities : [raid10] 
md0 : active raid10 sdb2[0] sdf2[5] sde2[4] sdd2[3] sdc2[2] sda2[1]
      2811801600 blocks super 1.2 512K chunks 2 near-copies [6/6] [UUUUUU]
      bitmap: 0/21 pages [0KB], 65536KB chunk

unused devices: <none>
~ # lvdisplay 
  --- Logical volume ---
  LV Path                /dev/vg0/swap
  LV Name                swap
  VG Name                vg0
  LV UUID                2TWWua-46AX-ckM8-bokB-uer3-mSWu-e84XGo
  LV Write Access        read/write
  LV Creation host, time kafka-main2006, 2024-08-15 08:12:46 +0000
  LV Status              available
  # open                 0
  LV Size                976.00 MiB
  Current LE             244
  Segments               1
  Allocation             inherit
  Read ahead sectors     auto
  - currently set to     8386560
  Block device           253:0
   
  --- Logical volume ---
  LV Path                /dev/vg0/root
  LV Name                root
  VG Name                vg0
  LV UUID                kB47RB-iMRc-zwxl-7O0X-qONV-OtkS-m9hQAb
  LV Write Access        read/write
  LV Creation host, time kafka-main2006, 2024-08-15 08:12:47 +0000
  LV Status              available
  # open                 0
  LV Size                74.50 GiB
  Current LE             19073
  Segments               1
  Allocation             inherit
  Read ahead sectors     auto
  - currently set to     8386560
  Block device           253:1
   
  --- Logical volume ---
  LV Path                /dev/vg0/srv
  LV Name                srv
  VG Name                vg0
  LV UUID                6SIWCc-Ay9g-f0DU-Pn8S-lhga-xCTQ-thBZUn
  LV Write Access        read/write
  LV Creation host, time kafka-main2006, 2024-08-15 08:12:47 +0000
  LV Status              available
  # open                 0
  LV Size                2.02 TiB
  Current LE             529861
  Segments               1
  Allocation             inherit
  Read ahead sectors     auto
  - currently set to     8386560
  Block device           253:2

Change #1296553 had a related patch set uploaded (by JMeybohm; author: JMeybohm):

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Use linux-swap as fs identifier

https://gerrit.wikimedia.org/r/1296553

The patch might be not enough (will test anyways). If not, we might need to add the workaround from T408777: Swap partition issue when installing a DB with Debian Trixie

Change #1296553 merged by JMeybohm:

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Use linux-swap as fs identifier

https://gerrit.wikimedia.org/r/1296553

Cookbook cookbooks.sre.hosts.reimage was started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie executed with errors:

  • kafka-main2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console kafka-main2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie

Change #1296597 had a related patch set uploaded (by JMeybohm; author: JMeybohm):

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Apply workaround to swap handling affecting trixie installations

https://gerrit.wikimedia.org/r/1296597

Change #1296597 merged by JMeybohm:

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Apply workaround to swap handling affecting trixie installations

https://gerrit.wikimedia.org/r/1296597

Cookbook cookbooks.sre.hosts.reimage started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie executed with errors:

  • kafka-main2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console kafka-main2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie

Change #1296613 had a related patch set uploaded (by JMeybohm; author: JMeybohm):

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Apply swap workaround

https://gerrit.wikimedia.org/r/1296613

Change #1296613 merged by JMeybohm:

[operations/puppet@production] partman/reuse-raid10-6dev.cfg: Apply swap workaround

https://gerrit.wikimedia.org/r/1296613

Cookbook cookbooks.sre.hosts.reimage started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie executed with errors:

  • kafka-main2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console kafka-main2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie

With the two workarounds the installer does on longer hang on asking for swap. Unfortunately it still failed on first try to create a filesystem for the root partition (without saying so ofc). Restarting the reimage did work but it did not include the swap partition in /etc/fstab (and did not activate it).
I will complete the reimage so we have the broker back and it can catch up with what it missed (since /srv seems to not have been formatted, yay) but before taking on the next broker we should figure out what's happening/not happening here.

Maybe it is possible to test the recipe with a 6 disk ganeti vm to have a workable feedback loop time and not interfere with the kafka-main cluster.

Cookbook cookbooks.sre.hosts.reimage started by jayme@cumin2002 for host kafka-main2006.codfw.wmnet with OS trixie completed:

  • kafka-main2006 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606021609_jayme_3605330_kafka-main2006.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1297201 had a related patch set uploaded (by JMeybohm; author: JMeybohm):

[operations/puppet@production] reuse-raid10-6dev.cfg: Fix swap reuse and grub-install on all disks

https://gerrit.wikimedia.org/r/1297201

Change #1297201 merged by JMeybohm:

[operations/puppet@production] reuse-raid10-6dev.cfg: Fix swap reuse and grub-install on all disks

https://gerrit.wikimedia.org/r/1297201

Change #1288918 merged by JMeybohm:

[operations/puppet@production] kafka-main2007: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288918

Cookbook cookbooks.sre.hosts.reimage was started by jayme@cumin1003 for host kafka-main2007.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jayme@cumin1003 for host kafka-main2007.codfw.wmnet with OS trixie completed:

  • kafka-main2007 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606041524_jayme_1135974_kafka-main2007.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1288919 merged by Jasmine:

[operations/puppet@production] kafka-main2008: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288919

The fix from T428078: Swap partition not used when reimaging to trixie with reuse-parts.cfg and mdadm + lvm works fine so we may continue with the reimages as planned @jasmine_

Thanks Janis! Much appreciated~ Proceeding with remaining brokers momentarily.

Edit: It appears that all remaining brokers in codfw kafka-main are on legacy vlan, captured in T428191: ServiceOps: Re-IP codfw private baremetal hosts to new per-rack vlans/subnets.
We will need to move vlan on these as with kafka-main100[8-9].

For the time being, will proceed with brokers on the non legacy vlan in eqiad, then to the brokers requiring vlan migration following.

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main1006.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main1006.eqiad.wmnet with OS trixie completed:

  • kafka-main1006 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606042340_jasmine_86465_kafka-main1006.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main1007.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main1007.eqiad.wmnet with OS trixie completed:

  • kafka-main1007 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606050040_jasmine_100324_kafka-main1007.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main1010.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main1010.eqiad.wmnet with OS trixie completed:

  • kafka-main1010 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606050139_jasmine_114239_kafka-main1010.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

The fix from T428078: Swap partition not used when reimaging to trixie with reuse-parts.cfg and mdadm + lvm works fine so we may continue with the reimages as planned @jasmine_

Thanks Janis! Much appreciated~ Proceeding with remaining brokers momentarily.

Edit: It appears that all remaining brokers in codfw kafka-main are on legacy vlan, captured in T428191: ServiceOps: Re-IP codfw private baremetal hosts to new per-rack vlans/subnets.
We will need to move vlan on these as with kafka-main100[8-9].

For the time being, will proceed with brokers on the non legacy vlan in eqiad, then to the brokers requiring vlan migration following.

All kafka-main brokers on non-legacy vlan are now on trixie in both DCs. The remaining brokers will require a vlan migration prior to the trixie reimage.

kafka-main100[8-9], kafka-main200[8-9], kafka-main2010

codfw: T428191: ServiceOps: Re-IP codfw private baremetal hosts to new per-rack vlans/subnets
eqiad: T421711: ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets

Since the vlan migrations will require a deployment, I plan to pick them up starting Monday.
Note: Additionally the cleanup for host level overrides for the eqiad cluster will be applied in kafka-main1009 as it will be the final broker to be reimaged to trixie (rather than kafka-main1010)

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main2008.codfw.wmnet with OS trixie

Change #1299570 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2008, following vlan migrations

https://gerrit.wikimedia.org/r/1299570

Change #1299570 merged by Jasmine:

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2008, following vlan migrations

https://gerrit.wikimedia.org/r/1299570

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2008.codfw.wmnet with OS trixie executed with errors:

  • kafka-main2008 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console kafka-main2008.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main2008.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2008.codfw.wmnet with OS trixie completed:

  • kafka-main2008 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606091858_jasmine_1492495_kafka-main2008.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main2009.codfw.wmnet with OS trixie

Change #1288920 merged by Jasmine:

[operations/puppet@production] kafka-main2009: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288920

Change #1299656 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2009, following vlan migrations

https://gerrit.wikimedia.org/r/1299656

Change #1299656 merged by Jasmine:

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2009, following vlan migrations

https://gerrit.wikimedia.org/r/1299656

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2009.codfw.wmnet with OS trixie completed:

  • kafka-main2009 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606100043_jasmine_1563623_kafka-main2009.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main1008.eqiad.wmnet with OS trixie

Change #1299661 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main1008, following vlan migrations

https://gerrit.wikimedia.org/r/1299661

Change #1299661 merged by Jasmine:

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main1008, following vlan migrations

https://gerrit.wikimedia.org/r/1299661

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main1008.eqiad.wmnet with OS trixie completed:

  • kafka-main1008 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606100149_jasmine_1578727_kafka-main1008.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main2010.codfw.wmnet with OS trixie

Change #1300208 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2010, following vlan migration

https://gerrit.wikimedia.org/r/1300208

Change #1300208 merged by Jasmine:

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main2010, following vlan migration

https://gerrit.wikimedia.org/r/1300208

Change #1288921 merged by Jasmine:

[operations/puppet@production] kafka-main2010: apply host-level override in advance of trixie upgrade [0]

https://gerrit.wikimedia.org/r/1288921

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main2010.codfw.wmnet with OS trixie completed:

  • kafka-main2010 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606101729_jasmine_1790524_kafka-main2010.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host kafka-main1009.eqiad.wmnet with OS trixie

Change #1300281 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main1009 following vlan migration

https://gerrit.wikimedia.org/r/1300281

Change #1300281 merged by Jasmine:

[operations/puppet@production] hieradata/common.yaml: add new IPs for kafka-main1009 following vlan migration

https://gerrit.wikimedia.org/r/1300281

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host kafka-main1009.eqiad.wmnet with OS trixie completed:

  • kafka-main1009 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606110102_jasmine_1891761_kafka-main1009.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Change #1300287 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main: clean up host level overrides for kafka-main jdk 21 in eqiad

https://gerrit.wikimedia.org/r/1300287

Change #1300288 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] kafka-main: clean up host level overrides for kafka-main jdk 21 in codfw

https://gerrit.wikimedia.org/r/1300288

Change #1300287 merged by Jasmine:

[operations/puppet@production] kafka-main: clean up host level overrides for kafka-main jdk 21 in eqiad

https://gerrit.wikimedia.org/r/1300287

Change #1300288 merged by Jasmine:

[operations/puppet@production] kafka-main: clean up host level overrides for kafka-main jdk 21 in codfw

https://gerrit.wikimedia.org/r/1300288

Resolving as all kafka-main brokers are now on Trixie and legacy vlan brokers have also been migrated :) Thanks @JMeybohm for supporting!