Page MenuHomePhabricator

Repurpose ganeti102[3456] for Zuul migration
Open, MediumPublic

Description

In T424680 ganeti102[3456] were decommissioned as Ganeti servers, but they will be repurposed for the Zuul migration. The node have been decommissioned on the OS/puppet level, so they need to be

  • re-provisioned under the new names (host names to be provided by SRE Collab)
  • Ganeti servers had special network requirements for the VLANs, these no longer exist for the Zuul servers, as such it would be also be possible to move these servers to other racks if that's helpful in terms of rack space and/or network switches
  • physical labels need to be updated
  • leave a comment in column H of the EOL Server Document once the renamed servers are added there, to help track of their current status
  • reimaged to the new OSes

Event Timeline

We have already established a zuul naming pattern for existing VMs and "1-3" are in use.

Please use zuul1004/zuul2004 and counting up from there.

Are these machines supposed to replace the main zuul VMs zuul1001/2001? I am missing the context a bit how we got to physical hardware being assigned now.

Change #1294347 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] site: add zuul[12]00[4-9] with insetup role

https://gerrit.wikimedia.org/r/1294347

Change #1294348 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] installserver: update partman for mixed VM/physical zuul machines

https://gerrit.wikimedia.org/r/1294348

@thcipriani @dduvall Was this requested by you because the existing VMs are too limited? Is the idea to replace (just) the "main" zuul VMs or also executors and trusted build nodes? Or is nothing being replaced and we just have this in addition?

Change #1294348 merged by Dzahn:

[operations/puppet@production] installserver: update partman for mixed VM/physical zuul machines

https://gerrit.wikimedia.org/r/1294348

Change #1294347 merged by Dzahn:

[operations/puppet@production] site: add zuul[12]00[4-9] with insetup role

https://gerrit.wikimedia.org/r/1294347

Hey @Dzahn thank you! Currently we can move forward with this. Currently where the servers sit right now may be the best spot for them unless you would like them in a different configuration? Other than that I will proceed to update the labels and netbox with the new information for them to be prepped for repurposing.

The machines are already in puppet site.pp and partman and can be installed with an OS any time.

@Dzahn I had a question. In the ticket it lists several ganeti servers and zuul servers.

ganeti1023 - zuul1004
ganeti1024 - zuul1005
ganeti1025 - zuul1006
ganeti1026 - zuul1007

but I also see zuul1008 and zuul1009? Could you please provide more clarification on this?

Physicalled relabeled zuul1004-1007. Awaiting clarification before proceeding.

but I also see zuul1008 and zuul1009? Could you please provide more clarification on this?

Hi @VRiley-WMF

I searched the puppet repo, netbox and Phabricator for zuul1008/zuul1009 but it seems like it is only mentioned in your comment above.

Where do you see those?

Oh, is it because I used node /^zuul([1-2]00[4-9])\.(codfw|eqiad)\./ in site.pp? I just did that to be future proof. I could have also ended that range at 7.

I think all is good with 1004 through 1007 as you listed them above.

LSobanski triaged this task as Medium priority.Jun 29 2026, 9:32 AM

Hi @VRiley-WMF let me know if there are still questions around this.

Change #1310656 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] site: add zuul1004/2004

https://gerrit.wikimedia.org/r/1310656

Change #1310667 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] site: limit regex for zuul physical machines to 4-7 range

https://gerrit.wikimedia.org/r/1310667

Change #1310667 merged by Dzahn:

[operations/puppet@production] site: limit regex for zuul physical machines to 4-7 range

https://gerrit.wikimedia.org/r/1310667

I made sure also in site.pp there is only 4 through 7 - no more 8 and 9 appearing there.

Ready to setup zuul1004/2004 on our end. Not particularly urgent but just for planning purposes and because I am going on vacation - could we expect these to be handed over this week? If not that's not a big deal but I will include it in a hand-over.

Hey @Dzahn thanks for the ping on this. I've recently been freed up to work on this a bit more as I was having a lot of difficulty with these servers at first.

Thanks @VRiley-WMF it's alright either way. It was just about who to assign it to on our end.

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1004.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1004.eqiad.wmnet with OS bullseye

Hey @MoritzMuehlenhoff is there a recommended OS for these devices? I've tried Bullseye and bookwork, but that didn't work.

Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1004.eqiad.wmnet with OS bullseye executed with errors:

  • zuul1004 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console zuul1004.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Mentioned in SAL (#wikimedia-operations) [2026-07-28T00:26:25Z] <mutante> attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell (T427353)

I tried to run reimage as well, with trixie, and then connected to the install-console.

Just getting BusyBox shell.

Something not right yet in netbox or special config because this used to be ganeti? Looking some more into it.

Change #1318354 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] installserver: switch zuul[12]004 to raid10-4dev partman recipe

https://gerrit.wikimedia.org/r/1318354

Change #1318354 merged by Dzahn:

[operations/puppet@production] installserver: switch zuul[12]004 to raid10-4dev partman recipe

https://gerrit.wikimedia.org/r/1318354

Change #1318362 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] installserver: switch zuul[12]004 to -efi partman recipes

https://gerrit.wikimedia.org/r/1318362

Change #1318362 merged by Dzahn:

[operations/puppet@production] installserver: switch zuul[12]004 to -efi partman recipes

https://gerrit.wikimedia.org/r/1318362

@VRiley-WMF Looks like I got it working after using different partman recipes. That was on me! I will take care of the second server tomorrow. Don't worry about it for now.

@VRiley-WMF So.. zuul1004 is done now. works for me:)

I was going to continue with zuul2004 but realized that does not exist yet in netbox and codfw is probably not you handling it. (Should I ping dcops-codfw separately about that?)

So I am doing zuul1005 and counting up next.

Oh, maybe codfw is not involved at all and I just assumed too much. Are all hosts we are repurposing eqiad and nothing in codfw?

@VRiley-WMF While zuul1004 works fine I got this when trying to do zuul1005 next:

spicerack.netbox.NetboxError: Server zuul1005 does not have any primary IP with a DNS name set.

Any other steps on your side that had already been done for 1004 but not the rest yet?

I'm currently working on those as we speak

Oh cool! Stepping back for now and checking in later. Thank you!

So, zuul1004-1007 would be in eqiad. If it starts with a 2, it would be for codfw

ACK! aware of that. Was just thinking the hardware we are repurposing might have existed in both DCs.

Currently, if you wanted to have zuul2004-7 we would need to find servers that are located in codfw. This would be on another ticket or a sub ticket for those devices. I apologize, I thought there was already a ticket for those devices.

Yea, that makes sense. Either way that would be another ticket.

@MoritzMuehlenhoff @thcipriani Has this always been an eqiad-only thing or is it both eqiad and codfw? Asking because I was not originally involved in the talks about using physical hardware for zuul. The VM-based version of zuul I setup in both DCs of course.

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1005.eqiad.wmnet with OS trixie

@Dzahn would you be able to take a look at Zuul1005? It looks like it was able to finish the installer but it looks like it's going to fail the re-image script

Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1005.eqiad.wmnet with OS trixie completed:

  • zuul1005 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607281808_vriley_3337297_zuul1005.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Oh, I could be wrong. It looks like it just finished

@VRiley-WMF Looks good to me :) I can ssh to zuul1005 and seems done. Also, thanks for using trixie.

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixie executed with errors:

  • zuul1006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console zuul1006.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixie

Change #1310656 merged by Dzahn:

[operations/puppet@production] site: add zuul1004

https://gerrit.wikimedia.org/r/1310656

Hey @Dzahn I may need some help with zuul1006. I believe all these servers will be using a 1 gig cable, but this unit is in a rack where they use fiber. I'm currently looking to see how to fix that. However, zuul1004 and zuul1005 are up

@VRiley-WMF Ok, thanks for the update. Take your time. I will be busy setting up zuul1004 at first regardless. It's ok if you prioritize zuul1006 a bit lower. Thanks for 1004 and 1005 :)

Change #1319528 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] zuul: add zuul1004 to zookeeper and list of main hosts

https://gerrit.wikimedia.org/r/1319528

Change #1319528 merged by Dzahn:

[operations/puppet@production] zuul: add zuul1004 to zookeeper and list of main hosts

https://gerrit.wikimedia.org/r/1319528

Change #1319547 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] site: add zuul1005 as a zuul executor

https://gerrit.wikimedia.org/r/1319547

Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1007.eqiad.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1007.eqiad.wmnet with OS trixie completed:

  • zuul1007 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202608051912_vriley_1087923_zuul1007.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully