Page MenuHomePhabricator

wikikube-ctrl2006 implementation tracking
Open, HighPublic

Description

This task is to track the service implementation of ServiceOps host(s) listed in the task description.

Once the linked racking task has been resolved, this task can be implemented.

This sub-task creation/update is per the request of ServiceOps; this task is assigned at creation to the 'Sub-team Technical Contact' provided in the initial ordering task.

Follow https://wikitech.wikimedia.org/wiki/Kubernetes/Clusters/Add_or_remove_control-planes#Add_stacked_control-plane then decom wikikube-ctrl2003

For decom please test-cookbook this change if it is still not merged (and report back on the CR), thanks!

Event Timeline

Change #1249321 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/puppet@production] wikikube: add wikikube-ctrl2006

https://gerrit.wikimedia.org/r/1249321

Change #1249423 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/dns@master] wmnet: add wikikube-ctrl2006 to etcd-server SRV record

https://gerrit.wikimedia.org/r/1249423

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

For visibility, I've intentionally aborted the reimage due to the following [0] which I will investigate & re-attempt.

==> Unable to verify that the host is inside the Debian installer, please verify manually with: sudo install-console wikikube-ctrl2006.codfw.wmnet

Cookbook cookbooks.sre.debmonitor.remove-hosts run by jmm: for 1 hosts: wikikube-ctrl2006.codfw.wmnet

Change #1285465 had a related patch set uploaded (by Jasmine; author: Jasmine):

[operations/dns@master] Add Kubernetes POD IP reverse range delegations for wikikube-ctrl2006

https://gerrit.wikimedia.org/r/1285465

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

When I try to run the reimage, it blows past PXE on the host without trying network boot and moves onto the disk bootloader.

@Jhancock.wm:

got the pxe issue fixed. but found a new one. @Clement_Goubert this server has to be uefi and it looks like the preseed is set up for bios. if i'm reading it right. could you update that for us?

We're having an issue here trying to boot this and it is entirely skipping PXE. Is this the same issue you had on the intial imaging?

Basically the reimage script reboots the host, but then rather than PXE it goes to the bootloader and loads the OS.

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

@Jhancock.wm:

got the pxe issue fixed. but found a new one. @Clement_Goubert this server has to be uefi and it looks like the preseed is set up for bios. if i'm reading it right. could you update that for us?

We're having an issue here trying to boot this and it is entirely skipping PXE. Is this the same issue you had on the intial imaging?

Basically the reimage script reboots the host, but then rather than PXE it goes to the bootloader and loads the OS.

Thanks RobH!

For additional visibility, we also attempted to dual re provision with then without legacy considering the previous bug [1], although PXE issue still withstanding.

[1] - https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1262196

Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Attempting to reimage and watching the serial console output, it showed no media present for PXE boot check before failing to the debian os already installed.

So PXE is betting flagged to the wrong NIC port or the cable isn't plugged into the proper port on the NIC (not this, info below).

I wanted to rule out the wrong NIC port being plugged in so I polled the MAC attached to the port via the switch and also confirmed the link was up:

robh@lsw1-d8-codfw> show interfaces descriptions 
Interface       Admin Link Description
xe-0/0/41       up    up   wikikube-ctrl2006

robh@lsw1-d8-codfw> show ethernet-switching table interface xe-0/0/41 

MAC database for interface xe-0/0/41

MAC database for interface xe-0/0/41.0

MAC flags (S - static MAC, D - dynamic MAC, L - locally learned, P - Persistent static
           SE - statistics enabled, NM - non configured MAC, R - remote PE MAC, O - ovsdb MAC)


Ethernet switching table : 1 entries, 1 learned
Routing instance : default-switch
   Vlan                MAC                 MAC      Logical                SVLBNH/      Active
   name                address             flags    interface              VENH Index   source
   private1-d8-codfw   90:5a:08:9e:74:2c   D        xe-0/0/41.0

I polled the ilom to determine the NIC port 1 (of 2) MAC info is: 90:5A:08:9E:74:2C

So this is an issue where the PXE flag isn't being set correctly by the provisioning script. It appears it sets it for the wrong port entirely on the NIC.

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS bookworm executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Steps taken:

  • Confirmed MAC of port 1 (of 2) is indeed connected to switch
  • successfully ran sudo cookbook sre.hosts.provision --no-dhcp --no-users wikikube-ctrl2006
  • ran reimage script with trixie option, during the PXE load I got some interesting quick screen shots of the failure:

Screenshot 2026-06-16 at 14.46.13.png (1,204×1,050 px, 460 KB)

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS bookworm

bookworm loader comes up fine, so this is a distro specific issue with trixie failing to pxe tftp load.

Cookbook cookbooks.sre.hosts.reimage started by robh@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS bookworm executed with errors:

  • wikikube-ctrl2006 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console wikikube-ctrl2006.codfw.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by pt1979@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by pt1979@cumin2002 for host wikikube-ctrl2006.codfw.wmnet with OS trixie completed:

  • wikikube-ctrl2006 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202606162344_pt1979_3729056_wikikube-ctrl2006.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

To recap the IRC discussion from earlier, it appears the PXE was enabled on both 10g nic's and disabling one appears to resolve the issue (thanks for doing so @Paputx ! Much appreciated!).
Following that, we were able to reimage the host successfully at least once.

However, while attempting to reimage once more for good measure before putting the host into production, running into a 401 unauthorized so pasting the response payload below while we investigate:

`Response payload: {'error': {'code': 'Base.1.10.3.GeneralError', 'message': 'A general error has occurred. See ExtendedInfo for more information.', '@Message.ExtendedInfo': [{'MessageId': 'Base.1.10.ResourceAtUriUnauthorized', 'Severity': 'Critical', 'Resolution': 'Ensure that the appropriate access is provided for the service in order for it to access the URI.', 'Message': "While accessing the resource at '/redfish/v1/Systems/1/Bios', the service received an authorization error 'unauthorized'.", 'MessageArgs': ['/redfish/v1/Systems/1/Bios', 'unauthorized'], 'RelatedProperties': ['unauthorized']}]}}`

Pinging here since this appears stale for almost a month. Is there anything we can do to get this sorted?

@Jhancock.wm,

This host was racked by you, but is currently reporting issues with reimage. Are you able to take a look at this?

To recap the IRC discussion from earlier, it appears the PXE was enable on both 10g nic's and disabling one appears to resolve the issue (thanks for doing so @Paputx ! Much appreciated!).
Following that, we were able to reimage the host successfully at least once.

However, while attempting to reimage once more for good measure before putting the host into production, running into a 401 unauthorized so pasting the response payload below while we investigate:

`Response payload: {'error': {'code': 'Base.1.10.3.GeneralError', 'message': 'A general error has occurred. See ExtendedInfo for more information.', '@Message.ExtendedInfo': [{'MessageId': 'Base.1.10.ResourceAtUriUnauthorized', 'Severity': 'Critical', 'Resolution': 'Ensure that the appropriate access is provided for the service in order for it to access the URI.', 'Message': "While accessing the resource at '/redfish/v1/Systems/1/Bios', the service received an authorization error 'unauthorized'.", 'MessageArgs': ['/redfish/v1/Systems/1/Bios', 'unauthorized'], 'RelatedProperties': ['unauthorized']}]}}`

@jasmine_ is it alright if we change the status of this server to failed and down time it? I can't seem to remote in and i can't fully re-provision it with the status as active. push comes to shove i can go physically plug a console into it, but that woud require a reboot anyway.

@jasmine_ is it alright if we change the status of this server to failed and down time it? I can't seem to remote in and i can't fully re-provision it with the status as active. push comes to shove i can go physically plug a console into it, but that woud require a reboot anyway.

Yes, that's totally fine as this server is yet to be commissioned. So absolutely, please do, thank you!

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin1003 for host wikikube-ctrl2006.codfw.wmnet with OS trixie