Page MenuHomePhabricator

hw troubleshooting: Comm Error: Backplane 0 on an-worker1172.eqiad.wmnet
Closed, ResolvedPublic

Assigned To
Authored By
RKemper
Mar 17 2026, 9:51 PM
Referenced Files
F73241550: image.png
Mar 20 2026, 9:34 AM
F73036246: Screenshot 2026-03-18 at 2.35.47 PM.png
Mar 18 2026, 6:36 PM
F72974315: image.png
Mar 17 2026, 11:50 PM
F72974098: image.png
Mar 17 2026, 11:50 PM
F72973897: image.png
Mar 17 2026, 11:50 PM

Description

FQDN: an-worker1172.eqiad.wmnet

Netbox: Set to Failed (https://netbox.wikimedia.org/dcim/devices/5037/)

Priority: Medium

Host can be worked on at will (it's already down)

Timeline
  • 2026-03-03 11:13 UTC — btullis ran sre.hosts.reboot-single (START)
  • 2026-03-03 12:08 UTC — cookbook returned END (PASS) after ~55 minutes, meaning the host came back up far enough to satisfy health checks
  • 2026-03-06 (or earlier) — host found SSH-unreachable; given the below backplane error, it prob very quickly fell back into failed state after the 2026-03-03 reboot
Record:      26
Date/Time:   03/03/2026 11:17:38
Source:      system
Severity:    Critical
Description: The System Configuration Check operation resulted in the following issue: Comm Error: Backplane 0.

Event Timeline

RKemper renamed this task from an-worker1172.eqiad.wmnet unreachable since 2026-03-03 to hw troubleshooting: Comm Error: Backplane 0 on an-worker1172.eqiad.wmnet.Mar 17 2026, 10:02 PM
RKemper reassigned this task from RKemper to Jclark-ctr.
RKemper added projects: DC-Ops, ops-eqiad.
RKemper updated the task description. (Show Details)
RKemper moved this task from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.

Switched this to a HW failure ticket, given racadm getsel revealed a backplane issue

RKemper triaged this task as Medium priority.Mar 17 2026, 10:06 PM
RKemper updated the task description. (Show Details)
RKemper updated the task description. (Show Details)

I confirmed that the backplane was giving this error.

image.png (2,388×1,245 px, 619 KB)

But then I tried cold resetting the BMC.

btullis@cumin1003:~$ sudo ipmitool -I lanplus -H "an-worker1172.mgmt.eqiad.wmnet" -U root -E shell
Unable to read password from environment
Password: 
ipmitool> bmc reset cold
Sent cold reset command to MC

Initially, the results looked good after reconnecting...

image.png (2,388×1,245 px, 455 KB)

But upon powering up, the error came back again.
image.png (2,388×1,245 px, 548 KB)

@BTullis
Performed firmware update on backplane seems to of cleared but is not booting might require reimage.

Screenshot 2026-03-18 at 2.35.47 PM.png (2,094×1,330 px, 243 KB)

Gehel subscribed.

Re-opening and moving to in progress to finalize reimage.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Change #1256270 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Temporarily set an-worker1172 into insetup mode

https://gerrit.wikimedia.org/r/1256270

The first reimage failed because of a partman issue.

image.png (899×536 px, 76 KB)

I'll put the host into insetup mode to carry out the reimage.

Change #1256270 merged by Btullis:

[operations/puppet@production] Temporarily set an-worker1172 into insetup mode

https://gerrit.wikimedia.org/r/1256270

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Change #1256426 had a related patch set uploaded (by Btullis; author: Btullis):

[operations/puppet@production] Update the partman recipe for an-worker1172

https://gerrit.wikimedia.org/r/1256426

Change #1256426 merged by Btullis:

[operations/puppet@production] Update the partman recipe for an-worker1172

https://gerrit.wikimedia.org/r/1256426

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by btullis@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye executed with errors:

  • an-worker1172 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console an-worker1172.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye

Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1003 for host an-worker1172.eqiad.wmnet with OS bullseye completed:

  • an-worker1172 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bullseye OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202603231359_jclark_1496331_an-worker1172.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is not optimal, downtime not removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status failed -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

@BTullis , I was able to reimage it. The an-workers always seem to have issues when the HDD virtual drives are present during imaging. Would you be able to add them back in with a single command instead of me having to do each one manually?

I belive that this is now fixed. Thanks @Jclark-ctr .

There is an outstanding puppet error, but I believe that this is unrelated to the hardware issue, so I'll close this ticket.