Page MenuHomePhabricator

db1246 crashed again
Closed, ResolvedPublic

Description

Same occurrence as T359940 T361968 T363119 T374215 T387673:

Record:      6
Date/Time:   04/08/2025 15:38:46
Source:      system
Severity:    Critical
Description: The system board BP1 PG voltage is outside of range.
-------------------------------------------------------------------------------
Record:      7
Date/Time:   04/08/2025 15:39:01
Source:      system
Severity:    Ok
Description: The system board BP1 PG voltage is within range.
-------------------------------------------------------------------------------

@VRiley-WMF @wiki_willy same host same crash. I guess we have to talk to Dell again, this is really not acceptable. We keep having the same crash with the same logs over and over again.

Event Timeline

Marostegui triaged this task as Medium priority.Apr 8 2025, 3:56 PM

Change #1135070 had a related patch set uploaded (by Marostegui; author: Marostegui):

[operations/puppet@production] db1246: Disable notifications

https://gerrit.wikimedia.org/r/1135070

Change #1135070 merged by Marostegui:

[operations/puppet@production] db1246: Disable notifications

https://gerrit.wikimedia.org/r/1135070

File system is corrupted so it was a hard crash (presumably storage?):

[ 1261.563104] XFS (dm-0): Metadata corruption detected at xfs_agi_verify+0x11a/0x170 [xfs], xfs_agi block 0x1e3a5e02
[ 1261.573758] XFS (dm-0): Unmount and run xfs_repair
[ 1261.578560] XFS (dm-0): First 128 bytes of corrupted metadata buffer:
[ 1261.585010] 00000000: 58 41 47 49 00 00 00 01 00 00 00 01 03 c7 4b c0  XAGI..........K.
[ 1261.593018] 00000010: 00 00 00 80 00 00 00 03 00 00 00 01 00 00 00 04  ................
[ 1261.601021] 00000020: 01 39 b1 c0 ff ff ff ff ff ff ff ff ff ff ff ff  .9..............
[ 1261.609028] 00000030: ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff  ................
[ 1261.617031] 00000040: ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff  ................
[ 1261.625041] 00000050: 00 00 00 8a ff ff ff ff ff ff ff ff ff ff ff ff  ................
[ 1261.633044] 00000060: ff ff ff ff 00 00 00 8f ff ff ff ff 00 00 00 91  ................
[ 1261.641045] 00000070: ff ff ff ff ff ff ff ff 00 00 00 94 ff ff ff ff  ................

The host won't boot up, but that's ok, we can reimage it once it is troubleshooted by Dell.

Change #1135158 had a related patch set uploaded (by Marostegui; author: Marostegui):

[operations/puppet@production] installserver: Add db1246

https://gerrit.wikimedia.org/r/1135158

Change #1135158 merged by Marostegui:

[operations/puppet@production] installserver: Add db1246

https://gerrit.wikimedia.org/r/1135158

Marostegui renamed this task from db1246 went down to db1246 crashed again.Apr 10 2025, 7:31 AM

Collecting a report on this and will update when I have a ticket with Dell.

Thank you - you can reboot and power off the host as much as you need, it's not accessible, data is corrupted and it's out of production

Dell work order number is 208331710. Currently this is under investigation

After talking to Dell about this ticket, they are escalating this toa higher tier of support. Will update when I hear back from them. Hopefully we can get this resolved once and for all.

For reference, here are the previous tickets that have been made for this specific unit

188297490 - April 5th 2024
197398410 - September 10th 2024
198075128 - September 23rd 2024
200579927 - November 7th 2024
206617456 - March 7th 2025
208331710 - April 10th 2025 (Current one)

Dell is currently with their level 3 engineers and looking at this ticket. They have laid out this plan of action on this server

"Plan of Action

Apply the latest iDRAC firmware update iDRAC 7.20.10.50

Once that has completed, pull a new SupportAssist collection with Debug Logs to provide to our level 3 engineers for review. "

@Papaul Would you happen to know if this is okay to apply to this server? I don't believe that we currently have any other servers with this level of iDRAC. Please advise. Thank you!

Thanks @VRiley-WMF - hopefully the plan is not to upgrade to that latest firmware and then wait again a few months to see exactly the same crash. Can you double check that their are not planning to make us upgrade and wait again? Because it's been more than one year since we are dealing with this server crashing exactly the same way all the time. And each crash is impacting live users.

Understood, I will be relaying this information to Dell to inquire if there are additional plans of action. As, I do know we have similar servers with similar configuration (if not the exact same) that have been running completely fine. Thank you for this input and I will give an update when I hear back from them.

@VRiley-WMF yes it is OK to apply 7.20 to the server. My personally opinion I don't think applying this latest IDRAC upgrade to the server will provide us with any information then what we had already. The main cause of this error we are getting is the main board or we have already replaced the main board and we are still seeing the same error. I might think that we received maybe another bad board from them but that is just a possibility. the second reason might be the CPU"S the third reason the power supplies and the last reason the server environment. so it is very difficult to tell .

Understood. I will be reaching out to them again to see if we can request that plan of action that you've recommended. I can ask them about the mainboard to see if they would replace it, but also, I'll ask about the CPU's and power supply. Hopefully they may be willing to replace all of those. I will update when I have more information

After working with Dell a bit more on this, I pushed back on their request regarding the iDRAC. They initially wanted to check if the newer firmware would collect more in-depth logs. During the conversation, they also suggested swapping the CPUs. However, given this server’s history, I pointed out that we’ve done similar steps before; Dell would say “everything looks clean,” and then a month or two later, we’d be right back here opening another ticket.

After I shared that concern, they agreed and instead requested another TSR with debug logs. I’ve gathered those and submitted them through the ticket. Now we’re waiting on a response from their Level 3 technician.

I have sent an email to them requesting an update on this. Awaiting response.

I received this as a response today

"After reviewing the debug logs and thermal data, we did not uncover any new information. It appears that the issue is self-correcting until it reoccurs, so gathering data at the time of the event would be ideal for a better understanding.

If the issue persists or we are unable to isolate it further, we recommend replacing the following parts:

Backplane
System board
Backplane power cabling
fPERC (if applicable)
PIB/PDB (if applicable)

Interestingly, the TSR logs from 4/5/2024 and 4/14/2025 show that the processors and DIMM A1 have remained in the same slot throughout testing. While there is no immediate indication of voltage errors from these components, it might be beneficial to swap the processors and DIMMs A1 & B5 if you are willing to test that first.

If swapping the processors and DIMMs is not an option, we can proceed with sending out the recommended parts. It is important to replace all the parts at once to avoid mixing old and new components, which could complicate tracing the voltage issues.

Please let us know how you would like to proceed. We are here to support you in resolving this issue."

I have requested someone to come onsite to replace the physical parts. They may replace some or all parts. Still waiting to hear back from them.

I got an update from Dell, and they said they would be replacing some of the parts on this. They have listed the Mainboard, Cables and power supplies as being possible issues and looking to resolve this issue once and for all. WE have them on schedual for coming onsite on the 28th.

Thank you!! I hope this will fix the issue for good!

Dell was onsite today and replaced the motherboard, moved DIMMs around, replaced cables and replaced a CPU. Heres hoping we can finally close this ticket, once and for all!

Cookbook cookbooks.sre.hosts.reimage was started by marostegui@cumin1002 for host db1246.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by marostegui@cumin1002 for host db1246.eqiad.wmnet with OS bookworm completed:

  • db1246 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202504290625_marostegui_4100326_db1246.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is not optimal, downtime not removed
    • Updated Netbox data from PuppetDB

I've reimaged the host, I had to reset the idrac password.

Ladsgroup reopened this task as Open.EditedMay 3 2025, 8:11 PM
Ladsgroup subscribed.

It paged again. This is the seventh time by my count. https://xkcd.com/2083/

Let's keep a task per crash, so I'm closing this and opened T393296