Page MenuHomePhabricator

db1245 crashed
Open, MediumPublic

Description

-------------------------------------------------------------------------------
Record:      45
Date/Time:   07/02/2026 21:07:24
Source:      system
Severity:    Critical
Description: CPU 1 MEMEFGH VPP PG voltage is outside of range.
-------------------------------------------------------------------------------
Record:      46
Date/Time:   07/02/2026 21:07:39
Source:      system
Severity:    Ok
Description: CPU 1 MEMEFGH VPP PG voltage is within range.
-------------------------------------------------------------------------------
Record:      47
Date/Time:   07/02/2026 21:07:59
Source:      system
Severity:    Critical
Description: CPU 1 MEMEFGH VPP PG voltage is outside of range.
-------------------------------------------------------------------------------
Record:      48
Date/Time:   07/02/2026 21:08:07
Source:      system
Severity:    Critical
Description: The system board Pfault fail-safe voltage is outside of range.
-------------------------------------------------------------------------------
Record:      49
Date/Time:   07/02/2026 21:27:23
Source:      system
Severity:    Non-Critical
Description: The CMOS battery is not ready either because the battery is not fully charged or a transient battery event is observed.
-------------------------------------------------------------------------------

CEST time: 23:09:39 <+icinga-wm> PROBLEM - Host db1245 is DOWN: PING CRITICAL - Packet loss = 100%

Event Timeline

jcrespo subscribed.

@dcops , I tried hard-resetting the host remotely, but it doesn't boot up, it shows multiple power issues to cpu, other board locations. Please see if you can at least drain power and be able to make it boot it up, but either there is a power issue, or the board may be fried. Please advice.

Change #1307436 had a related patch set uploaded (by Jcrespo; author: Jcrespo):

[operations/puppet@production] mariadb & mediabackups: Replace db1245 during ongoing hw issues

https://gerrit.wikimedia.org/r/1307436

Change #1307436 merged by Jcrespo:

[operations/puppet@production] mariadb & mediabackups: Replace db1245 during ongoing hw issues

https://gerrit.wikimedia.org/r/1307436

I forgot to mention, I consider data here lost/corrupted, the host was downtimed, feel free to do any operation with the server without asking.

I have opened up a ticket for this unit. It's good it failed now, as the service ends on this unit later this month.

Dell SR ticket number 228686846

The TSR report showed the following and Dell has provided some troubleshooting tips which I will be walking through

Post TSR analysis we can see multiple errors reported on system as listed below.

1.The system board Pfault fail-safe voltage is outside of range.

2.CPU 1 MEMEFGH VPP PG voltage is outside of range.

Please confirm if there were any recent changes/updates/power surge in the data center recently.

Please confirm if the server is able to POST(Power on self test)

If server is able to POST please confirm if it is loading into OS or not.

I would also like to request the following details:

  1. Can you confirm if the fans are spinning?
  2. Is there any UPS present in between?
  3. Please help me with the LED indications on the following components. a. Hard Drive: b. NIC Ports: c. Power Supplies: d. iDRAC Direct: e. System ID Button(Back) f. System ID Button(Front)

    Performing a power drain.
  1. Power down the server.
  2. Disconnect all power and network cables from the server.
  3. Hold down the power button for at least 35-40 seconds.

3(a).Clear NVRAM

Refer the service manual - https://dl.dell.com/content/manual43348652-dell-emc-poweredge-r650xs-installation-and-service-manual.pdf?language=en-us

  1. Plug the power and network cables back into the system.
  2. Wait approximately 2 minutes for iDRAC to initialize before turning on the server.
  3. Turn on the system.

Please help us with the requested details/ troubleshooting. I will wait for your further confirmation. Please help me with an approximate turn around time, so that I can follow up accordingly.

Worked with the server and here is what I have responded to Dell with

Here is the information you requested

Please confirm if there were any recent changes/updates/power surge in the data center recently. - No changes

Please confirm if the server is able to POST(Power on self test) - No, it doesn't POST

  1. Can you confirm if the fans are spinning? - Yes
  2. Is there any UPS present in between? - No
  3. Please help me with the LED indications on the following components. a. Hard Drive: No LED indications b. NIC Ports: Yes, amber and green c. Power Supplies: Yes, Green d. iDRAC Direct: Yes, Amber and green e. System ID Button(Back) - Amber f. System ID Button(Front) - Amber

Performing a power drain.

  1. Power down the server. - Done
  2. Disconnect all power and network cables from the server. - Done
  3. Hold down the power button for at least 35-40 seconds. - Done

3(a).Clear NVRAM - Cleared by moving the jumper as requested

  1. Plug the power and network cables back into the system. - Done
  2. Wait approximately 2 minutes for iDRAC to initialize before turning on the server. - Done
  3. Turn on the system. - Done

Awaiting their reply.

They recommened updating the firmware on the BIOS, but that's where I'm currently getting stuck at now, since it isn't powering back on. Reached back out to Dell to see if they have any further recommendations on this.

@jcrespo can we change the status of this server in netbox to failed? netbox reporting is complaining about it but it's also not critical in any way.

Of course, doing it, I didn't know this was going to take so long, I thought at first it was just a normal crash & restart.

After attemtping to update the firmware several times, it doesn't seem to be taking it. Awaiting on dell for further instructions.

They are sending an onsite tech to probably replace the mainboard. Will keep this updated.

The tech is schedualed to come onsite on Thursday July 23th.

Dell tech came onsite to check it out and they concluded to replace the mainboard on the server. They should be back tomorrow or monday depending on part availability

@VRiley-WMF Thank you for the updates. This helps me be aware of the changes and plan to be ready for software/os setup ASAP when it is back up. Thank you!

Dell called me today with an update. They are still having trouble obtaining the mainboard. However, they did assure me that a tech will be onsite Wednesday the 29th in order to replace it. Will update then.

VRiley-WMF changed the task status from Open to In Progress.Jul 29 2026, 2:13 PM

Dell is onsite and proceeding with the mainboard replacment.

One request, @VRiley-WMF if all goes well and it goes back to "normal" (bootable state), could you leave the host shutted down for the night (just powerable through remote management?). I will reimage it afterwards, but because it has skipped regular package upgrades, I'd prefer to not leave it powered on for several hours until I can service it. It is not super important, but I'd prefer if you could leave it down after any basic checks (e.g. that it posts or boots normally).

Of course @jcrespo I will make sure I do that. Thanks!

Dell has come on site and replaced the mainboard. I have logged the new MAC addresses (if needed) into Netbox. I am collecting a TSR report for Dell and then will power this unit off.

System is powered off, but it is reachable. Should be good to go. @jcrespo should we be okay to close this ticket?

VRiley-WMF changed the task status from In Progress to Open.Jul 29 2026, 3:47 PM
jcrespo triaged this task as Medium priority.

I've marked it as active on netbox and will be reimaging it to put it into service again.

Reopening because reimaging it failed:

Running IPMI command: ipmitool -I lanplus -H db1245.mgmt.eqiad.wmnet -U root -E chassis power status
Error: Unable to establish IPMI v2 / RMCP+ session
Exception raised while initializing the Cookbook sre.hosts.reimage:
Traceback (most recent call last):
  File "/usr/lib/python3/dist-packages/spicerack/ipmi.py", line 90, in command
    output = run(command + command_parts, env=self.env.copy(), stdout=PIPE, check=True).stdout.decode()
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.11/subprocess.py", line 571, in run
    raise CalledProcessError(retcode, process.args,
subprocess.CalledProcessError: Command '['ipmitool', '-I', 'lanplus', '-H', 'db1245.mgmt.eqiad.wmnet', '-U', 'root', '-E', 'chassis', 'power', 'status']' returned non-zero exit status 1.

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/usr/lib/python3/dist-packages/spicerack/_menu.py", line 205, in run
    runner = self.instance.get_runner(args)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/srv/deployment/spicerack/cookbooks/sre/hosts/reimage.py", line 121, in get_runner
    return ReimageRunner(args, self.spicerack)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/srv/deployment/spicerack/cookbooks/sre/hosts/reimage.py", line 248, in __init__
    self._validate()
  File "/srv/deployment/spicerack/cookbooks/sre/hosts/reimage.py", line 358, in _validate
    self.ipmi.check_connection()  # Will raise if unable to connect
    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3/dist-packages/spicerack/ipmi.py", line 105, in check_connection
    self.power_status()
  File "/usr/lib/python3/dist-packages/spicerack/ipmi.py", line 115, in power_status
    status = self.command(["chassis", "power", "status"], is_safe=True)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3/dist-packages/spicerack/ipmi.py", line 92, in command
    raise IpmiError(f"Remote IPMI for {self._target} failed (exit={e.returncode}): {e.output}") from e
spicerack.ipmi.IpmiError: Remote IPMI for db1245.mgmt.eqiad.wmnet failed (exit=1): b''

My guess is the host may be missing common configuration, such as ipmi enabling or automation (e.g. mgmt password) that was missed during its down time, but I have not debugged further, but it was a common thing that was skipped when the board was changed in the past, fully setting options back to workable.

While I was debugging, I lost network to the host AND management interface, none are reachable right now.

Hey @jcrespo,

I was able to log in via iDRAC and power it on. It was seemingly seeing an issue with tempature, which I may need to address. However, could you please test that out?

I am unable to access now through ssh or https, with both the old or new accounts. Could you check that the passwords have not been resetted to factory defaults, as well as IPMI is enabled, and other automations typical of a new host setup? I am blocked on reimaging on the IPMI access, and I cannot do it myself without remote access.

I've also lost ping again with the main OS host, it was only up for some minutes before losing connectivity.

For context, this is the second time it crashed for me- I was able to power it on normally as you left it, but crashed before I could reimage or debug the issues (they may not be real ones, as it could be he host crashing, as you said, due to temperature). And after you put it up again, I was able to log in but only for a few minutes before it crashed again.

I was given an action plan and will be following through with these steps. Reseat all fans, reseat the heatsink and cpu. update bios firmware. Firmware is now out of date due to the replacment.

Hey @jcrespo I updated the firmware and reseated some of the fans. Would you be able to try to test putting a load on it? I was looking for thermal paste, but can't seem to find any. If that doesn't fix it, we can put in an order for thermal paste.

Thank you @VRiley-WMF I am booting the server- The server should be able to be up and not crashing without thermal paste or with it "crusty", even if a bit over the norm, but I will keep an eye on reported temperature. Independently of that, could we buy some for this or another server? It shouldn't be an expensive item IMHO (your time is way more valuable!).

But I will report first what I find first.

Mentioned in SAL (#wikimedia-operations) [2026-08-07T09:15:46Z] <jynus> started stress testing db1245 dbs T431115

I did some unrealistic scenario, which is maximizing all cpu cores while also running a mysql benchmark, and temperature stabilized around 83 degrees Celsius, with no crashes:

image.png (2,480×392 px, 162 KB)

Screenshot_20260807_113150.png (1,230×1,158 px, 171 KB)

Screenshot_20260807_113158.png (1,230×1,158 px, 173 KB)

Change #1322690 had a related patch set uploaded (by Jcrespo; author: Jcrespo):

[operations/puppet@production] installserver: Set db1245 as full uefi reimage, remove db1265/85

https://gerrit.wikimedia.org/r/1322690

Change #1322690 merged by Jcrespo:

[operations/puppet@production] installserver: Set db1245 as full uefi reimage, remove db1265/85

https://gerrit.wikimedia.org/r/1322690

Sadly, it crashed again at 9:51 UTC, and it didn't seem it had anything to do with temperature, as it did after the stress test finalized:

image.png (834×916 px, 91 KB)
And when that happens I lose access to both the main OS and the out of band interface, so I don't think this is a simple CPU overheating. :-(

This is very strange. I just checked out this unit and it seems like it's up and active, however the iDRAC is now lighting up at all but nothing was moved or changed. I'm looking into this @jcrespo

@jcrespo Okay, so I just took a look at it, and there was a overheating incident with CPU 2 that the server was complaining about. I was able to find some thermal paste and decided to check into it. Upon opening and taking out CPU 2, I noticed that the heatsink/processor was very loose. I applied the thermal paste and set it back in. I went ahead an did the same to CPU 1 as well.

Upon booting it up, I noticed it came online a lot quicker than it did previously. Whenever I would reboot it, it would take at least a few minutes. After reseating and applying the thermal paste it seems extremely snappy. For good measure, I did give it a *very* stern talking to. Could you please test this out and see if it shuts down again?

@VRiley-WMF I am unable to connect to either the ssh or https management point of this server. I can do it with, e.g. unrelated db1244 with no issue. I don't know if it is because it crashed again or there is something missing, but I am getting timeouts to both endpoints. The host is down and I need that to be able to boot it.

Edit: based on your comment timestamp, I think it crashed ~3 hours later:

Host Down[2026-08-12 01:06:21] HOST ALERT: db1245;DOWN;HARD;2;PING CRITICAL - Packet loss = 100%
Host Down[2026-08-12 01:05:59] HOST ALERT: db1245;DOWN;SOFT;1;PING CRITICAL - Packet loss = 100%

There is no rush on it being stable- I will be on vacations for a month and it is currently not being in use. But it would be nice to try to save it back up, as backup service currently (with ongoing hardware failures) has very little to no redundancy, meaning that if more hosts fail, we may actually lose some some parts of the database recovery service.

@jcrespo Thanks, I do apologize for the delay. I assure you this is still on my radar and I'm trying to find an answer for this as soon as possible.

@jcrespo I was able to look further into this and it does seem that ipmi is enabled, and it does seem the passwords are there as well. I could try to reprovision this server? However, it does seem like all the information is there.

@VRiley-WMF Jaime is out of office for a few more weeks. What does reprovision implies here?

@Marostegui essentially, it's making sure that everything is set the way it's supposed to for booting up properly and reliably. We normally run the scripts for fresh installs. In some cases we do have to re-run them for tickets like this one

https://phabricator.wikimedia.org/T401441

I think it is fine to proceed @VRiley-WMF - thanks!

Rebooted the device, commenced a flea power drain and it seems to have come up healthy before oving forward with the provisioning. It seems to have come back up and should be ready to be put back into service. Could you check it out @Marostegui ?

@VRiley-WMF this host is unreachable:

[08:30:46] marostegui@cumin1003:~$ ping db1245
PING db1245.eqiad.wmnet (10.64.32.59) 56(84) bytes of data.

I've cloned and started replication on db1245 for both s4 and s5. Let's see if it crashes. I will reopen if that happens. If not, putting the back in service is being tracked at T437563

Leaving open until we hear back from Dell.