Page MenuHomePhabricator

cp5022 is unreachable
Open, HighPublic

Description

cp5022 went down on 2026-13-01 01:18:48 and remains unreachable on the main NIC and the management console:

vgutierrez@bast5004:~$ date ; nc -w 3 -zv cp5022.eqsin.wmnet 22 ; nc -w 3 -zv cp5022.mgmt.eqsin.wmnet 22
Tue Jan 13 09:13:31 AM UTC 2026
nc: connect to cp5022.eqsin.wmnet (2001:df2:e500:101:10:132:0:30) port 22 (tcp) failed: No route to host
nc: connect to cp5022.eqsin.wmnet (10.132.0.30) port 22 (tcp) timed out: Operation now in progress
nc: connect to cp5022.mgmt.eqsin.wmnet (10.132.128.18) port 22 (tcp) timed out: Operation now in progress

Troubleshooting

Host unresponsive to ping or ssh on both primary and idrac interfaces.

Jin attached his own bezel/LCD to confirm operation of host. It throws the error "The system board 5V SW PG voltage is outside of range." as well as a 'chassis open' error (it wasn't opened) so it appears the mainboard has gone bad. The power supplies show green leds (no errors) with only the main chassis throwing errors.

System out of warranty since October 2025. The 'self repair' options on the Dell site do not include replacement mainboards: https://www.dell.com/en-us/shop/pfydresults/273519/925F5S3

Repair / Replacement Options

Bad mainboard is difficult to swap out of warranty. PowerEdge R450 is under 5 years old, so we don't have any decommission servers of that mainboard in our core sites. The Dell parts website doesn't list core system parts like CPU/Mainboard for self order post warranty.

cp5022 had a hw failure via parent task T414411. Jin did basic troublshooting with a bezel with LCD and it reported bad 5volt to mainboard and refused to power on. A new power distro board was ordered via T417680 and a new mainboard ordered via T420367. While both of those have moved the repair along, the power failure also seems to have killed CPU1, so now the system is remotely accessible but with only 1 CPU, replacement CPU is now via child task T426985.

Details

Other Assignee
ssingh

Event Timeline

Vgutierrez triaged this task as Medium priority.Jan 13 2026, 9:14 AM

cp5022 is unresponsive to ping on its primary interface (expected with OS down) and idrac/mgmt interface (unexpected).

1-255962774671 entered should be completed by 2026-01-15 @ 13:00.

One of our servers has gone offline and we need the following done:

The server is cp5022 located in rack 604:U16, serial 925F5S3 and labeled cp5022.

  • Trace the power cables and tell us what power plugs the cables.
    • One should go to the PS1 and one to PS2 on the back, please tell us what power plug # they plug into.
  • Unplug all power from both power supplies for 30 seconds to cause a full power drain of only cp5022/925F5S3.
  • Plug back in both power cables, ensure system powers back on via power button in front.

This should reset the systems idrac and allow us to connect.

As the system is offline please do this work as soon as possible, no scheduling is required.

IBX Question:Dear Customer,We have traced both power cables, they are both connected at port 30 of PS1 and PS2.We have also unplugged and plug back both power cables as instructed.Kindly check on your end and advise if any further actions required.Thank you.

Still doesn't respond to ping on the idrac interface:

Support,

This host doesn't appear to be powering back up after the power cable removal. Would you be able to attach a crash cart and attempt to power on the system and let us know what you see?

The server is cp5022 located in rack 604:U16, serial 925F5S3 and labeled cp5022. Please attach a crash cart and diagnose why the host is not powering on.

Mentioned in SAL (#wikimedia-operations) [2026-01-18T15:51:38Z] <sukhe> downtime cp5022 as host is down: T414411

Hi @RobH: Thanks for following up on this. Any update from the eqsin folks?

RobH raised the priority of this task from Medium to High.Jan 20 2026, 4:32 PM

They've connected a crash cart and the host is hard down. Seems we have a bad mainboard or a bad PSU controller. I'm typing up directions for Jin@ DreamIIC for him to go out and really dig into this host for a Dell support request.

Chances are the bad controller board since we dont see any mgmt or power to the chassis at all.

RobH mentioned this in Unknown Object (Task).Jan 20 2026, 4:34 PM

Icinga downtime and Alertmanager silence (ID=85b11191-0733-4a6c-a314-a87c77eb102d) set by vgutierrez@cumin1003 for 10 days, 0:00:00 on 1 host(s) and their services with reason: cp5022 is unreacheable

cp5022.eqsin.wmnet

While I've contacted Jin to do this work (T415090) I'm hesitant to do so during the week of the SRE offsite. While I am attending remotely, the shift I'll have to make to attend in the virtual time zone is the exact opposite shift I'd need to make to work with Jin in Singapore.

Additionally, if you open and work in a rack, you can bump a cable. Normally this isn't as big a deal, but I don't want to possibly cause any other cascading failures in eqsin during the off hours of the SRE offsite. The potential for disruption seems greater than the benefit of bringing cp5022 back online 1 week earlier.

So I've suggested we delay Jin's work until the week following the offsite. He'll go out and diagnose the PDU ports to confirm working voltages, then move onto confirming the PDU cables to cp5022 work, then test the power supplies on the PDUS and check for LED output to the mainboard. Equinix remote hands did some of this, but in a less detailed manner (they just repeated they reseated the power cables and no LEDs on server when they press the buttons).

Hi @RobH: thanks for the update and for pursuing this. And yeah, that works; waiting for the week after the offsite is also fine since it is just one host. Thanks!

After having Jin check, this system has a failure of "The system board 5V SW PG voltage is outside of range." on the front LCD he plugged into it. The warranty expired in October 2025. I'll need to check with Willy how we can proceed.

The options Dell lists for 'self repair' off the website are not inclusive of mainboards, only user peripheals and easily swapped items: https://www.dell.com/en-us/shop/pfydresults/273519/925F5S3

RobH updated the task description. (Show Details)

Hey @RobH - did Jin say what kind of initial troubleshooting he did? Like did he do a power drain, reseat certain parts, etc? I think we can go ahead and purchase parts to see if it'll help fix this, though it'll be helpful knowing what was attempted so far. Thanks, Willy

He did full troubleshooting with photos with me in a google chat, it included the following:

  • confirming the power ports on the PDU towers were outputting power
  • confirming the power cables were fully seated
  • draining all power and reseating both PDUs
  • both PDUs show green LEDs denoting power being received from the servertech ps1/2 in the rack.
  • attempted to power on server, no response from power button, red error led.
  • attached a bezel (he owned) to the front of the host so the LCD could output error codes.
    • error for The system board 5V SW PG voltage is outside of range
    • error for Chassis having been opened (it had not been opened prior to host going offline so this appears to be due to mainboard failure as the sensor for this is on the mainboard)
    • error for cmos battery failure
  • while all the above errors show up, the host doesn't power up, and the idrac interface doesn't show any power LEDs.

It appears the system isn't getting proper power to the mainboard and thus cannot fire up any processes, which keeps the idrac offline so I cannot pull detailed failure logs.

We cannot purchase replacement parts via the Dell website I linked in, which is the Dell site linked when you try to open a case without warranty support.

The alternative is to try to get this back under warranty support (pay the fees) or go directly to our Dell Singapore team to see if they can source the parts for us. To be clear, I cannot be certain what exact part is broken due to this, I wanted to source and price both a replacement power distribution board and mainboard. I also need to confirm with Dell (or with one of our on-sites who can crack open a R450) that the power distro board is independent of the mainboard.

Sounds good @RobH, that plan works for me as well. Do you know if Jin has access to any of these parts by any chance? If he is able to get a hold of them, he could just add the cost onto our invoice.

The alternative is to try to get this back under warranty support (pay the fees) or go directly to our Dell Singapore team to see if they can source the parts for us. To be clear, I cannot be certain what exact part is broken due to this, I wanted to source and price both a replacement power distribution board and mainboard. I also need to confirm with Dell (or with one of our on-sites who can crack open a R450) that the power distro board is independent of the mainboard.

I'll ask, also going to ask in dc ops meeting if anyone has a spare r450 they can crack open to check for the part # of the power distribution board and the mainboard.

Ok, Jenn checked inside the R450 and it is indeed a stand alone power distro board.

IMG_8344.JPG (768×1,024 px, 139 KB)

IMG_8345.JPG (1,024×768 px, 84 KB)

John might have two hosts abandoned from T342455 (he is checking) and if so, we could steal the power distro board from one of them to ship to eqsin for repair of cp5022, then source a new power distro board via our Dell USA account team to fix the host in eqiad (the hosts on that order leave warranty support in August 2026.). This would be a non warranty out of pocket purchase, not submitting a USA based support case.

WMF request for Dell USA - Help determining the part # for R450 power distribution board

Dell Team,

I have an odd request, so I'll give you the background first. We have a host in Singapore (925F5S3) which left warranty in October 2025. On January 13th the host suddenly went offline without warning. We had a remote on-site attach a bezel to the front and perform some testing, and it has the error "The system board 5V SW PG voltage is outside of range" as well as a CMOS error and a board intrusion error. The two PSUs show green LEDs, and power from the utlity is good (we checked) and reseated the PDUs. It seems that either the power distribution board or the mainboard (or both) have gone bad in this out of warranty host.

When I look up options on https://www.dell.com/support/product-details/en-us/servicetag/925F5S3/overview the self repair doesn't include these parts. Is this something that I can source the part numbers from you, and then I can go to our Dell SG team and have them furnish a quote? I realize this is leveraging your expertise for an order you won't receive, but we only talk to the Dell SG team once every 3 years so figuring this out with them would add a lot of overhead.

Would you be able to provide me the sku for the power distribution board and its attaching cables, as well as the mainboard?

There also seems to be an extend option which I've submitted, but I suspect since the host is requiring repair and out of support coverage, extending support may no longer be an option?

Please advise,

They acknowledged the request on the 5th but no movement since then. I've sent a followup ping email yesterday and again today.

Ok, Summary of updates:

  • Dell attempted to get a special extension for this host, they were denied.
  • I attempted to extend the warranty on this host, was told I could not while it was in hardware failure.
  • Contacted Dell USA team, they have a few options:
  • Dell USA also was able to eventually provide me the Dell USA part numbers, GN3KY – Motherboard, 996M8 – Power Distribution Board
    • I'll request a quote from Dell SG dependent on the outcome of the support case just opened.
RobH mentioned this in Unknown Object (Task).Feb 17 2026, 5:16 PM
RobH added a subtask: Unknown Object (Task).

This host is still marked as Active in Netbox but disappeared from PuppetDB, if it's still broken please fix its status in Netbox.

Set to failed.

Ordering of replacement part will take place today or tomorrow, sorting the terms.

The order is placed and I'm currently scheduling the Unisys/Dell engineer to go onsite sometime between Friday-Wednesday of this/next week. Host is hard down, so no traffic intervention required.

Please note the maint window for this offline host is 2026-03-13 @ 07:00 AM Singapore / which is 5PM Thursday evening for me. I'll be online to remotely supervise the swap and attempt to login to the idrac when done.

Tech is running late, their dispatcher called me to let me know. They were set to be onsite at 7AM, but it will now be closer to 10:30AM / 19:30 Pacific

Tech is onsite and performing the hw power distro board swap on cp5022

The distro swap did not fix this host, it will require a mainboard swap via a procurement task (linked in)

RobH closed subtask Unknown Object (Task) as Resolved.Mar 17 2026, 2:50 PM
RobH added a parent task: Unknown Object (Task).Mar 17 2026, 2:55 PM
RobH removed a parent task: Unknown Object (Task).
RobH added a subtask: Unknown Object (Task).

Mainboard swap will occur on Wednesday, April 8th @ 10:00Singapore time which is Tuesday, Tuesday April 7th 18:00 Pacific.

I'll be online for the duration of the work and to ensure system is remotely accessible when complete for reimaging on Wednesday by Traffic.

I'll put a more detailed timeline and update tomorrow but as it stands now:

  • unisys engineer showed up at 10am singapore time
  • swapped mainboard, damaged the CPU bracket and mainboard leads to CPU1 socket in the process
  • had errors on attempted boot, determined what happened, updated the case with dell

Dell should now reach back out to me to reschedule a new visit with the repaired parts, as this was a Dell dispatch repair we're not going to be responsible for their damage to the host.

Dell has confirmed case update and will dispatch a new mainboard and cpu bracket. Once they do, they'll email/update with tracking and then dispatch will reach back out to schedule the third unisys site visit.

Hi @RobH. Any update on this from Dell's end?

I reached back out to them yesterday and I'm awaiting a reply. They were bugging us about the invoice for the mainboard they failed to install.

Update from email:

  • finally got an answer back after escalating both on the ticket, via our dell sg team, and via the accounts payable folks @ dell sg who want to be paid for the mainboard they broke.

We should see some movement, new case 226187428

Scheduled a new site visit for them to go out this Friday @ 8AM Singapore Time so my Thursday @ 4PM.

1-260037210462

Scheduled a new site visit for them to go out this Friday @ 8AM Singapore Time so my Thursday @ 4PM.

1-260037210462

Hi @RobH: Was this work undertaken at this time?

Apologies, this ran super late and I neglected to update the task accordingly.

The mainboard swap was successful but it appears of the two CPUs, one of them has failed. Dell SG is sending over a quote for replacement, as the system is currently remotely accessible with only one of two cpus installed.
I'll update the task with the new quote later today!

Apologies, this ran super late and I neglected to update the task accordingly.

The mainboard swap was successful but it appears of the two CPUs, one of them has failed. Dell SG is sending over a quote for replacement, as the system is currently remotely accessible with only one of two cpus installed.
I'll update the task with the new quote later today!

Thanks for the update Rob! Let's hope this gets fixed by Dell soon.

RobH closed subtask Unknown Object (Task) as Resolved.May 21 2026, 5:45 PM
RobH mentioned this in Unknown Object (Task).May 21 2026, 5:50 PM
RobH added a subtask: Unknown Object (Task).
RobH updated the task description. (Show Details)

Without getting into pricing on this public task the options are:

  • spend more money (see T426985) to replace the CPU
    • we have no money left in expendables for this, so it would likely have to kick to July (unless mgmt approves overspend)
  • shuffle all the memory to report to the working CPU and use this CP host with only half the CPU cores
    • no clue if this is viable, this is a question for @ssingh
  • decommission the host and use as spare parts for rest of eqsin fleet
    • worst case scenario since it means a depreciated fleet quantity in eqsin.

Without getting into pricing on this public task the options are:

  • spend more money (see T426985) to replace the CPU
    • we have no money left in expendables for this, so it would likely have to kick to July (unless mgmt approves overspend)
  • shuffle all the memory to report to the working CPU and use this CP host with only half the CPU cores
    • no clue if this is viable, this is a question for @ssingh
  • decommission the host and use as spare parts for rest of eqsin fleet
    • worst case scenario since it means a depreciated fleet quantity in eqsin.

Hi @RobH. Thanks for the update. I will discuss with the team and follow up.

In regards to buying a new CPU - we don't have any more budget available for FY25-26, but I'm ok with going over budget if this is the best route forward. We'll have four additional cp servers (as part of the initiative for AI scraping and DDoS) scheduled in Q2 as well, if that helps with things. Just let me know you're preference on how to best proceed. Thanks, Willy

  • spend more money (see T426985) to replace the CPU
    • we have no money left in expendables for this, so it would likely have to kick to July (unless mgmt approves overspend)

We discussed this and the general consensus seemed to be to just decomm the server and wait for the refresh which is happening shortly anyway. @ssingh Is that accurate, and if so, ready for me to open a task/decom it?

We discussed this and the general consensus seemed to be to just decomm the server and wait for the refresh which is happening shortly anyway. @ssingh Is that accurate, and if so, ready for me to open a task/decom it?

That is still accurate but I will talk to Willy once again.

Update: Based on the discussion in the Traffic meeting, we would like to pursue the option of getting the additional CPU in July 2026. CC @wiki_willy and @RobH -- please let us know if we can pursue this and thanks.

Update: Based on the discussion in the Traffic meeting, we would like to pursue the option of getting the additional CPU in July 2026. CC @wiki_willy and @RobH -- please let us know if we can pursue this and thanks.

Sounds reasonable to me. I've just dropped a reply to our Dell SG rep letting them know we'd like to look at placing the CPU replacement order to land in July and asking for the quote to be refreshed. Once we get the quote back and check pricing I'll escalate over to Willy (during the last week of June) for approvals.