Page MenuHomePhabricator

hw troubleshooting: memory errors on ml-serve2004.codfw.wmnet
Closed, ResolvedPublicRequest

Description

  • - Provide FQDN of system.
  • - If other than a hard drive issue, please depool the machine (and confirm that it’s been depooled) for us to work on it. If not, please provide time frame for us to take the machine down.

Machine is currently serving but can be taken out of service on very short notice.

  • - Provide urgency of request, along with justification (redundancy, dependencies, etc)

Not super urgent, since machine is still working with reduced capacity.

During reboot, BIOS will halt with error messages about failed DIMMs.

  • - Assign correct project tag and appropriate owner (based on above). Also, please ensure the service owners of the host(s) are added as subscribers to provide any additional input.

Like ml-serve2002, this is a repeat offender, and usually cold-booting it helps, but again, only for a while. So similarly, reseating may help.

The machine can be taken out of service/shutdown on short notice during CEST daytime. If that does not fit the DC work schedule, we can of course take it out for a few days.

Event Timeline

Mentioned in SAL (#wikimedia-operations) [2026-07-29T15:11:17Z] <klausman@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: T433478

Jhancock.wm claimed this task.
Jhancock.wm subscribed.

learned lesson from T433476. same errors on a reseat/reboot.
pulled the DIMM, set a side, and installed DIMM from decommed server. Will clean later and test to put back into stock. idrac and bios firmware have also been updated.