- - Provide FQDN of system.
- - If other than a hard drive issue, please depool the machine (and confirm that it’s been depooled) for us to work on it. If not, please provide time frame for us to take the machine down.
Machine is currently serving but can be taken out of service on very short notice.
- - Provide urgency of request, along with justification (redundancy, dependencies, etc)
Not super urgent, since machine is still working with reduced capacity.
- - Describe issue and/or attach hardware failure log. (Refer to https://wikitech.wikimedia.org/wiki/Dc-operations/Hardware_Troubleshooting_Runbook if you need help)
During reboot, BIOS will halt with error messages about failed DIMMs.
- - Assign correct project tag and appropriate owner (based on above). Also, please ensure the service owners of the host(s) are added as subscribers to provide any additional input.
racadm getsel log:
The DIMMs affected are A1 and B2. Usually, cold booting the machine helps for a while. The two DIMMs have been doing this about once a year since 2021. I doubt they're truly broken, but may just need a re-seat (or it's the PCB the memory sockets are on, but the "fix" is the same).
The machine can be taken out of service/shutdown on short notice during CEST daytime. If that does not fit the DC work schedule, we can of course take it out for a few days.