------------------------------------------------------------------------------- Record: 40 Date/Time: 04/05/2024 17:57:10 Source: system Severity: Critical Description: The system board BP1 PG voltage is outside of range. ------------------------------------------------------------------------------- Record: 41 Date/Time: 04/05/2024 17:57:30 Source: system Severity: Ok Description: The system board BP1 PG voltage is within range. -------------------------------------------------------------------------------
Description
Details
| Subject | Author | Repo | Branch | Lines +/- | |
|---|---|---|---|---|---|
| db1246: Disable notifications | Marostegui | operations/puppet | production | +1 -0 | |
| installserver: Format db1246 | Marostegui | operations/puppet | production | +1 -1 |
Related Objects
Event Timeline
Change #1017352 had a related patch set uploaded (by Marostegui; author: Marostegui):
[operations/puppet@production] db1246: Disable notifications
Hey @Marostegui I am currently looking at this unit. I checked both power supplied and logged into the machine. However, it seems to be healthy and it doesn't seem to be reporting any errors at this time. Would you be able to confirm? Let me know, thanks!
@VRiley-WMF the host is hard down for us, we cannot ssh to it. As I mentioned, this is the second time the host crashed for the same reason, do you think we could escalate this to dell?
Change #1017352 merged by Marostegui:
[operations/puppet@production] db1246: Disable notifications
Dell has suggested the following
Full reboot (Completed with no change)
Flea power drain reboot (Completed with no change)
Reseat all cables with another flea power drain reboot (Completed with no change)
Update iDRAC firmware (Complete with no change)
Update RAID controller firmware (Completed with no change)
Will continue to work with Dell on this ticket to see what other troubleshooting steps to proceed with.
Thank you! So far the host isn't available for ssh - maybe the whole filesystem is corrupted anyway as it happened here T359940. @VRiley-WMF let me know when I can proceed and reimage it.
Hi @Marostegui Thanks!
I just completed finishing up everything with Dell. Some other troubleshooting we did was
Strip the server down to bare bones (1 DIMM, 1 CPU, 1PSU) and tested it.
Reseated all the power connections (since the error was related to power) inside the server.
As of right now, since the unit is still showing healthy (however, with no boot), it seems like it should be cleared out according to Dell. If possible, please proceed with reimaging it as it seems that the filesystem has been corrupted.
Change #1017964 had a related patch set uploaded (by Marostegui; author: Marostegui):
[operations/puppet@production] installserver: Format db1246
Change #1017964 merged by Marostegui:
[operations/puppet@production] installserver: Format db1246
Cookbook cookbooks.sre.hosts.reimage was started by marostegui@cumin1002 for host db1246.eqiad.wmnet with OS bookworm
Cookbook cookbooks.sre.hosts.reimage started by marostegui@cumin1002 for host db1246.eqiad.wmnet with OS bookworm completed:
- db1246 (WARN)
- Downtimed on Icinga/Alertmanager
- Unable to disable Puppet, the host may have been unreachable
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- Host up (Debian installer)
- Add puppet_version metadata to Debian installer
- Checked BIOS boot parameters are back to normal
- Host up (new fresh bookworm OS)
- Generated Puppet certificate
- Signed new Puppet certificate
- Run Puppet in NOOP mode to populate exported resources in PuppetDB
- Found Nagios_host resource for this host in PuppetDB
- Downtimed the new host on Icinga/Alertmanager
- Removed previous downtime on Alertmanager (old OS)
- First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202404090555_marostegui_1502875_db1246.out
- configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
- Rebooted
- Automatic Puppet run was successful
- Forced a re-check of all Icinga services for the host
- Icinga status is not optimal, downtime not removed
- Updated Netbox data from PuppetDB
Host recloned. I am going to leave it running for 24h before repooling it back into production, just in case.
Thanks for the help @VRiley-WMF!