Page MenuHomePhabricator

VRiley-WMF (Valerie Riley)
User

Projects (3)

Today

  • No visible events.

Tomorrow

  • No visible events.

Tuesday

  • No visible events.

User Details

User Since
Aug 22 2023, 3:06 PM (154 w, 5 d)
Availability
Available
LDAP User
ValerieRiley
MediaWiki User
VRiley-WMF [ Global Accounts ]

Recent Activity

Thu, Aug 6

VRiley-WMF added a comment to T431115: db1245 crashed.

Hey @jcrespo I updated the firmware and reseated some of the fans. Would you be able to try to test putting a load on it? I was looking for thermal paste, but can't seem to find any. If that doesn't fix it, we can put in an order for thermal paste.

Thu, Aug 6, 6:40 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF closed T434118: Alert for device ps1-e2-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit as Resolved.

Will monitor this.

Thu, Aug 6, 6:26 PM · SRE, DC-Ops, ops-eqiad
VRiley-WMF closed T433970: decommission backup1003.eqiad.wmnet as Resolved.

This is completed.

Thu, Aug 6, 5:13 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence, bacula, Data-Persistence-Backup, decommission-hardware
VRiley-WMF closed T433970: decommission backup1003.eqiad.wmnet, a subtask of T420506: Setup backup[12]01[456789] & backup[12]020 and migrate data to them; prepare for decommission backup[12]00[34567], as Resolved.
Thu, Aug 6, 5:13 PM · Data-Persistence-Backup, media-backups, database-backups, bacula
VRiley-WMF updated the task description for T433970: decommission backup1003.eqiad.wmnet.
Thu, Aug 6, 5:13 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence, bacula, Data-Persistence-Backup, decommission-hardware
VRiley-WMF claimed T433970: decommission backup1003.eqiad.wmnet.
Thu, Aug 6, 4:49 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence, bacula, Data-Persistence-Backup, decommission-hardware
VRiley-WMF moved T433970: decommission backup1003.eqiad.wmnet from Backlog to Decommission on the ops-eqiad board.
Thu, Aug 6, 4:49 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence, bacula, Data-Persistence-Backup, decommission-hardware

Wed, Aug 5

VRiley-WMF moved T433030: Degraded RAID on an-presto1013 from Hardware Failure / Troubleshoot to Blocked on the ops-eqiad board.
Wed, Aug 5, 5:33 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), SRE, ops-eqiad, DC-Ops
VRiley-WMF added a comment to T433030: Degraded RAID on an-presto1013.

Currently waiting on HDD to arrive

Wed, Aug 5, 5:33 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), SRE, ops-eqiad, DC-Ops
VRiley-WMF updated subscribers of T431828: Degraded RAID on an-worker1191.

@brouberol Hey I wanted to check in with you on this device as I believe Ben may be out of the office for a bit. Would there be anything I need to help with from this point?

Wed, Aug 5, 5:30 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), SRE, ops-eqiad, DC-Ops
VRiley-WMF moved T434118: Alert for device ps1-e2-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Wed, Aug 5, 5:28 PM · SRE, DC-Ops, ops-eqiad
VRiley-WMF closed T433825: decommission db1150.eqiad.wmnet as Resolved.

This has been completed.

Wed, Aug 5, 5:27 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF closed T433825: decommission db1150.eqiad.wmnet, a subtask of T433475: Decommission db11[50-82,84], as Resolved.
Wed, Aug 5, 5:27 PM · DBA
VRiley-WMF updated the task description for T433825: decommission db1150.eqiad.wmnet.
Wed, Aug 5, 5:27 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF closed T433826: decommission db1171.eqiad.wmnet as Resolved.

This has been completed.

Wed, Aug 5, 5:26 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF closed T433826: decommission db1171.eqiad.wmnet, a subtask of T433475: Decommission db11[50-82,84], as Resolved.
Wed, Aug 5, 5:26 PM · DBA
VRiley-WMF updated the task description for T433826: decommission db1171.eqiad.wmnet.
Wed, Aug 5, 5:26 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF claimed T433825: decommission db1150.eqiad.wmnet.
Wed, Aug 5, 4:25 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF claimed T433826: decommission db1171.eqiad.wmnet.
Wed, Aug 5, 4:25 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF added a comment to T433348: Unresponsive management for an-worker1147.mgmt:22.

Thank you! I'm ready whenever they are ready to take it down.

Wed, Aug 5, 4:24 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), DC-Ops, ops-eqiad
VRiley-WMF moved T433825: decommission db1150.eqiad.wmnet from Backlog to Decommission on the ops-eqiad board.
Wed, Aug 5, 4:21 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware
VRiley-WMF moved T433826: decommission db1171.eqiad.wmnet from Backlog to Decommission on the ops-eqiad board.
Wed, Aug 5, 4:21 PM · SRE, DC-Ops, ops-eqiad, database-backups, Data-Persistence-Backup, Data-Persistence, decommission-hardware

Tue, Aug 4

VRiley-WMF added a comment to T431115: db1245 crashed.

I was given an action plan and will be following through with these steps. Reseat all fans, reseat the heatsink and cpu. update bios firmware. Firmware is now out of date due to the replacment.

Tue, Aug 4, 9:01 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF removed a project from T433494: decommission an-test-coord1001.eqiad.wmnet: ops-eqiad.
Tue, Aug 4, 8:15 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF removed a project from T433495: decommission an-test-master100[1-2]: ops-eqiad.
Tue, Aug 4, 8:14 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF updated subscribers of T433348: Unresponsive management for an-worker1147.mgmt:22.

@bking Hey, I wanted to reach out about this server. Is there a good time to reboot this? We're having some issues with the iDRAC and I'd like to try to power drain it. We can do it anytime today or tomorrow.

Tue, Aug 4, 3:52 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), DC-Ops, ops-eqiad

Mon, Aug 3

VRiley-WMF claimed T429267: cloudcephosd1044 boot issues.
Mon, Aug 3, 9:43 PM · SRE, ops-eqiad, DC-Ops, tools-infrastructure-team, cloud-services-team, Ceph, Cloud-VPS
VRiley-WMF claimed T433348: Unresponsive management for an-worker1147.mgmt:22.
Mon, Aug 3, 5:39 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), DC-Ops, ops-eqiad
VRiley-WMF updated the task description for T431682: Rebalance cloudvirts out of E4 and into C8.
Mon, Aug 3, 5:11 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF added a comment to T431682: Rebalance cloudvirts out of E4 and into C8.

@Andrew Awesome, sounds good. Would we like to start on 1049?

Mon, Aug 3, 5:11 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF claimed T433565: db1218 crashed.
Mon, Aug 3, 5:07 PM · SRE, DC-Ops, ops-eqiad, DBA

Fri, Jul 31

VRiley-WMF added a comment to T429267: cloudcephosd1044 boot issues.

opened ticket 229798904 with dell

Fri, Jul 31, 2:13 PM · SRE, ops-eqiad, DC-Ops, tools-infrastructure-team, cloud-services-team, Ceph, Cloud-VPS
VRiley-WMF added a comment to T429267: cloudcephosd1044 boot issues.

Updated firmware on iDRAC and BIOS. No change of yet. Grabbing TSR and will be opening up a ticket with Dell

Fri, Jul 31, 12:06 PM · SRE, ops-eqiad, DC-Ops, tools-infrastructure-team, cloud-services-team, Ceph, Cloud-VPS
VRiley-WMF changed the status of T429267: cloudcephosd1044 boot issues from Open to In Progress.
Fri, Jul 31, 11:18 AM · SRE, ops-eqiad, DC-Ops, tools-infrastructure-team, cloud-services-team, Ceph, Cloud-VPS

Thu, Jul 30

VRiley-WMF changed the status of T433565: db1218 crashed from In Progress to Open.

iDRAC is at 7.30 now

Thu, Jul 30, 5:03 PM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF moved T433565: db1218 crashed from High Priority Task to Hardware Failure / Troubleshoot on the ops-eqiad board.
Thu, Jul 30, 4:54 PM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

Hey @Dzahn I may need some help with zuul1006. I believe all these servers will be using a 1 gig cable, but this unit is in a rack where they use fiber. I'm currently looking to see how to fix that. However, zuul1004 and zuul1005 are up

Thu, Jul 30, 4:52 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF changed the status of T433565: db1218 crashed from Open to In Progress.
Thu, Jul 30, 4:38 PM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF added a comment to T433565: db1218 crashed.

BIOS is now at 1.21

Thu, Jul 30, 4:37 PM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF added a comment to T433565: db1218 crashed.

BIOS firmware installing now

Thu, Jul 30, 4:27 PM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF added a project to T431828: Degraded RAID on an-worker1191: Data-Platform-SRE (2026-07-03 - 2026-07-31).
Thu, Jul 30, 2:04 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), SRE, ops-eqiad, DC-Ops
VRiley-WMF added a comment to T431115: db1245 crashed.

I was able to log in via iDRAC and power it on. It was seemingly seeing an issue with tempature, which I may need to address. However, could you please test that out?

Thu, Jul 30, 9:36 AM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF added a comment to T433565: db1218 crashed.

Thanks, I'll take a look at this and update the firmware in just a moment

Thu, Jul 30, 8:38 AM · SRE, DC-Ops, ops-eqiad, DBA
VRiley-WMF closed T433368: Unresponsive management for ms-be1065.mgmt:22 as Resolved.

Reseated cable and there is now activity on it. It should be good to go.

Thu, Jul 30, 8:29 AM · SRE, DC-Ops, ops-eqiad
VRiley-WMF added a comment to T433494: decommission an-test-coord1001.eqiad.wmnet.

Hey @BTullis this is all yours. It has been physically removed.

Thu, Jul 30, 8:26 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF updated the task description for T433494: decommission an-test-coord1001.eqiad.wmnet.
Thu, Jul 30, 8:26 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF assigned T433494: decommission an-test-coord1001.eqiad.wmnet to BTullis.
Thu, Jul 30, 8:25 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF closed T433578: Physical decommission of an-test-coord1001 as Resolved.

This is completed.

Thu, Jul 30, 8:25 AM · SRE, DC-Ops, ops-eqiad, decommission-hardware
VRiley-WMF closed T433578: Physical decommission of an-test-coord1001, a subtask of T433494: decommission an-test-coord1001.eqiad.wmnet, as Resolved.
Thu, Jul 30, 8:25 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF created T433578: Physical decommission of an-test-coord1001.
Thu, Jul 30, 8:25 AM · SRE, DC-Ops, ops-eqiad, decommission-hardware
VRiley-WMF added a comment to T433495: decommission an-test-master100[1-2].

Hey @BTullis my part has been completed. I'm tossing this your way to finish it out, thanks!

Thu, Jul 30, 8:19 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF reassigned T433495: decommission an-test-master100[1-2] from VRiley-WMF to BTullis.
Thu, Jul 30, 8:19 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF closed T433576: Physical uninstall of an-test-master100[1-2], a subtask of T433495: decommission an-test-master100[1-2], as Resolved.
Thu, Jul 30, 8:19 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF closed T433576: Physical uninstall of an-test-master100[1-2] as Resolved.

This has been completed

Thu, Jul 30, 8:19 AM · SRE, ops-eqiad, DC-Ops, decommission-hardware
VRiley-WMF created T433576: Physical uninstall of an-test-master100[1-2].
Thu, Jul 30, 8:18 AM · SRE, ops-eqiad, DC-Ops, decommission-hardware
VRiley-WMF updated the task description for T433495: decommission an-test-master100[1-2].
Thu, Jul 30, 8:17 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF claimed T433495: decommission an-test-master100[1-2].
Thu, Jul 30, 8:11 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF moved T433565: db1218 crashed from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Thu, Jul 30, 8:07 AM · SRE, DC-Ops, ops-eqiad, DBA

Wed, Jul 29

VRiley-WMF added a comment to T431682: Rebalance cloudvirts out of E4 and into C8.

@fgiunchedi I was able to get through a good majority of this process, but I'm getting a hangup from the reimage. It seems like it's seeing all the correct ports and the fiber has activity. Should I try the reimage with something other than trixie?

Wed, Jul 29, 8:26 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF added a comment to T431682: Rebalance cloudvirts out of E4 and into C8.

Updating firmware on cloudvirt1048

Wed, Jul 29, 7:47 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF added a comment to T431682: Rebalance cloudvirts out of E4 and into C8.

Moved unit into

Wed, Jul 29, 6:57 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF changed the status of T431682: Rebalance cloudvirts out of E4 and into C8 from Open to In Progress.

Moving this now.

Wed, Jul 29, 6:27 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF changed the status of T431682: Rebalance cloudvirts out of E4 and into C8, a subtask of T424658: Ensure cloudvirt capacity is more evenly spread out among racks, from Open to In Progress.
Wed, Jul 29, 6:27 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF moved T433494: decommission an-test-coord1001.eqiad.wmnet from Backlog to Decommission on the ops-eqiad board.
Wed, Jul 29, 4:32 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF moved T433495: decommission an-test-master100[1-2] from Backlog to Decommission on the ops-eqiad board.
Wed, Jul 29, 4:32 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
VRiley-WMF changed the status of T431115: db1245 crashed from In Progress to Open.
Wed, Jul 29, 3:47 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF added a comment to T431115: db1245 crashed.

System is powered off, but it is reachable. Should be good to go. @jcrespo should we be okay to close this ticket?

Wed, Jul 29, 3:47 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF added a comment to T431115: db1245 crashed.

Dell has come on site and replaced the mainboard. I have logged the new MAC addresses (if needed) into Netbox. I am collecting a TSR report for Dell and then will power this unit off.

Wed, Jul 29, 3:33 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF moved T433368: Unresponsive management for ms-be1065.mgmt:22 from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Wed, Jul 29, 3:01 PM · SRE, DC-Ops, ops-eqiad
VRiley-WMF moved T429267: cloudcephosd1044 boot issues from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Wed, Jul 29, 3:01 PM · SRE, ops-eqiad, DC-Ops, tools-infrastructure-team, cloud-services-team, Ceph, Cloud-VPS
VRiley-WMF moved T433348: Unresponsive management for an-worker1147.mgmt:22 from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Wed, Jul 29, 3:01 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31), DC-Ops, ops-eqiad
VRiley-WMF added a comment to T431115: db1245 crashed.

Of course @jcrespo I will make sure I do that. Thanks!

Wed, Jul 29, 2:55 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF changed the status of T431115: db1245 crashed from Open to In Progress.

Dell is onsite and proceeding with the mainboard replacment.

Wed, Jul 29, 2:13 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup

Tue, Jul 28

VRiley-WMF added a comment to T431682: Rebalance cloudvirts out of E4 and into C8.

Sure thing, I'm planning on this tomorrow. Thank you!

Tue, Jul 28, 8:51 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

Oh, I could be wrong. It looks like it just finished

Tue, Jul 28, 7:07 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

@Dzahn would you be able to take a look at Zuul1005? It looks like it was able to finish the installer but it looks like it's going to fail the re-image script

Tue, Jul 28, 7:05 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

Currently, if you wanted to have zuul2004-7 we would need to find servers that are located in codfw. This would be on another ticket or a sub ticket for those devices. I apologize, I thought there was already a ticket for those devices.

Tue, Jul 28, 5:10 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

So, zuul1004-1007 would be in eqiad. If it starts with a 2, it would be for codfw

Tue, Jul 28, 5:08 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

I'm currently working on those as we speak

Tue, Jul 28, 5:07 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE

Mon, Jul 27

VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

Hey @MoritzMuehlenhoff is there a recommended OS for these devices?

Mon, Jul 27, 10:45 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T427353: Repurpose ganeti102[3456] for Zuul migration.

Hey @Dzahn thanks for the ping on this. I've recently been freed up to work on this a bit more as I was having a lot of difficulty with these servers at first.

Mon, Jul 27, 9:10 PM · Patch-For-Review, DC-Ops, ops-eqiad, collaboration-services, SRE
VRiley-WMF added a comment to T431115: db1245 crashed.

Dell called me today with an update. They are still having trouble obtaining the mainboard. However, they did assure me that a tech will be onsite Wednesday the 29th in order to replace it. Will update then.

Mon, Jul 27, 8:25 PM · SRE, DC-Ops, ops-eqiad, Data-Persistence-Backup
VRiley-WMF closed T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet as Resolved.
Mon, Jul 27, 7:00 PM · SRE, ops-eqiad, DC-Ops
VRiley-WMF closed T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet, a subtask of T431542: db1208 has a read-only file system / broken disk, as Resolved.
Mon, Jul 27, 7:00 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31)
VRiley-WMF updated the task description for T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet.
Mon, Jul 27, 6:58 PM · SRE, ops-eqiad, DC-Ops
VRiley-WMF changed the status of T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet, a subtask of T431542: db1208 has a read-only file system / broken disk, from In Progress to Open.
Mon, Jul 27, 6:55 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31)
VRiley-WMF changed the status of T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet from In Progress to Open.

Swapped out DIMM and matched what it has in there. @BTullis please see if there are any issues, it should be good to go.

Mon, Jul 27, 6:55 PM · SRE, ops-eqiad, DC-Ops
VRiley-WMF updated the task description for T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet.
Mon, Jul 27, 6:45 PM · SRE, ops-eqiad, DC-Ops
VRiley-WMF changed the status of T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet, a subtask of T431542: db1208 has a read-only file system / broken disk, from Open to In Progress.
Mon, Jul 27, 6:44 PM · Data-Platform-SRE (2026-07-03 - 2026-07-31)
VRiley-WMF changed the status of T432116: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet from Open to In Progress.

moving forward with this...

Mon, Jul 27, 6:44 PM · SRE, ops-eqiad, DC-Ops
VRiley-WMF closed T433284: arclamp1001 ram upgrade, a subtask of T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet), as Resolved.
Mon, Jul 27, 5:45 PM · Observability-Metrics, SRE, Wikimedia-production-error
VRiley-WMF closed T433284: arclamp1001 ram upgrade as Resolved.

Added requested DIMM into arclamp1001.

Mon, Jul 27, 5:45 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE
VRiley-WMF updated the task description for T433284: arclamp1001 ram upgrade.
Mon, Jul 27, 5:44 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE
VRiley-WMF changed the status of T433284: arclamp1001 ram upgrade, a subtask of T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet), from Open to In Progress.
Mon, Jul 27, 5:32 PM · Observability-Metrics, SRE, Wikimedia-production-error
VRiley-WMF changed the status of T433284: arclamp1001 ram upgrade from Open to In Progress.

Proceeding with this as @colewhite said we can take this unit down today.

Mon, Jul 27, 5:32 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE
VRiley-WMF updated the task description for T433284: arclamp1001 ram upgrade.
Mon, Jul 27, 5:32 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE
VRiley-WMF moved T433284: arclamp1001 ram upgrade from Backlog to Hardware Failure / Troubleshoot on the ops-eqiad board.
Mon, Jul 27, 5:31 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE

Fri, Jul 24

VRiley-WMF added a comment to T432317: Fix cables with placeholders names and planned status.

Updated several cables, will need to continue to update a few more.

Fri, Jul 24, 9:50 PM · SRE, DC-Ops, ops-eqsin, ops-eqdfw, ops-codfw, ops-magru, ops-eqiad, ops-esams, ops-drmrs, ops-ulsfo
VRiley-WMF updated the task description for T432317: Fix cables with placeholders names and planned status.
Fri, Jul 24, 9:48 PM · SRE, DC-Ops, ops-eqsin, ops-eqdfw, ops-codfw, ops-magru, ops-eqiad, ops-esams, ops-drmrs, ops-ulsfo
VRiley-WMF updated the task description for T432317: Fix cables with placeholders names and planned status.
Fri, Jul 24, 9:39 PM · SRE, DC-Ops, ops-eqsin, ops-eqdfw, ops-codfw, ops-magru, ops-eqiad, ops-esams, ops-drmrs, ops-ulsfo