TZ: UTC +1/+2
User Details
- User Since
- Sep 1 2016, 6:48 AM (518 w, 5 d)
- Availability
- Available
- IRC Nick
- marostegui
- LDAP User
- Marostegui
- MediaWiki User
- MArostegui (WMF) [ Global Accounts ]
Yesterday
@Jdforrester-WMF could you let us know when we can proceed?
Thanks
[Tue Aug 4 17:48:49 2026] {6}[Hardware Error]: DIMM location: not present. DMI handle: 0x0000
[Tue Aug 4 17:48:49 2026] EDAC MC3: 1 UE memory read error on CPU_SrcID#0_MC#3_Chan#1_DIMM#0 (channel:1 slot:0 page:0x15e95e1 offset:0x240 grain:32 - err_code:0x0000:0x009f SystemAddress:0x15e95e1240 ProcessorSocketId:0x0 MemoryControllerId:0x3 ChannelAddress:0x2ad2bc240 ChannelId:0x1 RankAddress:0x15695e240 PhysicalRankId:0x0 DimmSlotId:0x0 Row:0x9b4f Column:0x48 Bank:0x3 BankGroup:0x0 ChipSelect:0x0 ChipId:0x0)This is a few days old:
Description: A critical diagnostic event occurred in the memory device at A8. Contact your service provider for assistance in replacing the device. (Extended ID: 0x4E41).
Not yet ready
This is ready for DCOps.
All done!
All done
Not ready yet, this host was switched today as part of https://phabricator.wikimedia.org/T434288 so giving it a few days before decom
db1164 was promoted to master, db1213 is no longer a master and needs to be reimaged and then switched back.
This was done
Fri, Aug 7
[11:29:48] marostegui@cumin1003:~$ time sudo dbctl config commit -m "Adding db1271 to dbctl"
eqiad/hostsByName live eqiad/hostsByName generated
"db1260": "10.64.0.187", "db1260": "10.64.0.187",
"db1261": "10.64.16.77", "db1261": "10.64.16.77",
"db1262": "10.64.32.64", "db1262": "10.64.32.64",
"db1263": "10.64.48.47", "db1263": "10.64.48.47",
"db1264": "10.64.16.6", "db1264": "10.64.16.6",
"db1271": "10.64.0.155",
"es1035": "10.64.152.4", "es1035": "10.64.152.4",
"es1036": "10.64.154.3", "es1036": "10.64.154.3",
"es1037": "10.64.156.3", "es1037": "10.64.156.3",
"es1038": "10.64.160.4", "es1038": "10.64.160.4",
"es1039": "10.64.162.3", "es1039": "10.64.162.3",Yes, that is correct. I've +1ed it.
Ready for DC-Ops!
Thu, Aug 6
I've started replication on db1266 in ms2, this means ms2 topology is a bit strange at the moment, but I am going to leave it replicating for a few days before going ahead and setting up as a master
I've started replication on db1268 in ms3, this means ms3 topology is a bit strange at the moment, but I am going to leave it replicating for a few days before going ahead and setting up as a master
@jcrespo I will schedule this for Tuesday in the end, I'll do it after checking the backups table to makes sure nothing is in progress. The RO time would be around 5-10 seconds and I will merge your patch inmediately after.
Wed, Aug 5
I think if you document this in wikitech we can close this?
I've started replication on db1267 in ms1, this means ms1 topology is a bit strange at the moment, but I am going to leave it replicating for a few days before going ahead and setting up as a master
All new hosts are now serving traffic. Closing this task and will open a new one to decommission the old ones if all goes fine, in a month.
Ok thanks! I will adjust
Connectivity between proxies and db1164 has been tested.
@jcrespo could you let me know where it would be a BAD moment to do this switchover from a backups point of view so I can avoid that window?
Thanks