User Details
- User Since
- Jun 29 2021, 9:56 AM (267 w, 3 d)
- Availability
- Available
- IRC Nick
- btullis
- LDAP User
- Btullis
- MediaWiki User
- BTullis (WMF) [ Global Accounts ]
Today
Hi, sorry for the delay in getting back to you about this.
All of the missing disks have been fixed up, so this will unblock all remaining an-worker upgrades.
All workers have 12 data disks.
btullis@cumin1003:~$ sudo cumin A:hadoop-worker 'blkid |grep -c hadoop-' 91 hosts will be targeted: an-worker[1142-1147,1149-1151,1153-1187,1189-1190,1192-1236].eqiad.wmnet OK to proceed on 91 hosts? Enter the number of affected hosts to confirm or "q" to quit: 91 ===== NODE GROUP ===== (91) an-worker[1142-1147,1149-1151,1153-1187,1189-1190,1192-1236].eqiad.wmnet ----- OUTPUT for command #1: 'blkid |grep -c hadoop-' ----- 12 ================
Fixed up an-worker1208
btullis@an-worker1208:~$ sudo blkid|grep -c hadoop- 11
Fixed up an-worker1207.
Physical Drives = 14
Fixed up an-worker1205.
Physical Drives = 14
Fixed up an-worker1201.
Booted into emergency mode.
You are in emergency mode. AfterGive root password for maintenance
(or press Control-D to continue):
root@an-worker1201:~# lsblk
NAME MAJ:MIN RM SIZE RO TYPE MOUNTPOINT
sda 8:0 0 7.3T 0 disk
└─sda1 8:1 0 7.3T 0 part /var/lib/hadoop/data/l
sdb 8:16 0 7.3T 0 disk
└─sdb1 8:17 0 7.3T 0 part /var/lib/hadoop/data/k
sdc 8:32 0 7.3T 0 disk
sdd 8:48 0 7.3T 0 disk
└─sdd1 8:49 0 7.3T 0 part /var/lib/hadoop/data/m
sde 8:64 0 7.3T 0 disk
└─sde1 8:65 0 7.3T 0 part /var/lib/hadoop/data/j
sdf 8:80 0 7.3T 0 disk
└─sdf1 8:81 0 7.3T 0 part /var/lib/hadoop/data/h
sdg 8:96 0 7.3T 0 disk
└─sdg1 8:97 0 7.3T 0 part /var/lib/hadoop/data/g
sdh 8:112 0 7.3T 0 disk
└─sdh1 8:113 0 7.3T 0 part /var/lib/hadoop/data/f
sdi 8:128 0 7.3T 0 disk
sdj 8:144 0 7.3T 0 disk
└─sdj1 8:145 0 7.3T 0 part /var/lib/hadoop/data/d
sdk 8:160 0 7.3T 0 disk
└─sdk1 8:161 0 7.3T 0 part /var/lib/hadoop/data/c
sdl 8:176 0 7.3T 0 disk
└─sdl1 8:177 0 7.3T 0 part /var/lib/hadoop/data/b
sdm 8:192 0 446.6G 0 disk
├─sdm1 8:193 0 953M 0 part /boot
├─sdm2 8:194 0 1K 0 part
└─sdm5 8:197 0 445.7G 0 part
├─an--worker1201--vg-swap 254:0 0 9.3G 0 lvm [SWAP]
├─an--worker1201--vg-root 254:1 0 55.9G 0 lvm /
└─an--worker1201--vg-journalnode
254:2 0 10G 0 lvm /var/lib/hadoop/journalFixed up an-worker1231
Physical Drives = 14
Yesterday
Working on an-worker1201.
btullis@an-worker1201:~$ sudo perccli64 /c0 show
Now working on an-worker1200
Physical Drives = 14
Now working on an-worker 1199
Physical Drives = 14
an-worker1204 has three drives missing.
I have booted it but then it stopped and I have logged into over the serial console.
Physical Drives = 14
I am working to fix up all of the missing disk issues in T426610: Follow up on multiple RAID / drive issues
I did the first of these, an-worker1213.
Wed, Aug 12
We also have 9 servers where there are not 12 data disks showing up.
These need fixing before the reimage cookbook will run correctly.
btullis@cumin1003:~$ sudo cumin A:hadoop-worker 'blkid | grep -c hadoop-' 91 hosts will be targeted: an-worker[1142-1147,1149-1177,1179-1187,1189-1190,1192-1236].eqiad.wmnet OK to proceed on 91 hosts? Enter the number of affected hosts to confirm or "q" to quit: 91 ===== NODE GROUP ===== (5) an-worker[1200,1207-1208,1213,1231].eqiad.wmnet ----- OUTPUT for command #1: 'blkid | grep -c hadoop-' ----- 11 ===== NODE GROUP ===== (1) an-worker1204.eqiad.wmnet ----- OUTPUT for command #1: 'blkid | grep -c hadoop-' ----- 9 ===== NODE GROUP ===== (3) an-worker[1199,1201,1205].eqiad.wmnet ----- OUTPUT for command #1: 'blkid | grep -c hadoop-' ----- 10 ===== NODE GROUP ===== (82) an-worker[1142-1147,1149-1177,1179-1187,1189-1190,1192-1198,1202-1203,1206,1209-1212,1214-1230,1232-1236].eqiad.wmnet ----- OUTPUT for command #1: 'blkid | grep -c hadoop-' ----- 12 ================
We are down to 100 hosts in the analytics and presto clusters.
I'm going to do an in-place upgrade for Archiva, as it is a somewhat problematic VM with large data stores, that we are hoping to kill soon.
I don't want to spend the time getting a reuse-parts recipe working for this one-off host.
Tue, Aug 11
I am going to start upgrading the nameservers, starting with the standby, an-master1004.
Apologies for the delay.
I have now added you to the airflow-analytics-ops LDAP group, as discussed here.
It looks like 4 of the 15 hosts were detected this way around, so maybe it should work this way around.
btullis@cumin1003:~$ sudo cumin A:an-presto pvs 14 hosts will be targeted: an-presto[1006-1007,1009-1020].eqiad.wmnet OK to proceed on 14 hosts? Enter the number of affected hosts to confirm or "q" to quit: 14 ===== NODE GROUP ===== (3) an-presto[1011,1014,1016].eqiad.wmnet ----- OUTPUT for command #1: 'pvs' ----- PV VG Fmt Attr PSize PFree /dev/sda1 vg1 lvm2 a-- <21.83t <4.37t /dev/sdb2 vg0 lvm2 a-- 446.34g 89.27g ===== NODE GROUP ===== (11) an-presto[1006-1007,1009-1010,1012-1013,1015,1017-1020].eqiad.wmnet ----- OUTPUT for command #1: 'pvs' ----- PV VG Fmt Attr PSize PFree /dev/sda2 vg0 lvm2 a-- 446.34g 89.27g /dev/sdb1 vg1 lvm2 a-- <21.83t <4.37t ================
I might be able to force the detection order by deleting and recreating the RAID volumes, or running the sre.hosts.provision cookbook on them, perhaps.
an-presto1008 is giving a partman error.
It seems to be because the large RAID disk has been detected as /dev/sda, rather than /dev/sdb
~ # pvs PV VG Fmt Attr PSize PFree /dev/sda2 vg0 lvm2 a-- <21.83t <4.37t /dev/sdb1 vg1 lvm2 a-- 446.62g 446.62g
I have created the directory, as requested.
btullis@an-launcher1003:~$ sudo -u hdfs kerberos-run-command hdfs hdfs dfs -mkdir /wmf/data/wmf_apps
Wed, Jul 29
I still have a puppet cleanup patch to supply and some keytabs to remove, but please feel free to proceed with de-racking and disposing of these servers.
I still have a puppet cleanup patch to supply and some keytabs to remove, but please feel free to proceed with de-racking and disposing of this server.
I have modified the title to reflect that this is wikibase dumps, as opposed to XML/SQL dumps, or any of the other types.
I see that the log files on the relforge cluster seem to be generating ECS fortmatted logs now.
Thanks @bking for checking this out.
btullis@relforge1010:~$ tail -n 1 -f /var/log/opensearch/relforge-eqiad_server.json | jq . --unbuffered
{
"@timestamp": "2026-07-28T22:16:43.396Z",
"log.level": "INFO",
"message": "finalizing recovery took [12.3ms]",
"ecs.version": "1.2.0",
"service.name": "relforge-eqiad",
"event.dataset": "relforge-eqiad",
"process.thread.name": "opensearch[relforge1010-relforge-eqiad][generic][T#1]",
"log.logger": "org.opensearch.indices.recovery.RecoverySourceHandler"
}Was anything required other than a rolling restart of the servcies, or did this work without any issue?
The offending call to python is here.
btullis@an-test-coord1002:~$ head /usr/lib/presto/bin/launcher.py #!/usr/bin/env python
Oh, it turns out that there is another small issue with presto.
Jul 29 11:03:03 an-test-coord1002 presto-server[2265682]: /usr/bin/env: ‘python’: No such file or directory Jul 29 11:03:03 an-test-coord1002 systemd[1]: presto-server.service: Main process exited, code=exited, status=127/n/a Jul 29 11:03:03 an-test-coord1002 systemd[1]: presto-server.service: Failed with result 'exit-code'.
I might need to rebuild the package for this, but let's see.
root@krb1002:~# kadmin.local ktadd -norandkey -k /srv/kerberos/keytabs/an-test-coord1002.eqiad.wmnet/hive/hive.keytab hive/analytics-test-hive.eqiad.wmnet@WIKIMEDIA Entry for principal hive/analytics-test-hive.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-test-coord1002.eqiad.wmnet/hive/hive.keytab.
There's a small issue with starting the hive-metastore and presto-server services on the refreshed an-coord1002 host, but this is because I didn't add the required primitives to the keytabs.
Now I am about to switch the MariaDAB database on an-test-coord1002 to be a standalone master, instead of a replication slave of an-test-coord1001.
Merged all of the patches.
btullis@an-test-master1001:~$ sudo -u hdfs kerberos-run-command hdfs hdfs dfsadmin -safemode enter Safe mode is ON in an-test-master1001.eqiad.wmnet/10.64.5.39:8020 Safe mode is ON in an-test-master1002.eqiad.wmnet/10.64.36.112:8020
Now transitioning to standby.
btullis@an-test-master1001:~$ sudo -u hdfs hdfs haadmin -transitionToStandby an-test-master1001-eqiad-wmnet --forcemanual You have specified the --forcemanual flag. This flag is dangerous, as it can induce a split-brain scenario that WILL CORRUPT your HDFS namespace, possibly irrecoverably.

