Page MenuHomePhabricator

BTullis (Ben)
Staff SRE

Today

  • No visible events.

Tomorrow

  • No visible events.

Sunday

  • No visible events.

User Details

User Since
Jun 29 2021, 9:56 AM (267 w, 3 d)
Availability
Available
IRC Nick
btullis
LDAP User
Btullis
MediaWiki User
BTullis (WMF) [ Global Accounts ]

Recent Activity

Today

BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Fri, Aug 14, 11:00 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 10:56 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Fri, Aug 14, 10:45 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis moved T433348: Unresponsive management for an-worker1147.mgmt:22 from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-07 - 2026-08-28) board.

Hi, sorry for the delay in getting back to you about this.

Fri, Aug 14, 10:29 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), DC-Ops, ops-eqiad
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

All of the missing disks have been fixed up, so this will unblock all remaining an-worker upgrades.
All workers have 12 data disks.

btullis@cumin1003:~$ sudo cumin A:hadoop-worker 'blkid |grep -c hadoop-'
91 hosts will be targeted:
an-worker[1142-1147,1149-1151,1153-1187,1189-1190,1192-1236].eqiad.wmnet
OK to proceed on 91 hosts? Enter the number of affected hosts to confirm or "q" to quit: 91
===== NODE GROUP =====                                                                                                                                                  
(91) an-worker[1142-1147,1149-1151,1153-1187,1189-1190,1192-1236].eqiad.wmnet                                                                                           
----- OUTPUT for command #1: 'blkid |grep -c hadoop-' -----                                                                                                             
12                                                                                                                                                                      
================
Fri, Aug 14, 9:15 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 9:12 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Fixed up an-worker1208

btullis@an-worker1208:~$ sudo blkid|grep -c hadoop-
11
Fri, Aug 14, 9:12 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 9:09 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Fixed up an-worker1207.

Physical Drives = 14
Fri, Aug 14, 9:09 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Fixed up an-worker1205.

Physical Drives = 14
Fri, Aug 14, 9:05 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 9:04 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Fixed up an-worker1201.
Booted into emergency mode.

You are in emergency mode. AfterGive root password for maintenance
(or press Control-D to continue): 
root@an-worker1201:~# lsblk
NAME                           MAJ:MIN RM   SIZE RO TYPE MOUNTPOINT
sda                              8:0    0   7.3T  0 disk 
└─sda1                           8:1    0   7.3T  0 part /var/lib/hadoop/data/l
sdb                              8:16   0   7.3T  0 disk 
└─sdb1                           8:17   0   7.3T  0 part /var/lib/hadoop/data/k
sdc                              8:32   0   7.3T  0 disk 
sdd                              8:48   0   7.3T  0 disk 
└─sdd1                           8:49   0   7.3T  0 part /var/lib/hadoop/data/m
sde                              8:64   0   7.3T  0 disk 
└─sde1                           8:65   0   7.3T  0 part /var/lib/hadoop/data/j
sdf                              8:80   0   7.3T  0 disk 
└─sdf1                           8:81   0   7.3T  0 part /var/lib/hadoop/data/h
sdg                              8:96   0   7.3T  0 disk 
└─sdg1                           8:97   0   7.3T  0 part /var/lib/hadoop/data/g
sdh                              8:112  0   7.3T  0 disk 
└─sdh1                           8:113  0   7.3T  0 part /var/lib/hadoop/data/f
sdi                              8:128  0   7.3T  0 disk 
sdj                              8:144  0   7.3T  0 disk 
└─sdj1                           8:145  0   7.3T  0 part /var/lib/hadoop/data/d
sdk                              8:160  0   7.3T  0 disk 
└─sdk1                           8:161  0   7.3T  0 part /var/lib/hadoop/data/c
sdl                              8:176  0   7.3T  0 disk 
└─sdl1                           8:177  0   7.3T  0 part /var/lib/hadoop/data/b
sdm                              8:192  0 446.6G  0 disk 
├─sdm1                           8:193  0   953M  0 part /boot
├─sdm2                           8:194  0     1K  0 part 
└─sdm5                           8:197  0 445.7G  0 part 
  ├─an--worker1201--vg-swap    254:0    0   9.3G  0 lvm  [SWAP]
  ├─an--worker1201--vg-root    254:1    0  55.9G  0 lvm  /
  └─an--worker1201--vg-journalnode
                               254:2    0    10G  0 lvm  /var/lib/hadoop/journal
Fri, Aug 14, 8:49 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 8:35 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Fixed up an-worker1231

Physical Drives = 14
Fri, Aug 14, 7:48 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Fri, Aug 14, 7:45 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Fri, Aug 14, 7:39 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)

Yesterday

BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 2:59 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Working on an-worker1201.

btullis@an-worker1201:~$ sudo perccli64 /c0 show
Thu, Aug 13, 2:51 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 2:33 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Now working on an-worker1200

Physical Drives = 14
Thu, Aug 13, 2:32 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a parent task for T426610: Follow up on multiple RAID / drive issues: T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 2:24 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a subtask for T434494: Migrate production hadoop cluster to bookworm: T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 2:24 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 2:24 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 2:23 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

Now working on an-worker 1199

Physical Drives = 14
Thu, Aug 13, 2:21 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis triaged T434793: Replace dse-k8s-etcd2001 which was only available on ganeti2046 as High priority.
Thu, Aug 13, 2:10 PM · Patch-For-Review, Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis created T434793: Replace dse-k8s-etcd2001 which was only available on ganeti2046.
Thu, Aug 13, 2:09 PM · Patch-For-Review, Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 1:49 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

an-worker1204 has three drives missing.
I have booted it but then it stopped and I have logged into over the serial console.

Physical Drives = 14
Thu, Aug 13, 1:48 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 1:17 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 12:46 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 12:11 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Thu, Aug 13, 12:10 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

I am working to fix up all of the missing disk issues in T426610: Follow up on multiple RAID / drive issues

Thu, Aug 13, 11:35 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 11:32 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

I did the first of these, an-worker1213.

Thu, Aug 13, 11:32 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops
BTullis claimed T426610: Follow up on multiple RAID / drive issues.
Thu, Aug 13, 10:20 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), SRE, DC-Ops

Wed, Aug 12

BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

We also have 9 servers where there are not 12 data disks showing up.
These need fixing before the reimage cookbook will run correctly.

btullis@cumin1003:~$ sudo cumin A:hadoop-worker 'blkid | grep -c hadoop-'
91 hosts will be targeted:
an-worker[1142-1147,1149-1177,1179-1187,1189-1190,1192-1236].eqiad.wmnet
OK to proceed on 91 hosts? Enter the number of affected hosts to confirm or "q" to quit: 91
===== NODE GROUP =====                                                                                                                                                                                             
(5) an-worker[1200,1207-1208,1213,1231].eqiad.wmnet                                                                                                                                                                
----- OUTPUT for command #1: 'blkid | grep -c hadoop-' -----                                                                                                                                                       
11                                                                                                                                                                                                                 
===== NODE GROUP =====                                                                                                                                                                                             
(1) an-worker1204.eqiad.wmnet                                                                                                                                                                                      
----- OUTPUT for command #1: 'blkid | grep -c hadoop-' -----                                                                                                                                                       
9                                                                                                                                                                                                                  
===== NODE GROUP =====                                                                                                                                                                                             
(3) an-worker[1199,1201,1205].eqiad.wmnet                                                                                                                                                                          
----- OUTPUT for command #1: 'blkid | grep -c hadoop-' -----                                                                                                                                                       
10                                                                                                                                                                                                                 
===== NODE GROUP =====                                                                                                                                                                                             
(82) an-worker[1142-1147,1149-1177,1179-1187,1189-1190,1192-1198,1202-1203,1206,1209-1212,1214-1230,1232-1236].eqiad.wmnet                                                                                         
----- OUTPUT for command #1: 'blkid | grep -c hadoop-' -----                                                                                                                                                       
12                                                                                                                                                                                                                 
================
Wed, Aug 12, 5:01 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Wed, Aug 12, 3:35 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Wed, Aug 12, 3:12 PM · Data-Platform-SRE, Epic
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Wed, Aug 12, 12:07 PM · Data-Platform-SRE, Epic
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

We are down to 100 hosts in the analytics and presto clusters.

image.png (1,905×632 px, 104 KB)

Wed, Aug 12, 11:33 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.

I'm going to do an in-place upgrade for Archiva, as it is a somewhat problematic VM with large data stores, that we are hoping to kill soon.
I don't want to spend the time getting a reuse-parts recipe working for this one-off host.

Wed, Aug 12, 9:32 AM · Data-Platform-SRE, Epic
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Wed, Aug 12, 8:38 AM · Data-Platform-SRE, Epic
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Wed, Aug 12, 8:38 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Wed, Aug 12, 8:36 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Wed, Aug 12, 7:53 AM · Data-Platform-SRE, Epic
BTullis updated the task description for T433495: decommission an-test-master100[1-2].
Wed, Aug 12, 7:31 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware

Tue, Aug 11

BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Tue, Aug 11, 2:13 PM · Data-Platform-SRE, Epic
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

I am going to start upgrading the nameservers, starting with the standby, an-master1004.

Tue, Aug 11, 2:07 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Tue, Aug 11, 2:00 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T432766: Permission requested to kick off AirFlow jobs for Tchanders.

Apologies for the delay.
I have now added you to the airflow-analytics-ops LDAP group, as discussed here.

Tue, Aug 11, 1:56 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Tue, Aug 11, 1:31 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

It looks like 4 of the 15 hosts were detected this way around, so maybe it should work this way around.

btullis@cumin1003:~$ sudo cumin A:an-presto pvs
14 hosts will be targeted:
an-presto[1006-1007,1009-1020].eqiad.wmnet
OK to proceed on 14 hosts? Enter the number of affected hosts to confirm or "q" to quit: 14
===== NODE GROUP =====                                                                                                                                                                                             
(3) an-presto[1011,1014,1016].eqiad.wmnet                                                                                                                                                                          
----- OUTPUT for command #1: 'pvs' -----                                                                                                                                                                           
  PV         VG  Fmt  Attr PSize   PFree                                                                                                                                                                           
  /dev/sda1  vg1 lvm2 a--  <21.83t <4.37t
  /dev/sdb2  vg0 lvm2 a--  446.34g 89.27g
===== NODE GROUP =====                                                                                                                                                                                             
(11) an-presto[1006-1007,1009-1010,1012-1013,1015,1017-1020].eqiad.wmnet                                                                                                                                           
----- OUTPUT for command #1: 'pvs' -----                                                                                                                                                                           
  PV         VG  Fmt  Attr PSize   PFree                                                                                                                                                                           
  /dev/sda2  vg0 lvm2 a--  446.34g 89.27g
  /dev/sdb1  vg1 lvm2 a--  <21.83t <4.37t
================

I might be able to force the detection order by deleting and recreating the RAID volumes, or running the sre.hosts.provision cookbook on them, perhaps.

Tue, Aug 11, 11:18 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis added a comment to T434494: Migrate production hadoop cluster to bookworm.

an-presto1008 is giving a partman error.

image.png (898×541 px, 74 KB)

It seems to be because the large RAID disk has been detected as /dev/sda, rather than /dev/sdb

~ # pvs
  PV         VG  Fmt  Attr PSize   PFree  
  /dev/sda2  vg0 lvm2 a--  <21.83t  <4.37t
  /dev/sdb1  vg1 lvm2 a--  446.62g 446.62g
Tue, Aug 11, 11:09 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T434494: Migrate production hadoop cluster to bookworm.
Tue, Aug 11, 9:34 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Tue, Aug 11, 9:34 AM · Data-Platform-SRE, Epic
BTullis moved T434494: Migrate production hadoop cluster to bookworm from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-07 - 2026-08-28) board.
Tue, Aug 11, 9:29 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis triaged T434494: Migrate production hadoop cluster to bookworm as High priority.
Tue, Aug 11, 9:29 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Tue, Aug 11, 9:28 AM · Data-Platform-SRE, Epic
BTullis created T434494: Migrate production hadoop cluster to bookworm.
Tue, Aug 11, 9:27 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28)
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Tue, Aug 11, 8:59 AM · Data-Platform-SRE, Epic
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Tue, Aug 11, 8:26 AM · Data-Platform-SRE, Epic
BTullis closed T433495: decommission an-test-master100[1-2] as Resolved.
Tue, Aug 11, 8:25 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis closed T433495: decommission an-test-master100[1-2], a subtask of T406746: Ensure that our Hadoop stack works on bookworm, as Resolved.
Tue, Aug 11, 8:25 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis closed T433494: decommission an-test-coord1001.eqiad.wmnet as Resolved.
Tue, Aug 11, 8:25 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis closed T433494: decommission an-test-coord1001.eqiad.wmnet, a subtask of T406746: Ensure that our Hadoop stack works on bookworm, as Resolved.
Tue, Aug 11, 8:25 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis updated the task description for T433494: decommission an-test-coord1001.eqiad.wmnet.
Tue, Aug 11, 8:24 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T434445: Improve app instrument event data data lake management.
Tue, Aug 11, 8:23 AM · Patch-For-Review, Test Kitchen (Experiment Platform Sprint 28)
BTullis added a comment to T434445: Improve app instrument event data data lake management.

I have created the directory, as requested.

btullis@an-launcher1003:~$ sudo -u hdfs kerberos-run-command hdfs hdfs dfs -mkdir /wmf/data/wmf_apps
Tue, Aug 11, 8:23 AM · Patch-For-Review, Test Kitchen (Experiment Platform Sprint 28)

Wed, Jul 29

BTullis updated the task description for T433495: decommission an-test-master100[1-2].
Wed, Jul 29, 4:19 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 4:18 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T433495: decommission an-test-master100[1-2].
Wed, Jul 29, 4:18 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis placed T433495: decommission an-test-master100[1-2] up for grabs.

I still have a puppet cleanup patch to supply and some keytabs to remove, but please feel free to proceed with de-racking and disposing of these servers.

Wed, Jul 29, 4:17 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis placed T433494: decommission an-test-coord1001.eqiad.wmnet up for grabs.

I still have a puppet cleanup patch to supply and some keytabs to remove, but please feel free to proceed with de-racking and disposing of this server.

Wed, Jul 29, 4:17 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a comment to T433240: No new wikibase dump downloads being published..

I have modified the title to reflect that this is wikibase dumps, as opposed to XML/SQL dumps, or any of the other types.

Wed, Jul 29, 3:17 PM · Data-Platform-SRE, Data-Engineering (Q1 FS26/27 July 1st - September 30th), Wikidata data dumps, Dumps-Generation, Wikidata
BTullis renamed T433240: No new wikibase dump downloads being published. from No new dump downloads being published. to No new wikibase dump downloads being published..
Wed, Jul 29, 3:15 PM · Data-Platform-SRE, Data-Engineering (Q1 FS26/27 July 1st - September 30th), Wikidata data dumps, Dumps-Generation, Wikidata
BTullis moved T433495: decommission an-test-master100[1-2] from Backlog - project to In Progress on the Data-Platform-SRE (2026-07-03 - 2026-07-31) board.
Wed, Jul 29, 1:44 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a project to T433495: decommission an-test-master100[1-2]: Data-Platform-SRE (2026-07-03 - 2026-07-31).
Wed, Jul 29, 1:44 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis moved T433494: decommission an-test-coord1001.eqiad.wmnet from Backlog - project to In Progress on the Data-Platform-SRE (2026-07-03 - 2026-07-31) board.
Wed, Jul 29, 1:43 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a project to T433494: decommission an-test-coord1001.eqiad.wmnet: Data-Platform-SRE (2026-07-03 - 2026-07-31).
Wed, Jul 29, 1:43 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 1:43 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis claimed T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 1:42 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T433495: decommission an-test-master100[1-2].
Wed, Jul 29, 1:42 PM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a comment to T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format.

I see that the log files on the relforge cluster seem to be generating ECS fortmatted logs now.
Thanks @bking for checking this out.

btullis@relforge1010:~$ tail -n 1 -f /var/log/opensearch/relforge-eqiad_server.json | jq . --unbuffered
{
  "@timestamp": "2026-07-28T22:16:43.396Z",
  "log.level": "INFO",
  "message": "finalizing recovery took [12.3ms]",
  "ecs.version": "1.2.0",
  "service.name": "relforge-eqiad",
  "event.dataset": "relforge-eqiad",
  "process.thread.name": "opensearch[relforge1010-relforge-eqiad][generic][T#1]",
  "log.logger": "org.opensearch.indices.recovery.RecoverySourceHandler"
}

Was anything required other than a rolling restart of the servcies, or did this work without any issue?

Wed, Jul 29, 11:58 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review, Observability-Logging
BTullis added a subtask for T406746: Ensure that our Hadoop stack works on bookworm: T433495: decommission an-test-master100[1-2].
Wed, Jul 29, 11:55 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a parent task for T433495: decommission an-test-master100[1-2]: T406746: Ensure that our Hadoop stack works on bookworm.
Wed, Jul 29, 11:55 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis created T433495: decommission an-test-master100[1-2].
Wed, Jul 29, 11:54 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a subtask for T406746: Ensure that our Hadoop stack works on bookworm: T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 11:52 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a parent task for T433494: decommission an-test-coord1001.eqiad.wmnet: T406746: Ensure that our Hadoop stack works on bookworm.
Wed, Jul 29, 11:52 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis updated the task description for T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 11:52 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis created T433494: decommission an-test-coord1001.eqiad.wmnet.
Wed, Jul 29, 11:49 AM · SRE, DC-Ops, Data-Platform-SRE (2026-07-03 - 2026-07-31), decommission-hardware
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.

The offending call to python is here.

btullis@an-test-coord1002:~$ head /usr/lib/presto/bin/launcher.py
#!/usr/bin/env python
Wed, Jul 29, 11:20 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.

Oh, it turns out that there is another small issue with presto.

Jul 29 11:03:03 an-test-coord1002 presto-server[2265682]: /usr/bin/env: ‘python’: No such file or directory
Jul 29 11:03:03 an-test-coord1002 systemd[1]: presto-server.service: Main process exited, code=exited, status=127/n/a
Jul 29 11:03:03 an-test-coord1002 systemd[1]: presto-server.service: Failed with result 'exit-code'.

I might need to rebuild the package for this, but let's see.

Wed, Jul 29, 11:04 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.
root@krb1002:~# kadmin.local ktadd -norandkey -k /srv/kerberos/keytabs/an-test-coord1002.eqiad.wmnet/hive/hive.keytab hive/analytics-test-hive.eqiad.wmnet@WIKIMEDIA
Entry for principal hive/analytics-test-hive.eqiad.wmnet@WIKIMEDIA with kvno 1, encryption type aes256-cts-hmac-sha1-96 added to keytab WRFILE:/srv/kerberos/keytabs/an-test-coord1002.eqiad.wmnet/hive/hive.keytab.
Wed, Jul 29, 10:23 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.

There's a small issue with starting the hive-metastore and presto-server services on the refreshed an-coord1002 host, but this is because I didn't add the required primitives to the keytabs.

Wed, Jul 29, 10:16 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis updated the task description for T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later.
Wed, Jul 29, 10:08 AM · Data-Platform-SRE, Epic
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.

Now I am about to switch the MariaDAB database on an-test-coord1002 to be a standalone master, instead of a replication slave of an-test-coord1001.

Wed, Jul 29, 10:02 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.

Merged all of the patches.

Wed, Jul 29, 9:01 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review
BTullis added a comment to T406746: Ensure that our Hadoop stack works on bookworm.
btullis@an-test-master1001:~$ sudo -u hdfs kerberos-run-command hdfs hdfs dfsadmin -safemode enter
Safe mode is ON in an-test-master1001.eqiad.wmnet/10.64.5.39:8020
Safe mode is ON in an-test-master1002.eqiad.wmnet/10.64.36.112:8020

Now transitioning to standby.

btullis@an-test-master1001:~$ sudo -u hdfs hdfs haadmin -transitionToStandby an-test-master1001-eqiad-wmnet --forcemanual
You have specified the --forcemanual flag. This flag is dangerous, as it can induce a split-brain scenario that WILL CORRUPT your HDFS namespace, possibly irrecoverably.
Wed, Jul 29, 8:59 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review