Page MenuHomePhabricator

BTullis (Ben)
Staff SRE

Today

  • No visible events.

Tomorrow

  • No visible events.

Monday

  • No visible events.

User Details

User Since
Jun 29 2021, 9:56 AM (271 w, 4 d)
Availability
Available
IRC Nick
btullis
LDAP User
Btullis
MediaWiki User
BTullis (WMF) [ Global Accounts ]

Recent Activity

Yesterday

BTullis updated the task description for T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k.
Fri, Sep 11, 4:01 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis updated the task description for T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material.
Fri, Sep 11, 3:32 PM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis triaged T437717: ceph: Delete the unused CephX entities from the Data Platform clusters as Medium priority.
Fri, Sep 11, 2:34 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis created T437717: ceph: Delete the unused CephX entities from the Data Platform clusters.
Fri, Sep 11, 2:31 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis moved T437706: Liftwing studio chart from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Fri, Sep 11, 1:51 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Patch-For-Review
BTullis edited projects for T437706: Liftwing studio chart, added: Data-Platform-SRE (2026-08-28 - 2026-09-18); removed Data-Platform-SRE.
Fri, Sep 11, 1:51 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Patch-For-Review
BTullis updated the task description for T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k.
Fri, Sep 11, 12:46 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis removed a subtask for T266640: Decide whether to migrate from Presto to Trino: T324011: SPIKE: Spin up a Test Trino instance (Evaluate Trino).
Fri, Sep 11, 10:40 AM · Data-Platform-SRE
BTullis removed a subtask for T311738: [Iceberg] Debianize and install iceberg support for Spark, Presto, and optionally Hive: T324011: SPIKE: Spin up a Test Trino instance (Evaluate Trino).
Fri, Sep 11, 10:40 AM · Shared-Data-Infrastructure, Data Pipelines, Data-Engineering-Planning
BTullis removed parent tasks for T324011: SPIKE: Spin up a Test Trino instance (Evaluate Trino): T266640: Decide whether to migrate from Presto to Trino, T311738: [Iceberg] Debianize and install iceberg support for Spark, Presto, and optionally Hive.
Fri, Sep 11, 10:40 AM · Data-Platform-SRE
BTullis added a parent task for T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material: T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k.
Fri, Sep 11, 10:03 AM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a subtask for T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k: T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material.
Fri, Sep 11, 10:03 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis moved T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Fri, Sep 11, 10:01 AM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis triaged T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material as High priority.
Fri, Sep 11, 9:50 AM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material.

Here is the output when running the script on the two WMCS clusters.

Fri, Sep 11, 9:46 AM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material.

I have written and deployed the check script: /usr/local/sbin/verify-cephx-keys
It is deployed on these hosts:

Fri, Sep 11, 9:44 AM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)

Thu, Sep 10

BTullis added a comment to T437170: Degraded RAID on an-worker1200.

I think that this is is another one of those times when the disk hasn't actually failed, but it was disconnected by the RAID controller, temporarily.

Thu, Sep 10, 11:20 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), DC-Ops, SRE, ops-eqiad
BTullis moved T435044: [Epic] Improve defense in depth for higher-risk payloads running on kubernetes from Incoming to Watching on the Data-Platform-SRE board.
Thu, Sep 10, 6:19 AM · Data-Platform-SRE, ServiceOps (Next quarter), Kubernetes, Security, Epic, Prod-Kubernetes
BTullis added a project to T435044: [Epic] Improve defense in depth for higher-risk payloads running on kubernetes: Data-Platform-SRE.
Thu, Sep 10, 6:18 AM · Data-Platform-SRE, ServiceOps (Next quarter), Kubernetes, Security, Epic, Prod-Kubernetes
BTullis added a comment to T437170: Degraded RAID on an-worker1200.

This one keeps dropping disks. I've added the parent task where I'm doing a larger investigation into what seems like a batch of servers where this keeps happening.

Thu, Sep 10, 6:07 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), DC-Ops, SRE, ops-eqiad
BTullis added a parent task for T437170: Degraded RAID on an-worker1200: T426610: Follow up on multiple RAID / drive issues.
Thu, Sep 10, 6:05 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), DC-Ops, SRE, ops-eqiad
BTullis added a subtask for T426610: Follow up on multiple RAID / drive issues: T437170: Degraded RAID on an-worker1200.
Thu, Sep 10, 6:05 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, DC-Ops

Wed, Sep 9

BTullis triaged T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k as High priority.
Wed, Sep 9, 2:02 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis created T437447: ceph: Rotate the CephX keys on both Data Platform clusters to aes256k.
Wed, Sep 9, 1:27 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis triaged T426610: Follow up on multiple RAID / drive issues as Low priority.
Wed, Sep 9, 8:00 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, DC-Ops
BTullis updated the task description for T432206: Bring dse-k8s-worker[2004-2005] into service.
Wed, Sep 9, 7:59 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)

Tue, Sep 8

BTullis closed T432206: Bring dse-k8s-worker[2004-2005] into service as Resolved.

The workers are in their production role, puppet runs cleanly, and they have been uncordoned.

root@deploy1003:~# kubectl get nodes
NAME                                STATUS                     ROLES           AGE    VERSION
dse-k8s-ctrl2001.codfw.wmnet        Ready                      control-plane   358d   v1.31.14
dse-k8s-ctrl2002.codfw.wmnet        Ready                      control-plane   358d   v1.31.14
dse-k8s-wdqs-test2001.codfw.wmnet   Ready                      <none>          99d    v1.31.14
dse-k8s-wdqs2001.codfw.wmnet        Ready                      <none>          82d    v1.31.14
dse-k8s-wdqs2002.codfw.wmnet        Ready                      <none>          82d    v1.31.14
dse-k8s-wdqs2003.codfw.wmnet        Ready                      <none>          82d    v1.31.14
dse-k8s-wdqs2004.codfw.wmnet        Ready                      <none>          82d    v1.31.14
dse-k8s-worker2001.codfw.wmnet      Ready                      <none>          358d   v1.31.14
dse-k8s-worker2002.codfw.wmnet      Ready                      <none>          358d   v1.31.14
dse-k8s-worker2003.codfw.wmnet      Ready                      <none>          333d   v1.31.14
dse-k8s-worker2004.codfw.wmnet      Ready,SchedulingDisabled   <none>          20m    v1.31.14
dse-k8s-worker2005.codfw.wmnet      Ready,SchedulingDisabled   <none>          14m    v1.31.14
root@deploy1003:~# kubectl uncordon dse-k8s-worker2004.codfw.wmnet dse-k8s-worker2005.codfw.wmnet 
node/dse-k8s-worker2004.codfw.wmnet uncordoned
node/dse-k8s-worker2005.codfw.wmnet uncordoned
root@deploy1003:~# kubectl get nodes
NAME                                STATUS   ROLES           AGE    VERSION
dse-k8s-ctrl2001.codfw.wmnet        Ready    control-plane   358d   v1.31.14
dse-k8s-ctrl2002.codfw.wmnet        Ready    control-plane   358d   v1.31.14
dse-k8s-wdqs-test2001.codfw.wmnet   Ready    <none>          99d    v1.31.14
dse-k8s-wdqs2001.codfw.wmnet        Ready    <none>          82d    v1.31.14
dse-k8s-wdqs2002.codfw.wmnet        Ready    <none>          82d    v1.31.14
dse-k8s-wdqs2003.codfw.wmnet        Ready    <none>          82d    v1.31.14
dse-k8s-wdqs2004.codfw.wmnet        Ready    <none>          82d    v1.31.14
dse-k8s-worker2001.codfw.wmnet      Ready    <none>          358d   v1.31.14
dse-k8s-worker2002.codfw.wmnet      Ready    <none>          358d   v1.31.14
dse-k8s-worker2003.codfw.wmnet      Ready    <none>          333d   v1.31.14
dse-k8s-worker2004.codfw.wmnet      Ready    <none>          26m    v1.31.14
dse-k8s-worker2005.codfw.wmnet      Ready    <none>          20m    v1.31.14
root@deploy1003:~#
Tue, Sep 8, 5:19 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis updated the task description for T432206: Bring dse-k8s-worker[2004-2005] into service.
Tue, Sep 8, 4:26 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T432206: Bring dse-k8s-worker[2004-2005] into service.

Added the nodes to conftool-data and pooled them.

btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw' get|sort
{"dse-k8s-wdqs2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2002.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2003.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2004.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs-test2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2002.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2003.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2004.codfw.wmnet": {"weight": 0, "pooled": "inactive"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2005.codfw.wmnet": {"weight": 0, "pooled": "inactive"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet' set/weight=10
codfw/dse-k8s/kubesvc/dse-k8s-worker2004.codfw.wmnet: weight changed 0 => 10
WARNING:conftool.announce:conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet' set/weight=10
codfw/dse-k8s/kubesvc/dse-k8s-worker2005.codfw.wmnet: weight changed 0 => 10
WARNING:conftool.announce:conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet' set/pooled=yes
codfw/dse-k8s/kubesvc/dse-k8s-worker2004.codfw.wmnet: pooled changed inactive => yes
WARNING:conftool.announce:conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet' set/pooled=yes
codfw/dse-k8s/kubesvc/dse-k8s-worker2005.codfw.wmnet: pooled changed inactive => yes
WARNING:conftool.announce:conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
btullis@puppetserver1001:~$ sudo confctl select 'service=kubesvc,cluster=dse-k8s,dc=codfw' get|sort
{"dse-k8s-wdqs2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2002.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2003.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs2004.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-wdqs-test2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2001.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2002.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2003.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2004.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
{"dse-k8s-worker2005.codfw.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=codfw,cluster=dse-k8s,service=kubesvc"}
btullis@puppetserver1001:~$
Tue, Sep 8, 4:26 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

I restarted all radosgw daemons and everything seems OK.

Tue, Sep 8, 4:15 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Disabling standby replay.

btullis@cephosd2001:~$ sudo ceph fs ls
name: dpe, metadata pool: cephfs.dpe.meta, data pools: [cephfs.dpe.data-ssd ]
Tue, Sep 8, 1:00 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

I have restarted all OSD daemons by running the following an all three hosts, several minutes apart.

sudo systemctl restart ceph-osd.target

We can see the new versions being reported here: https://grafana-rw.wikimedia.org/goto/s4mw5r?orgId=default

image.png (931×374 px, 37 KB)

Tue, Sep 8, 12:52 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Restarting the mgr daemons on each host.

Tue, Sep 8, 11:39 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Going for it now.

Tue, Sep 8, 11:35 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Since ours is a co-located installation and all daemons are on the same hosts, it is not possible to defer the installation of the packages. When I try to install the ceph-mon package, all of the other relevant daemons are upgraded, too.

Tue, Sep 8, 11:18 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

The procedure set out here https://docs.ceph.com/en/latest/releases/squid/#upgrading-non-cephadm-clusters is broadly as follows:

Tue, Sep 8, 11:15 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T426610: Follow up on multiple RAID / drive issues.

A few weeks later, after our upgrade to bookworm on all of these an-worker hosts, I run the command again to see how many hosts have fewer than 12 hadoop-* data volumes.

btullis@cumin1003:~$ sudo cumin A:hadoop-worker 'blkid |grep -c hadoop-'
94 hosts will be targeted:
an-worker[1142-1147,1149-1236].eqiad.wmnet
OK to proceed on 94 hosts? Enter the number of affected hosts to confirm or "q" to quit: 94
===== NODE GROUP =====                                                                                                                                                                                             
(1) an-worker1200.eqiad.wmnet                                                                                                                                                                                      
----- OUTPUT for command #1: 'blkid |grep -c hadoop-' -----                                                                                                                                                        
10                                                                                                                                                                                                                 
===== NODE GROUP =====                                                                                                                                                                                             
(4) an-worker[1144,1198-1199,1204].eqiad.wmnet                                                                                                                                                                     
----- OUTPUT for command #1: 'blkid |grep -c hadoop-' -----                                                                                                                                                        
11                                                                                                                                                                                                                 
===== NODE GROUP =====                                                                                                                                                                                             
(89) an-worker[1142-1143,1145-1147,1149-1197,1201-1203,1205-1236].eqiad.wmnet                                                                                                                                      
----- OUTPUT for command #1: 'blkid |grep -c hadoop-' -----                                                                                                                                                        
12                                                                                                                                                                                                                 
================

5 of the 94 hadoop workers are currently showing as having dropped at least one disk.

Tue, Sep 8, 9:45 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, DC-Ops

Mon, Sep 7

BTullis updated the task description for T432206: Bring dse-k8s-worker[2004-2005] into service.
Mon, Sep 7, 3:14 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis triaged T432206: Bring dse-k8s-worker[2004-2005] into service as Medium priority.

Here is the homer diff, after enabling BGP for the two hosts.

btullis@cumin1003:~$ sudo homer lsw1-a?-codfw* diff
INFO:homer.devices:Initialized 129 devices
INFO:homer:Generating diff for query lsw1-a?-codfw*
INFO:homer:Gathering global Netbox data
INFO:homer.devices:Matched 7 device(s) for query 'lsw1-a?-codfw*'
INFO:homer:Generating configuration for lsw1-a2-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Running commit check on lsw1-a2-codfw.mgmt.codfw.wmnet
INFO:homer:Generating configuration for lsw1-a3-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Empty diff for lsw1-a3-codfw.mgmt.codfw.wmnet, skipping device.
INFO:homer:Generating configuration for lsw1-a4-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Empty diff for lsw1-a4-codfw.mgmt.codfw.wmnet, skipping device.
INFO:homer:Generating configuration for lsw1-a5-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Empty diff for lsw1-a5-codfw.mgmt.codfw.wmnet, skipping device.
INFO:homer:Generating configuration for lsw1-a6-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Empty diff for lsw1-a6-codfw.mgmt.codfw.wmnet, skipping device.
INFO:homer:Generating configuration for lsw1-a7-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Running commit check on lsw1-a7-codfw.mgmt.codfw.wmnet
INFO:homer:Generating configuration for lsw1-a8-codfw.mgmt.codfw.wmnet
INFO:homer.transports.junos:Empty diff for lsw1-a8-codfw.mgmt.codfw.wmnet, skipping device.
Changes for 1 devices: ['lsw1-a2-codfw.mgmt.codfw.wmnet']
Mon, Sep 7, 3:14 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

I have created T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material to fix the puppet management of CephX key data.

Mon, Sep 7, 3:01 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis created T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material.
Mon, Sep 7, 2:59 PM · Patch-For-Review, tools-infrastructure-team, Ceph, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

I don't think that the systemd unit name change is going to bite us, after all.
I still see the existing names of systemd units in https://github.com/ceph/ceph/tree/v19.2.0/systemd so I think that this actually affects cephadm clusters. Maybe it's a documentation bug.

Mon, Sep 7, 2:47 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis moved T399594: Proposed improvement: Manage CephX users via exported/collected Puppet resources from Backlog - project to Done on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Mon, Sep 7, 2:24 PM · tools-infrastructure-team, Data-Platform-SRE (2026-08-28 - 2026-09-18), Cloud-VPS, Ceph, cloud-services-team
BTullis closed T399594: Proposed improvement: Manage CephX users via exported/collected Puppet resources as Declined.

I think that I'm going to decline this one, having looked at it again in the context of T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Mon, Sep 7, 2:24 PM · tools-infrastructure-team, Data-Platform-SRE (2026-08-28 - 2026-09-18), Cloud-VPS, Ceph, cloud-services-team
BTullis added a comment to T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake.
  • Trying to precise the sentence "This is the closest match in shape to what Sqoop does for us." in the SeaTunnel section: To my understanding, SeaTunnel could handle both the data-movement aspect (sqoop) and the multi-tables/schema jobs and transforms (sqoop management script) using configuration files. If we were to use spark, it would be a replacement for sqoop itself and we would need to add an equivalent of the sqoop management script.

Thanks, yes I agree that this was a bit wooly and vague. How does this version of the SeaTunnel description sound, instead?

Mon, Sep 7, 1:02 PM · Data-Engineering, Data-Platform-SRE
BTullis updated the task description for T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake.
Mon, Sep 7, 12:52 PM · Data-Engineering, Data-Platform-SRE
BTullis updated the task description for T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake.
Mon, Sep 7, 12:51 PM · Data-Engineering, Data-Platform-SRE
BTullis closed T437206: Error from the cephosd cluster in eqiad - 1 scrub errors; Possible data damage: 1 pg inconsistent as Resolved.

This is now OK.

btullis@cephosd1001:~$ sudo ceph health detail
HEALTH_OK

I see this from the journal entry for the ceph-osd@80.service until on cephosd1005.

Sep 07 11:03:49 cephosd1005 ceph-osd[4055]: 2026-09-07T11:03:49.206+0000 7fc468dd26c0 -1 log_channel(cluster) log [ERR] : 19.3b0 repair 1 errors, 1 fixed

Not investigating any further, so I'll resolve this ticket.

Mon, Sep 7, 11:10 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T437206: Error from the cephosd cluster in eqiad - 1 scrub errors; Possible data damage: 1 pg inconsistent.

I have just executed this command.

btullis@cephosd1001:~$ sudo ceph pg repair 19.3b0
instructing pg 19.3b0 on osd.80 to repair
Mon, Sep 7, 10:39 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

Oh, this is also interesting. From here: https://docs.ceph.com/en/latest/releases/squid/#upgrading-non-cephadm-clusters

Mon, Sep 7, 10:23 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis added a comment to T428445: Upgrade Data Platform ceph clusters to version 19 - squid.

I'm reviewing the notable changes and critical upgrade steps for this upgrade.

Mon, Sep 7, 10:01 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis placed T436964: Q1:rack/setup/install dse-k8s-worker10[29-38] up for grabs.

I have updated preseed.yaml and site.pp for these new nodes, so I think that I'm done for now.

Mon, Sep 7, 9:47 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, ops-eqiad, DC-Ops
BTullis triaged T437206: Error from the cephosd cluster in eqiad - 1 scrub errors; Possible data damage: 1 pg inconsistent as Medium priority.
Mon, Sep 7, 9:40 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis created T437206: Error from the cephosd cluster in eqiad - 1 scrub errors; Possible data damage: 1 pg inconsistent.
Mon, Sep 7, 9:40 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis moved T436964: Q1:rack/setup/install dse-k8s-worker10[29-38] from In Progress to Done on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Mon, Sep 7, 9:03 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, ops-eqiad, DC-Ops

Sun, Sep 6

BTullis moved T436913: 7za is OOMkilled when dumps are running on a Bookworm-based image from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Sun, Sep 6, 2:49 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Data-Engineering, Dumps-Generation
BTullis edited projects for T436913: 7za is OOMkilled when dumps are running on a Bookworm-based image, added: Data-Platform-SRE (2026-08-28 - 2026-09-18); removed Data-Platform-SRE.
Sun, Sep 6, 2:49 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Data-Engineering, Dumps-Generation
BTullis moved T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake from Incoming to Quarterly Goals on the Data-Platform-SRE board.
Sun, Sep 6, 2:46 AM · Data-Engineering, Data-Platform-SRE
BTullis moved T437080: [Tech evaluation] How we ingest the Kafka streams into the S3 data lake from Incoming to Quarterly Goals on the Data-Platform-SRE board.
Sun, Sep 6, 2:45 AM · Data-Engineering, Data-Platform-SRE
BTullis edited projects for T435591: Allow unauthenticated access to the dse-k8s OIDC discovery and JWKS endpoints, added: Data-Platform-SRE (2026-08-28 - 2026-09-18); removed Data-Platform-SRE.
Sun, Sep 6, 2:02 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Data-Engineering

Fri, Sep 4

BTullis updated the task description for T435489: Data Lake ingestion after Hadoop: decide the future of Sqoop and Gobblin.
Fri, Sep 4, 5:35 PM · Data-Engineering, Data-Platform-SRE
BTullis updated the task description for T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake.
Fri, Sep 4, 5:32 PM · Data-Engineering, Data-Platform-SRE
BTullis created T437080: [Tech evaluation] How we ingest the Kafka streams into the S3 data lake.
Fri, Sep 4, 5:32 PM · Data-Engineering, Data-Platform-SRE
BTullis created T437079: [Tech evaluation] How we import the MediaWiki databases into the S3 data lake.
Fri, Sep 4, 5:19 PM · Data-Engineering, Data-Platform-SRE
BTullis moved T436964: Q1:rack/setup/install dse-k8s-worker10[29-38] from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Fri, Sep 4, 12:14 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, ops-eqiad, DC-Ops
BTullis edited projects for T436964: Q1:rack/setup/install dse-k8s-worker10[29-38], added: Data-Platform-SRE (2026-08-28 - 2026-09-18); removed Data-Platform-SRE.
Fri, Sep 4, 12:13 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), SRE, ops-eqiad, DC-Ops
BTullis claimed T432206: Bring dse-k8s-worker[2004-2005] into service.
Fri, Sep 4, 12:09 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis assigned T435604: [Tech evaluation] Table catalog options for the S3 data lake to JAllemandou.

Assigning to @JAllemandou for guidance at this stage.

Fri, Sep 4, 9:23 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Data-Engineering

Thu, Sep 3

BTullis closed T425087: Send JSON access logs for dumps.wikimedia.org to Kafka as Resolved.

I think that this is all done now. I have removed the development schema and the eventstream config for it.

Thu, Sep 3, 3:54 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis closed T425087: Send JSON access logs for dumps.wikimedia.org to Kafka, a subtask of T422880: Dumps user identification monitoring, as Resolved.
Thu, Sep 3, 3:54 PM · Patch-For-Review, Product-Analytics, MediaWiki-API-Platform-Team
BTullis updated the task description for T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.
Thu, Sep 3, 3:52 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis added a comment to T434046: CDN caching request for TTS v1 audio files.

As I understand it (from T435952 and T436758) , the back-end service hosting these files for the period of the experiment will be: analytics.wikimedia.org

Thu, Sep 3, 1:14 PM · Traffic, Machine-Learning-Team (Q1 FY2026-27)
BTullis added a comment to T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.

We now have data in the hive table.

spark-sql (default)> select count(*) from event.webrequest_dumps_v1 where year=2026 and month=9 and day=3;
count(1)
149722
Time taken: 12.268 seconds, Fetched 1 row(s)
Thu, Sep 3, 11:54 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis claimed T428445: Upgrade Data Platform ceph clusters to version 19 - squid.
Thu, Sep 3, 11:34 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis moved T436003: Migrate Matomo frontend to Kubernetes from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Thu, Sep 3, 11:33 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis moved T436186: Matomo K8s migration: Create Helm chart from Backlog - project to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Thu, Sep 3, 11:33 AM · Patch-For-Review, Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.

I checked on https://stream-internal.wikimedia.org/v2/ui/#/?streams=webrequest.dumps.v1 and we have live data in eventstreams from the dumps servers.
Also confirmed that we have a continuous dataset in logstash.

image.png (1,907×546 px, 120 KB)

Thu, Sep 3, 10:28 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services

Wed, Sep 2

BTullis added a comment to T435608: Create a Ceph administration VM and its Puppet role.

Naive q: does this need to be a VM or could it be a pod running in Kubernetes? Alternative idea, an operator (such as https://github.com/snapp-incubator/ceph-s3-operator) in charge of creating the users, roles and buckets?

Wed, Sep 2, 5:22 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis updated the task description for T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.
Wed, Sep 2, 3:41 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis added a comment to T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.

The new eventstreams are being ingested by gobblin.

image.png (1,907×831 px, 185 KB)

https://grafana-rw.wikimedia.org/goto/shfcn5?orgId=default

Wed, Sep 2, 3:03 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis moved T435952: Set up a VM with nginx to store and serve TTS v1 audio files as the CDN origin from Backlog - project to Done on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Wed, Sep 2, 1:08 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Machine-Learning-Team (Q1 FY2026-27)
BTullis edited projects for T435952: Set up a VM with nginx to store and serve TTS v1 audio files as the CDN origin, added: Data-Platform-SRE (2026-08-28 - 2026-09-18); removed Data-Platform-SRE.
Wed, Sep 2, 1:08 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Machine-Learning-Team (Q1 FY2026-27)
BTullis moved T434494: Migrate production hadoop cluster to bookworm from Backlog - project to Reported on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Wed, Sep 2, 12:33 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis closed T434494: Migrate production hadoop cluster to bookworm, a subtask of T401692: EPIC: Migrate Data Platform SRE-owned hosts to Bookworm or later, as Resolved.
Wed, Sep 2, 12:33 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Epic
BTullis closed T434494: Migrate production hadoop cluster to bookworm as Resolved.

an-presto1010 is still on Bullseye. The server has been garbage-collected from puppetdb in the mean time, but I logged in over the serial console to check the server's status:

For more info: https://wikitech.wikimedia.org/wiki/Analytics/Systems/Kerberos/UserGuide
The last Puppet run was at Tue Aug 11 13:58:00 UTC 2026 (18606 minutes ago). Puppet is disabled. Host reimage - btullis@cumin1003 - T434494
Last Puppet commit: (1425bddf65) Janis Meybohm - wikikube: Update staging eqiad to calico 3.30
Debian GNU/Linux 11 auto-installed on Fri Jul 1 21:45:20 UTC 2022.
Wed, Sep 2, 12:33 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis added a comment to T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.

Canary testing with Wikimedia Debug extension looks good to me.

image.png (1,016×706 px, 144 KB)

Wed, Sep 2, 12:22 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis moved T425087: Send JSON access logs for dumps.wikimedia.org to Kafka from Blocked/Waiting to In Progress on the Data-Platform-SRE (2026-08-28 - 2026-09-18) board.
Wed, Sep 2, 12:13 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis added a comment to T425087: Send JSON access logs for dumps.wikimedia.org to Kafka.

I have scheduled a backport deployment of the eventstreams patch for this afternoon.

Wed, Sep 2, 12:03 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-platform-team, Patch-For-Review, cloud-services-team, Data-Services
BTullis added a comment to T435952: Set up a VM with nginx to store and serve TTS v1 audio files as the CDN origin.

There are some instructions on how to publish to analytics.wikimedia.org here: https://wikitech.wikimedia.org/wiki/Data_Platform/Web_publication

Wed, Sep 2, 8:55 AM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Machine-Learning-Team (Q1 FY2026-27)

Fri, Aug 21

BTullis reassigned T403154: Upgrade clouddumps hosts to bookworm/trixie from BTullis to bking.

I believe that @bking will be able to run the upgrades next week.

Fri, Aug 21, 6:37 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), tools-infrastructure-team, Cloud-VPS, cloud-services-team
BTullis closed T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format as Resolved.

Boldly resolving.

Fri, Aug 21, 3:45 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review, Observability-Logging
BTullis closed T427399: Alert in need of triage: PuppetFailure (instance an-test-client1002:9100) as Resolved.

Pupet is running cleanly. Closing.

Fri, Aug 21, 3:38 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), sre-alert-triage
BTullis moved T396870: update WDQS UI dashboard to account for graph split from Backlog - project to Backlog - operations on the Data-Platform-SRE (2026-08-07 - 2026-08-28) board.
Fri, Aug 21, 3:36 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), WMDE Analytics, Essential-Work, Wikidata Query UI, Wikidata
BTullis moved T398445: Update datacenter switchover cookbook to reflect WDQS graph split changes from Backlog - project to Backlog - operations on the Data-Platform-SRE (2026-08-07 - 2026-08-28) board.
Fri, Aug 21, 3:36 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Essential-Work
BTullis moved T412309: WDQS: Remove Search Platform as alert recipients from Backlog - project to Backlog - operations on the Data-Platform-SRE (2026-08-07 - 2026-08-28) board.
Fri, Aug 21, 3:36 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Essential-Work, Wikidata, Wikidata-Query-Service
BTullis edited projects for T435608: Create a Ceph administration VM and its Puppet role, added: Data-Platform-SRE (2026-08-07 - 2026-08-28); removed Data-Platform-SRE.
Fri, Aug 21, 1:16 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18)
BTullis moved T435604: [Tech evaluation] Table catalog options for the S3 data lake from Quarterly Goals to 2026-08-07 - 2026-08-28 on the Data-Platform-SRE board.
Fri, Aug 21, 12:55 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Data-Engineering
BTullis moved T374531: Determine how we want to manage the radosgw users on the DPE Ceph cluster from Toil / Automation to Quarterly Goals on the Data-Platform-SRE board.
Fri, Aug 21, 12:51 PM · Ceph, Data-Platform-SRE
BTullis edited projects for T428445: Upgrade Data Platform ceph clusters to version 19 - squid, added: Data-Platform-SRE (2026-08-07 - 2026-08-28); removed Data-Platform-SRE.
Fri, Aug 21, 12:48 PM · Data-Platform-SRE (2026-08-28 - 2026-09-18), Ceph
BTullis moved T435610: Decide and implement a logging strategy for the Data Platform RADOS Gateways from Incoming to Observability on the Data-Platform-SRE board.
Fri, Aug 21, 12:22 PM · Ceph, Data-Platform-SRE