Page MenuHomePhabricator

Blake (Blake)
User

Projects (8)

Today

  • No visible events.

Tomorrow

  • No visible events.

Wednesday

  • No visible events.

User Details

User Since
Nov 3 2025, 12:35 PM (40 w, 6 d)
Availability
Available
LDAP User
Blake
MediaWiki User
BJensen-WMF [ Global Accounts ]

Recent Activity

Thu, Aug 13

Blake added a comment to T434567: expanddblist without arguments doesn't fail cleanly.

Hello! What would be required from serviceops on this task? A bit more context would help us with triage. Thanks!

Thu, Aug 13, 3:55 PM · ServiceOps-Mediawiki, ServiceOps, Patch-For-Review
Blake triaged T433684: Create a template-radar pipeline as Medium priority.
Thu, Aug 13, 3:53 PM · ServiceOps, observability
Blake moved T434399: Disable built-in session handler via php.ini in WMF production from Inbox to Radar (Awareness) on the ServiceOps board.
Thu, Aug 13, 3:52 PM · MediaWiki-Core-Platform-Team (Kanban), MediaWiki-Core-AuthManager, ServiceOps
Blake updated subscribers of T421418: Upgrade Envoy to v1.39.0.

Working through https://wikitech.wikimedia.org/wiki/Envoy#Validate_the_new_version, and I've successfully deployed, tested, and subsequently rolled back envoy-future on mw-debug. @Scott_French, do you happen to have a suggestion for a "low traffic non-Mediawiki" service I might try to upgrade the Envoy of, to continue validation? Thanks!

Thu, Aug 13, 10:09 AM · ServiceOps, ServiceOps-Services-Oids, envoy

Wed, Aug 12

Blake renamed T421418: Upgrade Envoy to v1.39.0 from Upgrade Envoy to v1.38.3 to Upgrade Envoy to v1.39.0.
Wed, Aug 12, 12:03 PM · ServiceOps, ServiceOps-Services-Oids, envoy
Blake closed T434052: Release httpbb 0.0.6 as Resolved.

All set:

blake@cumin1003:~$ cat 2026-08-12-httpbb.yaml 
comment: ''
fixes:
  bookworm: 0.0.6-1
  bullseye: ''
  trixie: 0.0.6-1+deb13u1
libraries: []
source: httpbb
transitions: {}
update_type: tool
blake@cumin1003:~$ sudo debdeploy deploy -u 2026-08-12-httpbb.yaml -Q C:httpbb
Rolling out httpbb:
Non-daemon update, no service restart needed
Wed, Aug 12, 8:35 AM · ServiceOps

Tue, Aug 11

Blake moved T434392: Package and deploy php-zstd from Inbox to Backlog on the ServiceOps board.
Tue, Aug 11, 12:43 PM · ServiceOps
Blake triaged T434392: Package and deploy php-zstd as Medium priority.
Tue, Aug 11, 12:43 PM · ServiceOps
Blake added a comment to T434392: Package and deploy php-zstd.

Hi Amir! Would it be possible to get more context around this request? Is it currently blocking anything, or is it particularly urgent? Thanks!

Tue, Aug 11, 11:30 AM · ServiceOps

Fri, Aug 7

Blake added a comment to T421418: Upgrade Envoy to v1.39.0.

We are now at the step where we build the envoy-future Docker image.

Fri, Aug 7, 2:50 PM · ServiceOps, ServiceOps-Services-Oids, envoy
Blake added a comment to T434052: Release httpbb 0.0.6.

I've now built these, and they're uploaded, not sure how excited I am to debdeploy on Friday, though. I'll wait until Monday.

Fri, Aug 7, 11:44 AM · ServiceOps

Wed, Aug 5

Blake added a comment to T421418: Upgrade Envoy to v1.39.0.

After a review of the changelogs, I'm not seeing anything that jumps out as being particularly likely to break. It also seems as though we'd be able to proceed to 1.39, and I'm inclined to get us on the latest available version, given that there are no additional concerning changes. I'll proceed to attempt to build a new version.

Wed, Aug 5, 9:52 AM · ServiceOps, ServiceOps-Services-Oids, envoy
Blake updated the task description for T421418: Upgrade Envoy to v1.39.0.
Wed, Aug 5, 9:51 AM · ServiceOps, ServiceOps-Services-Oids, envoy
Blake triaged T434052: Release httpbb 0.0.6 as Medium priority.
Wed, Aug 5, 9:16 AM · ServiceOps
Blake created T434052: Release httpbb 0.0.6.
Wed, Aug 5, 8:22 AM · ServiceOps

Mon, Jul 27

Blake added a comment to T430290: Silence NodeBGPSessionStatusNotEstablished during reimages.

There's a tradeoff here, I think. The silence is currently time-limited, and we do not have any guarantees that the cookbook run will proceed within any given timeframe (it waits for operator input). If the second homer run hasn't been completed by the time the silence expires, we're going to get an alert. I don't think we want to silence for an arbitrarily long time, but we could consider increasing the duration if this is a common occurrence.

Mon, Jul 27, 10:05 AM · ServiceOps

Thu, Jul 23

Blake claimed T432460: Remove api-gateway kubernetes service components.
Thu, Jul 23, 10:39 AM · ServiceOps, ServiceOps-SharedInfra, MediaWiki-API-Platform-Team

Mon, Jul 20

Blake added a comment to T427405: Add kube-state-metrics 2.18.

Had a chat with Janis about this - we'll need to diff upstream at 7.3.0 (the most current version we can use) and 5.10.1 (the last time we brought in upstream), and verify that there are no chart changes that need to be made based on that diff. I'll be taking another look at this in a few weeks.

Mon, Jul 20, 2:26 PM · ServiceOps, Prod-Kubernetes, Kubernetes

Jul 17 2026

Blake added a comment to T432445: Remove api-gateway service traffic.

Patches have been uploaded and are ready for review. Thanks @jijiki for the overview today!

Jul 17 2026, 1:53 PM · Traffic, ServiceOps, ServiceOps-SharedInfra, Epic, MediaWiki-API-Platform-Team

Jul 15 2026

Blake closed T430290: Silence NodeBGPSessionStatusNotEstablished during reimages as Resolved.
Jul 15 2026, 12:33 PM · ServiceOps
Blake added a comment to T430290: Silence NodeBGPSessionStatusNotEstablished during reimages.

I don't think there's anything left to do here - there's now a flag to pass the renumber-node cookbook (--run_homer_inline) (e.g. sudo cookbook sre.k8s.renumber-node -t T421711 wikikube-worker1067.eqiad.wmnet --os=trixie --run_homer_inline) to help ensure that we don't miss the step.

Jul 15 2026, 12:32 PM · ServiceOps
Blake added a comment to T427668: Turn up the Pretrain MVP environment.

Okay, we're happy in codfw after fixing the host typo:

Jul 15 2026, 9:49 AM · MW-on-K8s, ServiceOps-Mediawiki, ServiceOps

Jul 14 2026

Blake added a comment to T421711: ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets.

Ah, okay, thanks Cathal!

Jul 14 2026, 12:34 PM · ServiceOps
Blake added a comment to T421711: ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets.

A diff I wasn't expecting during the reip for 1067:

Jul 14 2026, 11:16 AM · ServiceOps

Jul 13 2026

Blake added a comment to T431836: NodeBGPSessionStatusNotEstablished.

I don't think there's any follow-up remaining here, and agree that, had the silence succeeded for the new host, everything would have been fine. Thanks, Scott!

Jul 13 2026, 9:40 AM · ServiceOps, serviceops-deprecated

Jul 8 2026

Blake added a comment to T427668: Turn up the Pretrain MVP environment.

I am excited by this :)

Jul 8 2026, 3:17 PM · MW-on-K8s, ServiceOps-Mediawiki, ServiceOps
Blake added a comment to T431443: NodeBGPSessionStatusNotEstablished.

I ran a renumber-node with the patch, and it looks like the 4h silence created by pool-depool-node is long enough to cover the reimage and the homer runs afterwards, so I don't think that duration necessarily needs to change.

Jul 8 2026, 2:08 PM · ServiceOps-Upgrades-Hardware, ServiceOps, serviceops-deprecated
Blake added a comment to T431443: NodeBGPSessionStatusNotEstablished.

I think the reason I skipped the later homer run for wikikube-worker1036 is because I saw a homer run earlier in the output, and assumed that that was the run the cookbook was asking me about. sre.hosts.reimage calls sre.hosts.move-vlan, which calls sre.network.configure-switch-interfaces, which actively runs homer and outputs to the cookbook log, but only for the TOR switch, and not for the core router.

Jul 8 2026, 1:37 PM · ServiceOps-Upgrades-Hardware, ServiceOps, serviceops-deprecated
Blake added a comment to T431443: NodeBGPSessionStatusNotEstablished.

Arzhel mentioned that it would be acceptable to make an inline call to run_homer, so I'll start work on that.

Jul 8 2026, 10:23 AM · ServiceOps-Upgrades-Hardware, ServiceOps, serviceops-deprecated
Blake added a comment to T431443: NodeBGPSessionStatusNotEstablished.

In pool-depool-node, the Icinga downtime appears to be 4h:

Jul 8 2026, 10:12 AM · ServiceOps-Upgrades-Hardware, ServiceOps, serviceops-deprecated
Blake added a comment to T431443: NodeBGPSessionStatusNotEstablished.

I suspect I failed to run homer during the reimage of wikikube-worker1036. Looking at my shell history on cumin1003, it looks like I ran homer for wikikube-worker1037 (homer lsw1-d8-eqiad* commit 'T421711'), but I do not see a matching command for wikikube-worker1036 (and would expect homer lsw1-c6-eqiad* commit 'T421711'). It appears that I didn't !log this in either case, and will ensure to do that in the future. It also appears that I missed running homer for the cr in both cases, even though I can see it asking me to quite clearly in the cookbook logs:

Jul 8 2026, 8:47 AM · ServiceOps-Upgrades-Hardware, ServiceOps, serviceops-deprecated

Jul 7 2026

Blake added a comment to T430290: Silence NodeBGPSessionStatusNotEstablished during reimages.

It looks like downtimes are correctly created, and the matcher in question will catch instances which happen to have port numbers at the end. I think nothing needs to change here, and we should just be sure to use sre.k8s.renumber-node to reimage nodes which require a vlan move.

Jul 7 2026, 2:20 PM · ServiceOps

Jul 6 2026

Blake closed T423251: Remove Kubernetes 1.23 support, a subtask of T341984: Update Kubernetes clusters to 1.31, as Resolved.
Jul 6 2026, 2:16 PM · Data-Platform-SRE (2026.01.05 - 2026.01.23), Epic, ServiceOps, Patch-For-Review, Collaboration-Services, Kubernetes, Prod-Kubernetes
Blake closed T423251: Remove Kubernetes 1.23 support, a subtask of T427069: Update Kubernetes clusters to 1.34, as Resolved.
Jul 6 2026, 2:16 PM · Epic, ServiceOps, Prod-Kubernetes, Kubernetes
Blake closed T423251: Remove Kubernetes 1.23 support as Resolved.

Ah, okay, thanks. I'll close this out, then.

Jul 6 2026, 2:16 PM · ServiceOps, Kubernetes, Prod-Kubernetes
Blake added a comment to T423251: Remove Kubernetes 1.23 support.

@JMeybohm Does that mean I ought to remove all of the general-*.yaml files from the repo?

Jul 6 2026, 1:42 PM · ServiceOps, Kubernetes, Prod-Kubernetes

Jul 3 2026

Blake added a comment to T423851: Collect calico BGP metrics.

Ah, sorry I missed this - IMO, because we have one side of this BGP connection monitored, this is more of a feature request than an imminent production risk. I don't think it's a lot of work, though, and would be nice to have.

Jul 3 2026, 1:58 PM · ServiceOps (Next quarter), Sustainability (Incident Followup), ServiceOps-good-first-task, observability, Prod-Kubernetes, Kubernetes
Blake closed T430848: MediaWiki periodic job update-special-pages-s6 failed as Resolved.
Jul 3 2026, 9:43 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake closed T430848: MediaWiki periodic job update-special-pages-s6 failed, a subtask of T422486: MediaWiki periodic job failures due to timeouts, as Resolved.
Jul 3 2026, 9:43 AM · Data-Persistence, ServiceOps (Next quarter)
Blake added a comment to T430848: MediaWiki periodic job update-special-pages-s6 failed.

I've deleted the job, and started a manual run of the same.

Jul 3 2026, 9:42 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake added a comment to T427405: Add kube-state-metrics 2.18.

@JMeybohm, is there anything else that needs to be done before this is usable? Thanks!

Jul 3 2026, 9:06 AM · ServiceOps, Prod-Kubernetes, Kubernetes

Jun 30 2026

Blake closed T430631: Deployment rights to production-images for Blake as Invalid.

Hm, that's very strange - as far as I can tell, I'm using the user 'blake@wikimedia.org' on Gerrit. I don't think I've ever created or used a different account.

Jun 30 2026, 10:45 AM · Gerrit, Release-Engineering-Team
Blake added a comment to T430631: Deployment rights to production-images for Blake.

Incidentally, it looks like the repo might be locked down exclusively to gerrit admins (https://gerrit.wikimedia.org/r/admin/repos/operations/docker-images,access). I'm not sure whether that's intended.

Jun 30 2026, 10:00 AM · Gerrit, Release-Engineering-Team
Blake created T430631: Deployment rights to production-images for Blake.
Jun 30 2026, 9:31 AM · Gerrit, Release-Engineering-Team
Blake added a comment to T427405: Add kube-state-metrics 2.18.

That's now https://wikitech.wikimedia.org/wiki/Kube-state-metrics#Building_a_new_version, please let me know if there are any other details you'd like included.

Jun 30 2026, 8:38 AM · ServiceOps, Prod-Kubernetes, Kubernetes

Jun 26 2026

Blake triaged T430290: Silence NodeBGPSessionStatusNotEstablished during reimages as Medium priority.
Jun 26 2026, 10:17 AM · ServiceOps
Blake created T430290: Silence NodeBGPSessionStatusNotEstablished during reimages.
Jun 26 2026, 10:17 AM · ServiceOps

Jun 25 2026

Blake triaged T430133: Add "how_to_switch" field to the service catalog as Medium priority.
Jun 25 2026, 10:27 AM · ServiceOps
Blake created T430133: Add "how_to_switch" field to the service catalog.
Jun 25 2026, 10:26 AM · ServiceOps

Jun 23 2026

Blake claimed T427405: Add kube-state-metrics 2.18.
Jun 23 2026, 4:24 PM · ServiceOps, Prod-Kubernetes, Kubernetes
Blake added a comment to T423251: Remove Kubernetes 1.23 support.

Hm, looking at the description of https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1179654, it looks like the fixtures file was added intentionally. From what I understand from https://wikitech.wikimedia.org/wiki/Kubernetes/Deployment_Charts#Testing_a_Chart, it looks like that fixtures file might be necessary, though it's very possible that I'm misunderstanding something here.

Jun 23 2026, 11:32 AM · ServiceOps, Kubernetes, Prod-Kubernetes
Blake closed T429487: MediaWiki periodic job update-special-pages-s5 failed, a subtask of T422486: MediaWiki periodic job failures due to timeouts, as Resolved.
Jun 23 2026, 10:17 AM · Data-Persistence, ServiceOps (Next quarter)
Blake closed T429487: MediaWiki periodic job update-special-pages-s5 failed as Resolved.
Jun 23 2026, 10:17 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake added a comment to T429487: MediaWiki periodic job update-special-pages-s5 failed.

A more recent run has completed, and the old job has been cleaned up.

Jun 23 2026, 10:17 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake triaged T428750: Move poolcounter definitions to etcdconfig as Medium priority.
Jun 23 2026, 10:14 AM · ServiceOps
Blake placed T428750: Move poolcounter definitions to etcdconfig up for grabs.
Jun 23 2026, 10:14 AM · ServiceOps
Blake closed T424942: wikikube-worker13[75-84] implementation tracking, a subtask of T423719: Repurpose tools-k8s-ctrl[1001-1002],tools-k8s-worker[1001-1008] to wikikube-worker13{75-84}, as Resolved.
Jun 23 2026, 9:28 AM · ServiceOps, ServiceOps-Upgrades-Hardware, SRE, ops-eqiad, DC-Ops
Blake closed T424942: wikikube-worker13[75-84] implementation tracking as Resolved.
Jun 23 2026, 9:28 AM · ServiceOps, ServiceOps-Upgrades-Hardware, DC-Ops

Jun 19 2026

Blake added a comment to T429487: MediaWiki periodic job update-special-pages-s5 failed.

Same issue as in the parent. A new run has already started for this job.

Jun 19 2026, 10:06 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake added a subtask for T422486: MediaWiki periodic job failures due to timeouts: T429487: MediaWiki periodic job update-special-pages-s5 failed.
Jun 19 2026, 10:05 AM · Data-Persistence, ServiceOps (Next quarter)
Blake added a parent task for T429487: MediaWiki periodic job update-special-pages-s5 failed: T422486: MediaWiki periodic job failures due to timeouts.
Jun 19 2026, 10:05 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake claimed T429487: MediaWiki periodic job update-special-pages-s5 failed.
Jun 19 2026, 10:00 AM · ServiceOps, Wikimedia-production-error, MediaWiki-Special-pages
Blake triaged T429480: Update helmfile to supported v1.x version as Medium priority.
Jun 19 2026, 9:11 AM · ServiceOps, Prod-Kubernetes, Kubernetes
Blake lowered the priority of T429156: EtcdConfig failed to fetch data: (curl error: 28) Timeout was reached from High to Medium.

Reprioritizing this to medium, as the immediate issue has been mitigated.

Jun 19 2026, 9:07 AM · Kubernetes, ServiceOps, Product Safety and Integrity
Blake changed the status of T429599: Create an llms.txt where honest robots can read our API Policy from Open to In Progress.
Jun 19 2026, 9:00 AM · Patch-For-Review, ServiceOps
Blake changed the status of T429599: Create an llms.txt where honest robots can read our API Policy, a subtask of T426157: RfC: Investigate the use of llms.txt as a polite-crawler instruction manual for our sites, from Open to In Progress.
Jun 19 2026, 9:00 AM · Patch-For-Review, ServiceOps

Jun 18 2026

Blake updated the task description for T426044: Migrate Mediawiki memcached to Debian Trixie.
Jun 18 2026, 9:33 AM · ServiceOps

Jun 17 2026

Blake triaged T421418: Upgrade Envoy to v1.39.0 as Medium priority.
Jun 17 2026, 9:03 AM · ServiceOps, ServiceOps-Services-Oids, envoy
Blake added a subtask for T421418: Upgrade Envoy to v1.39.0: Unknown Object (Task).
Jun 17 2026, 9:01 AM · ServiceOps, ServiceOps-Services-Oids, envoy

Jun 15 2026

Blake placed T423852: Add calico network alerting up for grabs.

I suspect this task, in addition to T423851, would probably take 3-4w of non-continuous effort. Unassigning myself for now, as this work is unlikely to be scheduled for this coming quarter.

Jun 15 2026, 12:17 PM · ServiceOps (Next quarter), Sustainability (Incident Followup), ServiceOps-good-first-task, observability, Prod-Kubernetes, Kubernetes
Blake closed T428772: Serve mediawiki keys.txt with UTF-8 charset as Resolved.
Jun 15 2026, 10:20 AM · Wikimedia-production-error, ServiceOps-good-first-task, MediaWiki-Core-Platform-Team (Radar), ServiceOps, ServiceOps-Mediawiki, Wikimedia-Apache-configuration

Jun 12 2026

Blake updated subscribers of T428772: Serve mediawiki keys.txt with UTF-8 charset.

After a spelunking session with @Clement_Goubert (thanks very much!), we found that Apache configuration for mw-on-k8s ought to be modified in hieradata/common/mediawiki.yaml. I've sent a patch for review, and we'll look to deploy this next week.

Jun 12 2026, 10:54 AM · Wikimedia-production-error, ServiceOps-good-first-task, MediaWiki-Core-Platform-Team (Radar), ServiceOps, ServiceOps-Mediawiki, Wikimedia-Apache-configuration
Blake claimed T428772: Serve mediawiki keys.txt with UTF-8 charset.
Jun 12 2026, 9:08 AM · Wikimedia-production-error, ServiceOps-good-first-task, MediaWiki-Core-Platform-Team (Radar), ServiceOps, ServiceOps-Mediawiki, Wikimedia-Apache-configuration
Blake added a comment to T428772: Serve mediawiki keys.txt with UTF-8 charset.

I'm inclined to try adding .txt to the UTF-8 AddCharset directive here, unless anyone else has a more informed opinion about how to proceed.

Jun 12 2026, 9:08 AM · Wikimedia-production-error, ServiceOps-good-first-task, MediaWiki-Core-Platform-Team (Radar), ServiceOps, ServiceOps-Mediawiki, Wikimedia-Apache-configuration

Jun 11 2026

Blake added a comment to T428772: Serve mediawiki keys.txt with UTF-8 charset.

Looks like my patch didn't work - the charset parameter wasn't added to the Content-Type header. I'm currently in the process of reverting, and will explore this more tomorrow.

Jun 11 2026, 5:32 PM · Wikimedia-production-error, ServiceOps-good-first-task, MediaWiki-Core-Platform-Team (Radar), ServiceOps, ServiceOps-Mediawiki, Wikimedia-Apache-configuration
Blake added a comment to T423251: Remove Kubernetes 1.23 support.

The files in deployment-charts that appear to reference 1.23 directly are:

Jun 11 2026, 10:50 AM · ServiceOps, Kubernetes, Prod-Kubernetes

Jun 10 2026

Blake moved T428750: Move poolcounter definitions to etcdconfig from Inbox to Backlog on the ServiceOps board.
Jun 10 2026, 1:24 PM · ServiceOps
Blake created T428750: Move poolcounter definitions to etcdconfig.
Jun 10 2026, 1:23 PM · ServiceOps

Jun 9 2026

Blake updated the task description for T426044: Migrate Mediawiki memcached to Debian Trixie.
Jun 9 2026, 5:25 PM · ServiceOps
Blake closed T427065: Update Kubernetes clusters to 1.31.14, a subtask of T427069: Update Kubernetes clusters to 1.34, as Resolved.
Jun 9 2026, 3:17 PM · Epic, ServiceOps, Prod-Kubernetes, Kubernetes
Blake closed T427065: Update Kubernetes clusters to 1.31.14 as Resolved.
Jun 9 2026, 3:17 PM · ServiceOps, Prod-Kubernetes, Kubernetes
Blake added a comment to T427065: Update Kubernetes clusters to 1.31.14.

This is complete, the the process has been documented at https://wikitech.wikimedia.org/wiki/Kubernetes/Point_upgrades.

Jun 9 2026, 3:17 PM · ServiceOps, Prod-Kubernetes, Kubernetes

Jun 4 2026

Blake added a comment to T427065: Update Kubernetes clusters to 1.31.14.

I've rebuilt the package (kubernetes_1.31.14-2) and deployed it to apt. After an apt update and 'apt install kubernetes-node', kubestage2004 correctly updated and restarted kubelet and kube-proxy automatically. I'll upgrade the rest of the staging clusters today.

Jun 4 2026, 11:16 AM · ServiceOps, Prod-Kubernetes, Kubernetes

Jun 2 2026

Blake added a comment to T426044: Migrate Mediawiki memcached to Debian Trixie.

All of the memcached hosts in the main pool in eqiad have been reimaged to Trixie.

Jun 2 2026, 4:32 PM · ServiceOps
Blake updated the task description for T426044: Migrate Mediawiki memcached to Debian Trixie.
Jun 2 2026, 4:32 PM · ServiceOps

May 29 2026

Blake claimed T426044: Migrate Mediawiki memcached to Debian Trixie.
May 29 2026, 11:39 AM · ServiceOps
Blake added a comment to T427065: Update Kubernetes clusters to 1.31.14.

During a discussion with @JMeybohm a couple days ago, it sounded like we should be able to update kubernetes-node, then kubernetes-client, then kubernetes-master, sequentially. That seems to have worked correctly on the workers of staging-codfw, but when I got to the controllers, debdeploy shows this (regular node and debdeploy spec included for context):

May 29 2026, 9:34 AM · ServiceOps, Prod-Kubernetes, Kubernetes

May 28 2026

Blake added a comment to T426044: Migrate Mediawiki memcached to Debian Trixie.

mc1054 has been added to the pool, and things look good (also, memkeys was tested, and that worked!). i'll reimage mc1055 to trixie on monday, and then swap these servers back so we can proceed with the decom of 1054.

May 28 2026, 11:15 AM · ServiceOps

May 27 2026

Blake closed T418927: wikikube-worker23[57-74] implementation tracking as Resolved.
May 27 2026, 1:20 PM · ServiceOps-Upgrades-Hardware, ServiceOps, SRE
Blake closed T418927: wikikube-worker23[57-74] implementation tracking, a subtask of T418925: Q3:rack/setup/install wikikube-worker23[57-74], as Resolved.
May 27 2026, 1:20 PM · ops-codfw, ServiceOps-Upgrades-Hardware, ServiceOps, SRE, DC-Ops
Blake added a comment to T418927: wikikube-worker23[57-74] implementation tracking.

if i'm ever looking at this task for history, the docs are here.

May 27 2026, 1:20 PM · ServiceOps-Upgrades-Hardware, ServiceOps, SRE
Blake claimed T424942: wikikube-worker13[75-84] implementation tracking.
May 27 2026, 1:09 PM · ServiceOps, ServiceOps-Upgrades-Hardware, DC-Ops
Blake added a comment to T427065: Update Kubernetes clusters to 1.31.14.

the packages have been built, and uploaded (and the docs have been updated). next step is to roll this out to staging-codfw.

May 27 2026, 10:26 AM · ServiceOps, Prod-Kubernetes, Kubernetes
Blake added a comment to T426044: Migrate Mediawiki memcached to Debian Trixie.

i've verified that mc1054 is running the versions of memkeys and prometheus-memcached-exporter we expect, so i believe we can now add it to the pool.

May 27 2026, 8:40 AM · ServiceOps

May 26 2026

Blake closed T426049: Package prometheus-memcached-exporter for Debian Trixie as Resolved.
May 26 2026, 3:21 PM · ServiceOps
Blake closed T426049: Package prometheus-memcached-exporter for Debian Trixie, a subtask of T426044: Migrate Mediawiki memcached to Debian Trixie, as Resolved.
May 26 2026, 3:21 PM · ServiceOps
Blake added a comment to T426049: Package prometheus-memcached-exporter for Debian Trixie.

trixie-packaging-wikimedia now contains the latest upstream code, and has been built and uploaded to apt.

May 26 2026, 3:21 PM · ServiceOps

May 21 2026

Blake changed the status of T418927: wikikube-worker23[57-74] implementation tracking, a subtask of T418925: Q3:rack/setup/install wikikube-worker23[57-74], from Open to In Progress.
May 21 2026, 10:51 AM · ops-codfw, ServiceOps-Upgrades-Hardware, ServiceOps, SRE, DC-Ops
Blake changed the status of T418927: wikikube-worker23[57-74] implementation tracking from Open to In Progress.
May 21 2026, 10:51 AM · ServiceOps-Upgrades-Hardware, ServiceOps, SRE
Blake closed T426047: Package memkeys for Debian Trixie, a subtask of T426044: Migrate Mediawiki memcached to Debian Trixie, as Resolved.
May 21 2026, 9:55 AM · ServiceOps