Page MenuHomePhabricator

Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie
Open, Needs TriagePublic

Related Objects

Event Timeline

Change #1302202 had a related patch set uploaded (by Ahmon Dancy; author: Ahmon Dancy):

[operations/mediawiki-config@master] beta: Replace deployment-db14 with deployment-db15

https://gerrit.wikimedia.org/r/1302202

Change #1302202 merged by jenkins-bot:

[operations/mediawiki-config@master] beta: Replace deployment-db14 with deployment-db15

https://gerrit.wikimedia.org/r/1302202

Change #1302257 had a related patch set uploaded (by Ahmon Dancy; author: Ahmon Dancy):

[operations/mediawiki-config@master] beta: Add deployment-db16 as a second read replica

https://gerrit.wikimedia.org/r/1302257

Change #1302257 merged by jenkins-bot:

[operations/mediawiki-config@master] beta: Add deployment-db16 as a second read replica

https://gerrit.wikimedia.org/r/1302257

bd808 added subscribers: Eevans, KOfori.

I missed the email sent on Jun 17 (though I'm not sure it would have changed things much); I only became aware of this after Wednesday's (Jul 29) email.

Of the instances that (I assume) are associated with me: deployment-echostore02.deployment-prep, deployment-restbase05.deployment-prep, and deployment-sessionstore06.deployment-prep (at least) are definitely going to cause deployment-prep breakage if you shut them down, and I will definitely not have any time to do anything with them before the deadline. And, I'm not comfortable committing to when I would be able to; My first priority at the moment is production machines still running Bullseye (I'm sure you can understand why I cannot pause that to spend time here).

More broadly, I think we need to get clarity on the ownership of these. It's seems —at least implicitly— that this falls to me (or Data Persistence), but I don't think that has been made explicit. I'm always happy to help, but I'm not sure that best-effort, or as-needed, is cutting it here.

More broadly, I think we need to get clarity on the ownership of these. It's seems —at least implicitly— that this falls to me (or Data Persistence), but I don't think that has been made explicit. I'm always happy to help, but I'm not sure that best-effort, or as-needed, is cutting it here.

TL;DR:

  • DevEx is "responsible"
  • DevEx may ask you for help
  • DevEx may ask your management chain to prioritize your help if you are WMF staff
  • DevEx will be looking for a contractor/vendor to help

In T215217: deployment-prep (beta cluster): Code stewardship request @Lferreira declared that the Developer Experience group he leads are the stewards of Beta Cluster. As part of that group I have been trying to work through what that means in practice. To date I have settled on an interpretation that we are ultimately the responsible parties for keeping the project running, but that we also cannot do that work without extensive support from other teams and individuals inside and outside of the Wikimedia Foundation.

To that end, I can and do take the initiative to poke people for help when I see something that is breaking in ways that are disruptive to the larger project. I know you have been one of the folks I have poked @Eevans, and that you have made time to help when you could. I appreciate that. I also agree that there is no well established organizational responsibility for you to always be the responder. I doubt that is going to change honestly after seeing multiple attempts find consensus on Beta Cluster as WMF pre-production staging project maintained by the same teams and roles as maintain the Wikimedia Foundation core services replicated there fail.

A big part of the reason the Developer Experience is working on several projects to make new spaces for workflows that historically have been performed in Beta Cluster is that we have proven to ourselves that the current ad hoc management of Beta Cluster services is not scalable, equitable, or responsive. Patch Demo is being enhanced in part to make it a stable demonstration environment for pre-release features and experiments. Pretrain is being implemented to provide a more stable environment for post-merge validation of features and experiments. Once these projects have progressed technically and socially to be able to handle those workflows we plan to revisit the remaining Beta Cluster use cases to find the next thing to rehome. Our ultimate goal is to be able to tell everyone where these various workflows are better supported as we decommission Beta Cluster entirely. There is wiggle room in that end state to transfer the project to a group of technical volunteers if such a group arises with a meaningful desire and commitment to carry the project forward in support of some unfulfilled workflow, but the intent is that "critical" workflows all find a better home.

There are a bunch of future looking statements above that will not be realized before there is extreme pressure to remove support for Debian Bullseye from operations/puppet.git and the entire Wikimedia operated compute environment. So we are back to needing to talk about the shorter term needs that Eric has rightly brought forward. I had hopes for my own efforts in the migration to be more than writing phab tasks by this point. Those hopes did not turn into work yet however. Levi and I have been talking about the problem and have an idea about seeing if we can make a bit better migration than the last one which was very heavily the labor of one volunteer who just jumped in and did the work. The idea is to find someone, probably from the volunteer technical community, who has the skills and time to do the work as contract project. We have not yet started the search for that vendor, but it is our next step in that direction. This is still a bit of a firefighting hero plan, but it at least wants to provide some compensation beyond thanks and barnstars for the labor.

I don't think that we will find someone so magically aligned with the technical needs of every bit of the Wikimedia stack in Beta Cluster that no help will be needed from SREs and others who maintain parallel services in production. I do hope however that the needed support will be manageable with some priority negotiation vs other commitments.

I would also add as a post script here that we should all feel empowered to resist social pressures to help when that help is really asking us to over commit. Those of us asking for help must listen for those no answers and respect them.

[ ... ]

TL;DR:

  • DevEx is "responsible"
  • DevEx may ask you for help
  • DevEx may ask your management chain to prioritize your help if you are WMF staff
  • DevEx will be looking for a contractor/vendor to help

[ ... ]

A big part of the reason the Developer Experience is working on several projects to make new spaces for workflows that historically have been performed in Beta Cluster is that we have proven to ourselves that the current ad hoc management of Beta Cluster services is not scalable, equitable, or responsive. Patch Demo is being enhanced in part to make it a stable demonstration environment for pre-release features and experiments. Pretrain is being implemented to provide a more stable environment for post-merge validation of features and experiments. Once these projects have progressed technically and socially to be able to handle those workflows we plan to revisit the remaining Beta Cluster use cases to find the next thing to rehome. Our ultimate goal is to be able to tell everyone where these various workflows are better supported as we decommission Beta Cluster entirely. There is wiggle room in that end state to transfer the project to a group of technical volunteers if such a group arises with a meaningful desire and commitment to carry the project forward in support of some unfulfilled workflow, but the intent is that "critical" workflows all find a better home.

TIL & this is great to hear!

[ ... ]

I don't think that we will find someone so magically aligned with the technical needs of every bit of the Wikimedia stack in Beta Cluster that no help will be needed from SREs and others who maintain parallel services in production. I do hope however that the needed support will be manageable with some priority negotiation vs other commitments.

I would also add as a post script here that we should all feel empowered to resist social pressures to help when that help is really asking us to over commit. Those of us asking for help must listen for those no answers and respect them.

Thanks @bd808!