< T327742: Migrate deployment-prep away from Debian Buster to Bullseye/Bookworm | NOTYETCREATED >
Instances left to migrate: live report
Related:
< T327742: Migrate deployment-prep away from Debian Buster to Bullseye/Bookworm | NOTYETCREATED >
Instances left to migrate: live report
Related:
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Open | None | T401839 Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie | |||
| Resolved | Krinkle | T380881 Re-create poolcounter instance in Beta Cluster (deployment-prep) | |||
| Resolved | dancy | T412975 Replace deployment-mx03 with a bookworm-based instance (was Puppet failure: "Unable to locate package spamd") | |||
| Declined | Feature | bd808 | T421244 Replace deployment-deploy04 with a Bookworm instance with Java 21 | ||
| Open | None | T421449 Clean up broken legacy maps stack in Beta Cluster | |||
| Open | Feature | None | T426326 Replace deployment-mx04 with a newer OS and MX service stack | ||
| Open | None | T428052 Beta cluster haproxy does not support `warn-blocked-traffic-after` keyword | |||
| Resolved | dancy | T428910 Beta Cluster MariaDB is still 10.6.17, MW now requires 10.11 | |||
| Resolved | dancy | T428930 Set up deployment-db15 with Trixie and wmf-mariadb1011 | |||
| Open | tchin | T429497 Upgrade (replace) beta|deployment-prep event platform VPS instances | |||
| In Progress | bd808 | T433998 Planning for a plan to migrate Beta Cluster again |
I would like this to involve T394316: Use infrastructure as code techniques to rebuild the Beta Cluster. I have been poking at T398643: [tofu-cloudvps] cloudvps_puppet_prefix.hiera settings show dirty diffs based on YAML canonicalization which I feel is the major blocker to tofu adoption in Beta Cluster.
Change #1302202 had a related patch set uploaded (by Ahmon Dancy; author: Ahmon Dancy):
[operations/mediawiki-config@master] beta: Replace deployment-db14 with deployment-db15
Change #1302202 merged by jenkins-bot:
[operations/mediawiki-config@master] beta: Replace deployment-db14 with deployment-db15
Change #1302257 had a related patch set uploaded (by Ahmon Dancy; author: Ahmon Dancy):
[operations/mediawiki-config@master] beta: Add deployment-db16 as a second read replica
Change #1302257 merged by jenkins-bot:
[operations/mediawiki-config@master] beta: Add deployment-db16 as a second read replica
TL;DR:
In T215217: deployment-prep (beta cluster): Code stewardship request @Lferreira declared that the Developer Experience group he leads are the stewards of Beta Cluster. As part of that group I have been trying to work through what that means in practice. To date I have settled on an interpretation that we are ultimately the responsible parties for keeping the project running, but that we also cannot do that work without extensive support from other teams and individuals inside and outside of the Wikimedia Foundation.
To that end, I can and do take the initiative to poke people for help when I see something that is breaking in ways that are disruptive to the larger project. I know you have been one of the folks I have poked @Eevans, and that you have made time to help when you could. I appreciate that. I also agree that there is no well established organizational responsibility for you to always be the responder. I doubt that is going to change honestly after seeing multiple attempts find consensus on Beta Cluster as WMF pre-production staging project maintained by the same teams and roles as maintain the Wikimedia Foundation core services replicated there fail.
A big part of the reason the Developer Experience is working on several projects to make new spaces for workflows that historically have been performed in Beta Cluster is that we have proven to ourselves that the current ad hoc management of Beta Cluster services is not scalable, equitable, or responsive. Patch Demo is being enhanced in part to make it a stable demonstration environment for pre-release features and experiments. Pretrain is being implemented to provide a more stable environment for post-merge validation of features and experiments. Once these projects have progressed technically and socially to be able to handle those workflows we plan to revisit the remaining Beta Cluster use cases to find the next thing to rehome. Our ultimate goal is to be able to tell everyone where these various workflows are better supported as we decommission Beta Cluster entirely. There is wiggle room in that end state to transfer the project to a group of technical volunteers if such a group arises with a meaningful desire and commitment to carry the project forward in support of some unfulfilled workflow, but the intent is that "critical" workflows all find a better home.
There are a bunch of future looking statements above that will not be realized before there is extreme pressure to remove support for Debian Bullseye from operations/puppet.git and the entire Wikimedia operated compute environment. So we are back to needing to talk about the shorter term needs that Eric has rightly brought forward. I had hopes for my own efforts in the migration to be more than writing phab tasks by this point. Those hopes did not turn into work yet however. Levi and I have been talking about the problem and have an idea about seeing if we can make a bit better migration than the last one which was very heavily the labor of one volunteer who just jumped in and did the work. The idea is to find someone, probably from the volunteer technical community, who has the skills and time to do the work as contract project. We have not yet started the search for that vendor, but it is our next step in that direction. This is still a bit of a firefighting hero plan, but it at least wants to provide some compensation beyond thanks and barnstars for the labor.
I don't think that we will find someone so magically aligned with the technical needs of every bit of the Wikimedia stack in Beta Cluster that no help will be needed from SREs and others who maintain parallel services in production. I do hope however that the needed support will be manageable with some priority negotiation vs other commitments.
I would also add as a post script here that we should all feel empowered to resist social pressures to help when that help is really asking us to over commit. Those of us asking for help must listen for those no answers and respect them.
TIL & this is great to hear!
[ ... ]
I don't think that we will find someone so magically aligned with the technical needs of every bit of the Wikimedia stack in Beta Cluster that no help will be needed from SREs and others who maintain parallel services in production. I do hope however that the needed support will be manageable with some priority negotiation vs other commitments.
I would also add as a post script here that we should all feel empowered to resist social pressures to help when that help is really asking us to over commit. Those of us asking for help must listen for those no answers and respect them.
Thanks @bd808!