User Details
- User Since
- Nov 3 2025, 12:35 PM (40 w, 6 d)
- Availability
- Available
- LDAP User
- Blake
- MediaWiki User
- BJensen-WMF [ Global Accounts ]
Thu, Aug 13
Hello! What would be required from serviceops on this task? A bit more context would help us with triage. Thanks!
Working through https://wikitech.wikimedia.org/wiki/Envoy#Validate_the_new_version, and I've successfully deployed, tested, and subsequently rolled back envoy-future on mw-debug. @Scott_French, do you happen to have a suggestion for a "low traffic non-Mediawiki" service I might try to upgrade the Envoy of, to continue validation? Thanks!
Wed, Aug 12
All set:
blake@cumin1003:~$ cat 2026-08-12-httpbb.yaml
comment: ''
fixes:
bookworm: 0.0.6-1
bullseye: ''
trixie: 0.0.6-1+deb13u1
libraries: []
source: httpbb
transitions: {}
update_type: tool
blake@cumin1003:~$ sudo debdeploy deploy -u 2026-08-12-httpbb.yaml -Q C:httpbb
Rolling out httpbb:
Non-daemon update, no service restart neededTue, Aug 11
Hi Amir! Would it be possible to get more context around this request? Is it currently blocking anything, or is it particularly urgent? Thanks!
Fri, Aug 7
We are now at the step where we build the envoy-future Docker image.
I've now built these, and they're uploaded, not sure how excited I am to debdeploy on Friday, though. I'll wait until Monday.
Wed, Aug 5
After a review of the changelogs, I'm not seeing anything that jumps out as being particularly likely to break. It also seems as though we'd be able to proceed to 1.39, and I'm inclined to get us on the latest available version, given that there are no additional concerning changes. I'll proceed to attempt to build a new version.
Mon, Jul 27
There's a tradeoff here, I think. The silence is currently time-limited, and we do not have any guarantees that the cookbook run will proceed within any given timeframe (it waits for operator input). If the second homer run hasn't been completed by the time the silence expires, we're going to get an alert. I don't think we want to silence for an arbitrarily long time, but we could consider increasing the duration if this is a common occurrence.
Thu, Jul 23
Mon, Jul 20
Had a chat with Janis about this - we'll need to diff upstream at 7.3.0 (the most current version we can use) and 5.10.1 (the last time we brought in upstream), and verify that there are no chart changes that need to be made based on that diff. I'll be taking another look at this in a few weeks.
Jul 17 2026
Patches have been uploaded and are ready for review. Thanks @jijiki for the overview today!
Jul 15 2026
I don't think there's anything left to do here - there's now a flag to pass the renumber-node cookbook (--run_homer_inline) (e.g. sudo cookbook sre.k8s.renumber-node -t T421711 wikikube-worker1067.eqiad.wmnet --os=trixie --run_homer_inline) to help ensure that we don't miss the step.
Okay, we're happy in codfw after fixing the host typo:
Jul 14 2026
Ah, okay, thanks Cathal!
A diff I wasn't expecting during the reip for 1067:
Jul 13 2026
I don't think there's any follow-up remaining here, and agree that, had the silence succeeded for the new host, everything would have been fine. Thanks, Scott!
Jul 8 2026
I am excited by this :)
I ran a renumber-node with the patch, and it looks like the 4h silence created by pool-depool-node is long enough to cover the reimage and the homer runs afterwards, so I don't think that duration necessarily needs to change.
I think the reason I skipped the later homer run for wikikube-worker1036 is because I saw a homer run earlier in the output, and assumed that that was the run the cookbook was asking me about. sre.hosts.reimage calls sre.hosts.move-vlan, which calls sre.network.configure-switch-interfaces, which actively runs homer and outputs to the cookbook log, but only for the TOR switch, and not for the core router.
Arzhel mentioned that it would be acceptable to make an inline call to run_homer, so I'll start work on that.
In pool-depool-node, the Icinga downtime appears to be 4h:
I suspect I failed to run homer during the reimage of wikikube-worker1036. Looking at my shell history on cumin1003, it looks like I ran homer for wikikube-worker1037 (homer lsw1-d8-eqiad* commit 'T421711'), but I do not see a matching command for wikikube-worker1036 (and would expect homer lsw1-c6-eqiad* commit 'T421711'). It appears that I didn't !log this in either case, and will ensure to do that in the future. It also appears that I missed running homer for the cr in both cases, even though I can see it asking me to quite clearly in the cookbook logs:
Jul 7 2026
It looks like downtimes are correctly created, and the matcher in question will catch instances which happen to have port numbers at the end. I think nothing needs to change here, and we should just be sure to use sre.k8s.renumber-node to reimage nodes which require a vlan move.
Jul 6 2026
Ah, okay, thanks. I'll close this out, then.
@JMeybohm Does that mean I ought to remove all of the general-*.yaml files from the repo?
Jul 3 2026
Ah, sorry I missed this - IMO, because we have one side of this BGP connection monitored, this is more of a feature request than an imminent production risk. I don't think it's a lot of work, though, and would be nice to have.
I've deleted the job, and started a manual run of the same.
@JMeybohm, is there anything else that needs to be done before this is usable? Thanks!
Jun 30 2026
Hm, that's very strange - as far as I can tell, I'm using the user 'blake@wikimedia.org' on Gerrit. I don't think I've ever created or used a different account.
Incidentally, it looks like the repo might be locked down exclusively to gerrit admins (https://gerrit.wikimedia.org/r/admin/repos/operations/docker-images,access). I'm not sure whether that's intended.
That's now https://wikitech.wikimedia.org/wiki/Kube-state-metrics#Building_a_new_version, please let me know if there are any other details you'd like included.
Jun 26 2026
Jun 25 2026
Jun 23 2026
Hm, looking at the description of https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1179654, it looks like the fixtures file was added intentionally. From what I understand from https://wikitech.wikimedia.org/wiki/Kubernetes/Deployment_Charts#Testing_a_Chart, it looks like that fixtures file might be necessary, though it's very possible that I'm misunderstanding something here.
A more recent run has completed, and the old job has been cleaned up.
Jun 19 2026
Same issue as in the parent. A new run has already started for this job.
Reprioritizing this to medium, as the immediate issue has been mitigated.
Jun 18 2026
Jun 17 2026
Jun 15 2026
I suspect this task, in addition to T423851, would probably take 3-4w of non-continuous effort. Unassigning myself for now, as this work is unlikely to be scheduled for this coming quarter.
Jun 12 2026
After a spelunking session with @Clement_Goubert (thanks very much!), we found that Apache configuration for mw-on-k8s ought to be modified in hieradata/common/mediawiki.yaml. I've sent a patch for review, and we'll look to deploy this next week.
I'm inclined to try adding .txt to the UTF-8 AddCharset directive here, unless anyone else has a more informed opinion about how to proceed.
Jun 11 2026
Looks like my patch didn't work - the charset parameter wasn't added to the Content-Type header. I'm currently in the process of reverting, and will explore this more tomorrow.
The files in deployment-charts that appear to reference 1.23 directly are:
Jun 10 2026
Jun 9 2026
This is complete, the the process has been documented at https://wikitech.wikimedia.org/wiki/Kubernetes/Point_upgrades.
Jun 4 2026
I've rebuilt the package (kubernetes_1.31.14-2) and deployed it to apt. After an apt update and 'apt install kubernetes-node', kubestage2004 correctly updated and restarted kubelet and kube-proxy automatically. I'll upgrade the rest of the staging clusters today.
Jun 2 2026
All of the memcached hosts in the main pool in eqiad have been reimaged to Trixie.
May 29 2026
During a discussion with @JMeybohm a couple days ago, it sounded like we should be able to update kubernetes-node, then kubernetes-client, then kubernetes-master, sequentially. That seems to have worked correctly on the workers of staging-codfw, but when I got to the controllers, debdeploy shows this (regular node and debdeploy spec included for context):
May 28 2026
mc1054 has been added to the pool, and things look good (also, memkeys was tested, and that worked!). i'll reimage mc1055 to trixie on monday, and then swap these servers back so we can proceed with the decom of 1054.
May 27 2026
if i'm ever looking at this task for history, the docs are here.
the packages have been built, and uploaded (and the docs have been updated). next step is to roll this out to staging-codfw.
i've verified that mc1054 is running the versions of memkeys and prometheus-memcached-exporter we expect, so i believe we can now add it to the pool.
May 26 2026
trixie-packaging-wikimedia now contains the latest upstream code, and has been built and uploaded to apt.
