User Details
- User Since
- Jan 5 2016, 9:54 PM (553 w, 1 d)
- Availability
- Busy Busy until Aug 14.
- LDAP User
- Unknown
- MediaWiki User
- LToscano (WMF) [ Global Accounts ]
Fri, Jul 31
Had a chat with @CWilliams-WMF, the patch is merged and it will be published in the next spicerack release. I'll be afk for a couple of weeks, so for any tests etc.. it should be sufficient to use the PYTHONPATH env variable with the current spicerack's HEAD :)
To keep archives happy - in order to reimage the host, you'd need to use:
After a chat with Matthew it seems that all Supermicro config J hosts (see https://netbox.wikimedia.org/dcim/devices/?device_type_id=338) have some problem with IPMI: both ADMIN and wmfroot cannot authenticate correctly to execute basic commands. Interestingly, the new config J hosts with single CPU socket (sretest2010) don't show this behavior. It is not an immediate issue since all of the config j hosts use UEFI, and hence ipmi is not needed to reimage.
So the Bios endpoint issue seemed to be just waiting for a host reboot: IIUC on Supermicro X12 the BMC is just a proxy for the mb's BIOS, and forcing a complete host startup causes a resync.
Done!
@MatthewVernon I tried to reset the BMC via Redfish but it didn't help, and everything boils down to the /redfish/v1/Systems/1/Bios endpoint returning 503 for every try. I would suggest to reach out to dcops to ask them to shutdown and drain the BMC completely, and then restart it.
@Dzahn technically we are adding a sudo rule so the I/F team should +1 it, but I already anticipate it is a +1. I'll try to make a quick poll today in my team's chat so we can proceed.
Thu, Jul 30
@jcrespo I somehow missed it, most probably because it wasn't part of cumin aliases anymore! I re-run the script for it now, it should work!
https://gerrit.wikimedia.org/r/1319427 fixed the idrac 10 hosts, so the ipmi rollout should be fine.
Failures to set the IPMI after another round of cookbooks runs:
Wed, Jul 29
Copied the image:tag combinations specified in T432829#12144758, plus all the more recent tags up to the latest ones.
@CWilliams-WMF the easiest could be to just bump the max tries to something more, I think the 15 value was probably something established in the past that worked. Maybe we could just bump it to 20/25 and restart using it?
Old VMs deleted!
Raising the priority since T433246 is a clear sign that we should proceed in this direction sooner rather than later :)
Tue, Jul 28
The host is now ready for a test by data persistence!
The wikikube-worker2088 host seems to have a different firmware from the rest of the Supermicro's config c:
Mon, Jul 27
Sorry for the late reply, there is a weirdness in Redfish that seemed related to the BMC's health, but you are right it seems working. Thanks!
Fixed also ml-serve[2009,2011].codfw.wmnet or wikikube-worker[2248,2254,2256-2257,2278,2298,2300-2301,2320].codfw.wmnet with the same fix!
And I managed to fix:
Managed to shrink the last list down to:
Fri, Jul 24
- Fixed cloudelastic1012.eqiad.wmnet,db1257.eqiad.wmnet (renamed "root" back to "ADMIN")
- Re-ran on ml-serve[1012,1014-1015].eqiad.wmnet,rdb[2011-2012].codfw.wmnet,wikikube-worker2332.codfw.wmnet, cirrussearch[1112,1115,1119,1121-1123].eqiad.wmnet,clouddb[1026-1033].eqiad.wmnet,cloudelastic1011.eqiad.wmnet,dbproxy[2005-2006].codfw.wmnet and it completed without issues.
- Tried to cold-reset the BMC on dbproxy[2007-2008].codfw.wmnet,dbproxy[1028-1029].eqiad.wmnet,druid-internal1006.eqiad.wmnet,ganeti-jumbo[1001-1002].eqiad.wmnet,ml-build1001.eqiad.wmnet,rdb2013.codfw.wmnet,wdqs[1033-1035].eqiad.wmnet,wikikube-ctrl2006.codfw.wmnet,wikikube-worker[2334,2346,2369].codfw.wmnet,wikikube-worker[1328-1334,1360-1374].eqiad.wmnet, with the hope that then next run of the cookbook will work better.
After a chat with Supermicro it seems that new firmwares (we don't know it if it is for all models, and the timeline) need something like the following http boot configuration sent to Redfish:
I reduced the "root" accounts to be removed to:
Thu, Jul 23
And the Trixie host:
I merged a couple of changes to skip the eventstreams image in docker report / debmonitor, we ca remove them once the new code is deployed.
I was able to cold-restart some Supermicro BMCs and the cookbook ran fine after that (namely, the root account was removed successfully). Remaining hosts to check:
Wed, Jul 22
@dduvall the alternative that I can think of is to get the whole list of images on the registry, remove the ones that we know from their names are really old, and copy all image:tags combinations. We'll surely copy some old stuff but potentially we'll avoid troubles.
@dduvall Hi! I'd need some help from your team :) Ideally we'd need to see if in the above list there is something that we want to copy over anyway, because a CI run may not have triggered in the past two weeks. I see a lot of stale images (at least from the names), what do you think?
I took the images listed in the Docker Registry's web UI and compared with the ones above, these are the images NOT PRESENT in the last two weeks worth of logs:
The task is not super critical but we have multiple regular reports about what Debian packages are installed for various clusters that are failing because of this image, so ideally the sooner we fix this the better :)
I used the following one liner to get images from the Docker Registry's logs on registry200[4,5] (the active/pooled ones):
Very nice results for the /v2/wikimedia/machinelearning.* prefix in T428022#12141795
@RobH sadly the host is still WIP, I have an open email thread with Supermicro about some weird firmware issues, and it doesn't seem near to a resolution. I pinged them again, maybe we can try to expedite the resolution but it has been ongoing for months. Is there a maximum deadline for the decision?
Tue, Jul 21
Worst case scenario, rollback: https://gerrit.wikimedia.org/r/1312538
I just deployed the nginx change and performed the following tests:
- Renamed root to ADMIN on clouddb1025.eqiad.wmnet and dse-k8s-worker1014.eqiad.wmnet
- dse-k8s-wdqs1002.eqiad.wmnet, logging-hd[1004-1005].eqiad.wmnet just needed a re-run
pki-db-1 and pki-root-1 replaced the last buster VMs. I already configured everything and made sure that the new pki-root vm contains the same files under /etc/cfss as the old VM. I am going to leave the old VMs around for a couple of weeks just to be sure, and then I'll delete them.
Second run on bookworm hosts:
@bking my understanding is that the host changes IP address so IPVS/Pybal need to be fixed to list the new value. What is the use case of not using depool before a reimage?
Mon, Jul 20
- Manually removed the root account from aqs1022.eqiad.wmnet,deploy1003.eqiad.wmnet,ms-be[2084-2087].codfw.wmnet,ms-be1091.eqiad.wmnet,restbase[1043-1045].eqiad.wmnet
- The new version of the cookbook fixed aqs1023.eqiad.wmnet,parsoidtest1001.eqiad.wmnet,restbase2039.codfw.wmnet
First run on all Bullseye nodes:
Fri, Jul 17
Got pinged by Aiko on this one, I think the work is already completed. The specific gerrit patch in the task's description was built in T423459#11843102, but it wasn't enough to solve the problem.
The cookbook is ready to be used to rollout the wmfroot user, and to do the other following misc things:
Hey folks! IIUC this task is to track both Bookworm's and Trixie's updates, but I don't see a lot of hosts mentioned in T431659 (and a lot of ML ones, CC @klausman @DPogorzelski-WMF)
Wed, Jul 15
Today I tried the skopeo tool on registry1004, that seems really nice:
Jul 14 2026
There is still one subtask opened but it is a cookbook follow up. All the clusters in deployment-prep and production are now running Kafka 3.7!
Jul 8 2026
I tried the new cookbook with an-test* and this is the first issue:
@klausman for RR wikidata sometimes it fails and sometimes it doesn't, but I noticed this in the logs:
Back to this! I created https://gitlab.wikimedia.org/elukey/registry-clone to support the move of the Docker images. My plan is the following:
@klausman recommendation-api-ng.discovery.wmnet is a CNAME to k8s-ingress-ml-serve.discovery.wmnet, it is not part of the inference.discovery.wmnet's endpoints (see the TLS error). You'd need to instruct httpb to use a different --host for that use case, the ingress endpoint should be fine. No idea for rr-wikidata :)
Jul 7 2026
Jul 6 2026
@BTullis I haven't really stated that the current setup is a poc, but it started as one and we spent many hours refining it. And it has been working really well so far :)
Everything rolled out, I also tested a reimage and it worked nicely and as expected. I am going to ping Filippo for Cloud cumin.
Didn't realize this was closed. Reopening it to complete the above :)
@Lferreira Hi! I noticed that Tyler is on sabbatical, can I ping to you for this task?
Jul 3 2026
Hi Ben! Thanks for starting the conversation :)
Yesterday I met with Tiziano in person and we had the chance to discuss this project, and the requirements from Alert Manager. I am going to add some highlights to avoid forgetting them:
Jul 2 2026
To keep archives happy, with Puppet 7 on Debian Bookworm:
Jul 1 2026
3.0.0 deployed fleetwide, and 3.1.0 is ready to go. @MoritzMuehlenhoff we can do the 3.1.0 rollout next week, hopefully this will give enough time to people to find bugs etc.. with 3.0.0. Does it sound good?
Ok this was a pebcak due to me not knowing how bdist_wheel works. I thought I needed to add the wheel package on one of the pywmflib's venvs specified in setup.py, to support bdist_wheel, and in doing so setup.py was modified. Since bdist_wheel noticed a diff compared to the last git tag, it assumed it was a dev version adding the suffix. I just needed to add the wheels package in the python-release's venv (see above patch).
Jun 30 2026
Pretty sure this is not an issue anymore, resolving, please re-open if I am mistaken.
Resoling this, since most of the work is done and only a couple of tasks are left to do.
Declining, please re-open if needed!
Tentatively declining, please reopen if necessary!