We're testing Debian Trixie in our DB infra, so far we've tested it on db*, es* and currently doing dbproxy*
I believe we should go ahead and get one of the new hosts at T409162 to be reimaged with Trixie before we get them with data, so we know if everything will work out of the box
Description
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Open | None | T428047 Upgrade clouddb* hosts to Debian Trixie | |||
| Unknown Object (Task) | |||||
| Resolved | Jclark-ctr | T409162 Q2:rack/setup/install clouddb1026-1033 | |||
| Resolved | Request | Jhancock.wm | T408407 decommission es2028 | ||
| Resolved | Marostegui | T422365 Migration to Debian Trixie of production database-related hosts | |||
| Resolved | Marostegui | T406981 Compile and test MariaDB 10.11 on Debian Trixie | |||
| Resolved | Marostegui | T407472 Install a testing db with Debian Trixie | |||
| Resolved | fnegri | T415165 Install a clouddb host with Debian Trixie | |||
| Resolved | Jclark-ctr | T422813 clouddb1019 down | |||
| Resolved | Request | Jclark-ctr | T423151 decommission clouddb1019.eqiad.wmnet | ||
| Resolved | Marostegui | T426842 Package wmf-pt-kill to debian trixie | |||
| Resolved | Marostegui | T410369 Install Debian Trixie on one s1 host | |||
| Resolved | Jclark-ctr | T410388 PXE failing on db1169 |
Event Timeline
If this goes well, any reason not to install all of those on Trixie from the start to avoid having to re-image them later?
I would suggest:
- reimage one host
- install mariadb and set up replication on that host
- if it all goes well, reimage following hosts
fnegri moved this task from Inbox to Watching on the cloud-services-team board.
According to https://wikitech.wikimedia.org/wiki/Portal:Data_Services/Admin/Wiki_Replicas#Who_admins_what the reimage itself is done by WMCS, so I moved it to the wrong column.
Change of plans, @Marostegui will reimage clouddb1019 instead, as that one needs to be reimaged anyway because of a hardware issue: T422813: clouddb1019 down.
Unfortunately I don't think clouddb1019 will be back (or for long): https://phabricator.wikimedia.org/T422813#11807424 - should we go back to reimage clouddb1015 and if 1019 ever comes back, we can also do that too.
should we go back to reimage clouddb1015
clouddb1015 is the only clouddb with s4 and s6, until we get the replacement for clouddb1019.
We can announce a downtime for the reimage, but I was thinking that maybe we could first test the trixie upgrade on another clouddb which is less critical, e.g. clouddb1022.
A temporary replacement for clouddb1019 was found (see T409557: Productionize new clouddb* hosts (clouddb1022-1033)) so we can go back again to the original plan of reimaging clouddb1015 to trixie, without requiring downtime. I will do it this week or the next.
Mentioned in SAL (#wikimedia-operations) [2026-05-20T09:57:42Z] <fnegri@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1015.eqiad.wmnet with reason: Rebooting clouddb1015 T415165
Cookbook cookbooks.sre.hosts.reimage was started by fnegri@cumin1003 for host clouddb1015.eqiad.wmnet with OS trixie
Cookbook cookbooks.sre.hosts.reimage started by fnegri@cumin1003 for host clouddb1015.eqiad.wmnet with OS trixie executed with errors:
- clouddb1015 (FAIL)
- Downtimed on Icinga/Alertmanager
- Disabled Puppet
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console clouddb1015.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.
Cookbook cookbooks.sre.hosts.reimage was started by fnegri@cumin1003 for host clouddb1015.eqiad.wmnet with OS trixie
Reimage completed, Mariadb is running, but puppet is failing with:
E: Unable to locate package wmf-pt-kill
I'm gonna copy that package from bookworm to trixie.
fnegri@apt1002:~$ sudo -i reprepro copy trixie-wikimedia bookworm-wikimedia wmf-pt-kill
There's a broken dependency:
The following packages have unmet dependencies:
wmf-pt-kill : Depends: sysuser-helper (< 1.4) but it is not going to be installed
E: Unable to correct problems, you have held broken packages.
E: The following information from --solver 3.0 may provide additional context:
Unable to satisfy dependencies. Reached two conflicting decisions:
1. wmf-pt-kill:amd64=3.1.0-1+wmf6 is selected for install
2. wmf-pt-kill:amd64 Depends sysuser-helper (< 1.4)
but none of the choices are installable:
[no choices]Percona Toolkit, which this package is a fork of, doesn't depend on that for trixie: https://packages.debian.org/trixie/percona-toolkit so I don't know but my guess is it could be ignored on the package, although it would could just in case some update + repackage.
I found the source at https://gerrit.wikimedia.org/r/q/project:operations/debs/wmf-pt-kill but I don't know what procedure should be followed for building it, could you or @Marostegui please rebuild the package for trixie?
wikireplicas-utils was also missing in trixie, in this case a simple copy worked:
fnegri@apt1002:~$ sudo -i reprepro copy trixie-wikimedia bookworm-wikimedia wikireplicas-utilsCookbook cookbooks.sre.hosts.reimage started by fnegri@cumin1003 for host clouddb1015.eqiad.wmnet with OS trixie completed:
- clouddb1015 (PASS)
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- Host up (Debian installer)
- Checked BIOS boot parameters are back to normal
- Host up (new fresh trixie OS)
- Generated Puppet certificate
- Signed new Puppet certificate
- Run Puppet in NOOP mode to populate exported resources in PuppetDB
- Found Nagios_host resource for this host in PuppetDB
- Downtimed the new host on Icinga/Alertmanager
- First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202605201044_fnegri_1941956_clouddb1015.out
- configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
- Rebooted
- Automatic Puppet run was successful
- Forced a re-check of all Icinga services for the host
- Icinga status is optimal
- Icinga downtime removed
- Updated Netbox data from PuppetDB
Cookbook cookbooks.sre.hosts.reimage started by fnegri@cumin1003 for host clouddb1015.eqiad.wmnet with OS trixie executed with errors:
- clouddb1015 (FAIL)
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- Host up (Debian installer)
- Checked BIOS boot parameters are back to normal
- Host up (new fresh trixie OS)
- Generated Puppet certificate
- Signed new Puppet certificate
- Run Puppet in NOOP mode to populate exported resources in PuppetDB
- Found Nagios_host resource for this host in PuppetDB
- Downtimed the new host on Icinga/Alertmanager
- First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202605201044_fnegri_1941956_clouddb1015.out
- configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
- Rebooted
- Automatic Puppet run was successful
- Forced a re-check of all Icinga services for the host
- Icinga status is optimal
- Icinga downtime removed
- Updated Netbox data from PuppetDB
- The reimage failed, see the cookbook logs for the details. You can also try typing "sudo install-console clouddb1015.eqiad.wmnet" to get a root shell, but depending on the failure this may not work.
The cookbook did actually PASS, but an exception was raised while writing the PASS comment (that you can see above), which caused the following FAIL message to be posted.
It was probably a timeout while reading the response from Phabricator.
File "/usr/lib/python3.11/ssl.py", line 1167, in read
return self._sslobj.read(len, buffer)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
TimeoutError: The read operation timed out
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/usr/lib/python3/dist-packages/spicerack/_menu.py", line 265, in _run
raw_ret = runner.run()
^^^^^^^^^^^^
File "/srv/deployment/spicerack/cookbooks/sre/hosts/reimage.py", line 910, in run
self.phabricator.task_comment(
File "/usr/lib/python3/dist-packages/wmflib/phabricator.py", line 274, in task_comment
raise PhabricatorError(message) from e
wmflib.phabricator.PhabricatorError: Unable to update Phabricator task 'T415165'Full spicerack logs: https://phabricator.wikimedia.org/P92684