- db2235 master
- db2160
- db1217
- db1164 old master
- db1250 future new master
Description
Description
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Resolved | Marostegui | T422365 Migration to Debian Trixie of production database-related hosts | |||
| Resolved | Marostegui | T430902 Migrate m5 to Debian Trixie | |||
| Resolved | Marostegui | T430903 Move db1228 to m5 | |||
| Resolved | Marostegui | T430934 db1228 crashed | |||
| Resolved | Marostegui | T432967 Switchover m5 master db1164 -> db1228 |
Event Timeline
Comment Actions
Cookbook cookbooks.sre.hosts.reimage was started by marostegui@cumin1003 for host db2235.codfw.wmnet with OS trixie
Comment Actions
Cookbook cookbooks.sre.hosts.reimage started by marostegui@cumin1003 for host db2235.codfw.wmnet with OS trixie completed:
- db2235 (PASS)
- Downtimed on Icinga/Alertmanager
- Disabled Puppet
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- Host up (Debian installer)
- Checked BIOS boot parameters are back to normal
- Host up (new fresh trixie OS)
- Generated Puppet certificate
- Signed new Puppet certificate
- Run Puppet in NOOP mode to populate exported resources in PuppetDB
- Found Nagios_host resource for this host in PuppetDB
- Downtimed the new host on Icinga/Alertmanager
- Removed previous downtime on Alertmanager (old OS)
- First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607231133_marostegui_2124714_db2235.out
- configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
- Rebooted
- Automatic Puppet run was successful
- Forced a re-check of all Icinga services for the host
- Icinga status is optimal
- Icinga downtime removed
- Updated Netbox data from PuppetDB
Comment Actions
The master was switched, I am going to give the new master a few hours before I reimage the old one.
Comment Actions
Cookbook cookbooks.sre.hosts.reimage was started by marostegui@cumin1003 for host db1164.eqiad.wmnet with OS trixie
Comment Actions
Cookbook cookbooks.sre.hosts.reimage started by marostegui@cumin1003 for host db1164.eqiad.wmnet with OS trixie completed:
- db1164 (WARN)
- Downtimed on Icinga/Alertmanager
- Disabled Puppet
- Removed from Puppet and PuppetDB if present and deleted any certificates
- Removed from Debmonitor if present
- Forced PXE for next reboot
- Host rebooted via IPMI
- Host up (Debian installer)
- Checked BIOS boot parameters are back to normal
- Host up (new fresh trixie OS)
- Generated Puppet certificate
- Signed new Puppet certificate
- Run Puppet in NOOP mode to populate exported resources in PuppetDB
- Found Nagios_host resource for this host in PuppetDB
- Downtimed the new host on Icinga/Alertmanager
- Removed previous downtime on Alertmanager (old OS)
- First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607290605_marostegui_3571245_db1164.out
- configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
- Rebooted
- Automatic Puppet run was successful
- Forced a re-check of all Icinga services for the host
- Icinga status is not optimal, downtime not removed
- Updated Netbox data from PuppetDB