Page MenuHomePhabricator

decommission db1171.eqiad.wmnet
Closed, ResolvedPublicRequest

Description

This task will track the decommission-hardware of server db1171.eqiad.wmnet.

With the launch of updates to the decom cookbook, the majority of these steps can be handled by the service owners directly. The DC Ops team only gets involved once the system has been fully removed from service and powered down by the decommission cookbook.

db1171

Steps for service owner:

  • - all system services confirmed offline from production use
  • - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place. (likely done by script)
  • - remove system from all lvs/pybal active configuration
  • - any service group puppet/hiera/dsh config removed
  • - login to cumin host and run the decom cookbook: cookbook sre.hosts.decommission <host fqdn> -t <phab task>. This does: bootloader wipe, host power down, netbox update to decommissioning status, puppet node clean, puppet node deactivate, debmonitor removal, and run homer.
  • - remove all remaining puppet references and all host entries in the puppet repo
  • - reassign task from service owner to no owner and ensure the site project (ops-sitename depending on site of server) is assigned.

End service owner steps / Begin DC-Ops team steps:

  • - system disks removed (by onsite)
  • - determine system age, under 5 years are reclaimed to spare, over 5 years are decommissioned.
  • - IF DECOM: system unracked and decommissioned (by onsite), update netbox with result and set state to offline
  • - IF DECOM: mgmt dns entries removed.

Related Objects

Event Timeline

jcrespo claimed this task.
jcrespo raised the priority of this task from Medium to High.

Mentioned in SAL (#wikimedia-operations) [2026-08-04T07:55:59Z] <jynus> running extra backups to test db1285 T433826

Change #1320684 had a related patch set uploaded (by Jcrespo; author: Jcrespo):

[operations/puppet@production] mariadb: Set db1171 as insetup for decommissioning

https://gerrit.wikimedia.org/r/1320684

Icinga downtime and Alertmanager silence (ID=646e8b85-582e-4091-9c21-56bc4d544ebc) set by jynus@cumin1003 for 2 days, 1:00:00 on 1 host(s) and their services with reason: decom

db1171.eqiad.wmnet

Change #1320684 merged by Jcrespo:

[operations/puppet@production] mariadb: Set db1171 as insetup for decommissioning

https://gerrit.wikimedia.org/r/1320684

Host set as insetup, removed from orchestrator and zarcillo, waiting a bit to ensure new backups flow normally before nuking it permanently.

Change #1320923 had a related patch set uploaded (by Jcrespo; author: Jcrespo):

[operations/puppet@production] mariadb: Remove all references on puppet to db1150 & db1171

https://gerrit.wikimedia.org/r/1320923

Change #1320923 merged by Jcrespo:

[operations/puppet@production] mariadb: Remove all references on puppet to db1150 & db1171

https://gerrit.wikimedia.org/r/1320923

cookbooks.sre.hosts.decommission executed by jynus@cumin1003 for hosts: db1171.eqiad.wmnet

  • db1171.eqiad.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found physical host
    • Downtimed management interface on Alertmanager
    • Wiped all swraid, partition-table and filesystem signatures
    • Powered off
    • [Netbox] Set status to Decommissioning, deleted all non-mgmt IPs, updated switch interfaces (disabled, removed vlans, etc)
    • Configured the linked switch interface(s)
    • Removed from DebMonitor
    • Removed from Puppet server and PuppetDB
jcrespo updated the task description. (Show Details)
jcrespo added a project: ops-eqiad.

This is ready for dcops unracking or repurpose.

jcrespo lowered the priority of this task from High to Medium.Wed, Aug 5, 8:47 AM
VRiley-WMF updated the task description. (Show Details)

This has been completed.