This task will track the #decommission of server db2067.codfw.wmnet
The first 5 steps should be completed by the service owner that is returning the server to DC-ops (for reclaim to spare or decommissioning
With the launch of updates to the decom cookbook, dependent on server configuration and age.)the majority of these steps can be handled by the service owners directly. The DC Ops team only gets involved once the system has been fully removed from service and powered down by the decommission cookbook.
db2067
**Steps for service owner:**
[] - all system services confirmed offline from production use
[x] - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place. https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/552198/
[x] - remove system from all lvs/pybal active configuration
[x] - any service group puppet/hiera/dsh config removed
[x] - remove site.pp, replace with role(spare::system) recommended to ensure services offline but not 100% required as long as the decom script is IMMEDIATELY run below. replace with role(spare::system)https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/553444/
[] - unassign service owner from this task, check off completed steps, and assign to @robh for followup on below steps.
Steps for DC-Ops:
The following steps cannot be interruptedlogin to cumin host and run the decom cookbook: cookbook sre.hosts.decommission <host fqdn> -t <phab task>. This does: bootloader wipe, as it will leave the system in an unfinished state.
**Start non-interrupt steps:**
[] - disable puppet on host
[] - power down host
[] - update netbox status to Inventory (if decom) or Planned (if spare)host power down, netbox update to decommissioning status, puppet node clean, puppet node deactivate, debmonitor removal.
[] - disable switch port
[] - switch port assignment noted on this task (for later removal)- remove all remaining puppet references (include role::spare) and all host entries in the puppet repo
[] - remove all remaining puppet references (include role::spare)ALL dns entries except the asset tag mgmt entries.
[] - remove production dns entries
[] - puppet node cleanassign task from service owner to DC ops team member depending on site of server: codfw = @papaul, puppet node deactivate (handled by wmf-decommission-host)
[] - remove dbmonitor entries on neodymium/sarin: sudo curl -X DELETE https://debmonitor.discovery.wmnet/hosts/${HOST_FQDN} --cert /etc/debmonitor/ssl/cert.pem --key /etc/debmonitor/ssl/server.key (handled by wmf-decommission-host)eqiad = @Jclark-ctr, all other sites = @robh.
**End non-interruptservice owner steps / Begin DC-Ops team steps.**:**
[] - Label disk #8 as broken so it doe- disable switch port / set to asset tag if host isn't get re-usedbeing unracked / remove from switch if being unracked.
[] - system disks wiped (by onsite)
[] - IF DECOM:- determine system unracked andage, under 5 years are reclaimed to spare, over 5 years are decommissioned (by onsite). If uncertain, update ask @wiki_willy.
[] - IF DECOM: system unracktables with resulted and decommissioned (by onsite), update netbox with result and set state to offline
[] - IF DECOM: switch port configration removed from switch once system is unracked.
[] - IF DECOM: add system to decommission tracking google sheet
[] - IF DECOM: mgmt dns entries removed.
[] - IF RECLAIM: system added backet netbox state to spares tracking (by onsite) 'inventory' and hostname to asset tag