The following hosts will be deprecated thanks to the purchase of the new dbprov* and dedicated dumps replicas:
[] dbstore1001
Steps for service owner:
[] - all system services confirmed offline from production use
[] - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place.
[] - remove system from all lvs/pybal active configuration
[] - any service group puppet/hiera/dsh config removed
[] - remove site.pp, replace with role(spare::system)
[] - unassign service owner from this task, check off completed steps, and assign to @robh for followup on below steps.
Steps for DC-Ops:
The following steps cannot be interrupted, as it will leave the system in an unfinished state.
**Start non-interrupt steps:**
[] - disable puppet on host
[] - power down host
[] - update netbox status to Inventory (if decom) or Planned (if spare)
[] - disable switch port
[] - switch port assignment noted on this task (for later removal)
[] - remove all remaining puppet references (include role::spare)
[] - remove production dns entries
[] - puppet node clean, puppet node deactivate (handled by wmf-decommission-host)
[] - remove dbmonitor entries on neodymium/sarin: sudo curl -X DELETE https://debmonitor.discovery.wmnet/hosts/${HOST_FQDN} --cert /etc/debmonitor/ssl/cert.pem --key /etc/debmonitor/ssl/server.key (handled by wmf-decommission-host)
**End non-interrupt steps.**
[] - system disks wiped (by onsite)
[] - IF DECOM: system unracked and decommissioned (by onsite), update racktables with result
[] - IF DECOM: switch port configration removed from switch once system is unracked.
[] - IF DECOM: add system to decommission tracking google sheet
[] - IF DECOM: mgmt dns entries removed.
[] - IF RECLAIM: system added back to spares tracking (by onsite)
[] dbstore2001
Steps for service owner:
[] - all system services confirmed offline from production use
[] - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place.
[] - remove system from all lvs/pybal active configuration
[] - any service group puppet/hiera/dsh config removed
[] - remove site.pp, replace with role(spare::system)
[] - unassign service owner from this task, check off completed steps, and assign to @robh for followup on below steps.
Steps for DC-Ops:
The following steps cannot be interrupted, as it will leave the system in an unfinished state.
**Start non-interrupt steps:**
[] - disable puppet on host
[] - power down host
[] - update netbox status to Inventory (if decom) or Planned (if spare)
[] - disable switch port
[] - switch port assignment noted on this task (for later removal)
[] - remove all remaining puppet references (include role::spare)
[] - remove production dns entries
[] - puppet node clean, puppet node deactivate (handled by wmf-decommission-host)
[] - remove dbmonitor entries on neodymium/sarin: sudo curl -X DELETE https://debmonitor.discovery.wmnet/hosts/${HOST_FQDN} --cert /etc/debmonitor/ssl/cert.pem --key /etc/debmonitor/ssl/server.key (handled by wmf-decommission-host)
**End non-interrupt steps.**
[] - system disks wiped (by onsite)
[] - IF DECOM: system unracked and decommissioned (by onsite), update racktables with result
[] - IF DECOM: switch port configration removed from switch once system is unracked.
[] - IF DECOM: add system to decommission tracking google sheet
[] - IF DECOM: mgmt dns entries removed.
[] - IF RECLAIM: system added back to spares tracking (by onsite)
[] dbstore2002
Steps for service owner:
[] - all system services confirmed offline from production use
[] - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place.
[] - remove system from all lvs/pybal active configuration
[] - any service group puppet/hiera/dsh config removed
[] - remove site.pp, replace with role(spare::system)
[] - unassign service owner from this task, check off completed steps, and assign to @robh for followup on below steps.
Steps for DC-Ops:
The following steps cannot be interrupted, as it will leave the system in an unfinished state.
**Start non-interrupt steps:**
[] - disable puppet on host
[] - power down host
[] - update netbox status to Inventory (if decom) or Planned (if spare)
[] - disable switch port
[] - switch port assignment noted on this task (for later removal)
[] - remove all remaining puppet references (include role::spare)
[] - remove production dns entries
[] - puppet node clean, puppet node deactivate (handled by wmf-decommission-host)
[] - remove dbmonitor entries on neodymium/sarin: sudo curl -X DELETE https://debmonitor.discovery.wmnet/hosts/${HOST_FQDN} --cert /etc/debmonitor/ssl/cert.pem --key /etc/debmonitor/ssl/server.key (handled by wmf-decommission-host)
**End non-interrupt steps.**
[] - system disks wiped (by onsite)
[] - IF DECOM: system unracked and decommissioned (by onsite), update racktables with result
[] - IF DECOM: switch port configration removed from switch once system is unracked.
[] - IF DECOM: add system to decommission tracking google sheet
[] - IF DECOM: mgmt dns entries removed.
[] - IF RECLAIM: system added back to spares tracking (by onsite)
Also, es2001,2,3,4 are being used as temporary disk hosts to hold short-term backups before sending them to bacula. They should be able to disappear (although those may require the completion of the extra bacula shelf for backups), so it is being handled at at later time at T222592
This task is to identify blockers for its decommission (e.g. transfer away all valuable data, such as recent backups, new hw setup, etc.) and then hand it to dc ops.