Page MenuHomePhabricator

Reclaim parsoidtest1001 and testreduce1001
Open, In Progress, MediumPublicRequest

Description

This task will track the RETIREMENT and RECLAIM of server parsoidtest1001.eqiad.wmnet and testreduce1001.eqiad.wmnet

With the launch of updates to the decom cookbook, the majority of these steps can be handled by the service owners directly. The DC Ops team only gets involved once the system has been fully removed from service and powered down by the decommission cookbook.

parsoidtest1001

Steps for service owner:

  • - all system services confirmed offline from production use
  • - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place. (likely done by script)
  • - remove system from all lvs/pybal active configuration
  • - any service group puppet/hiera/dsh config removed
  • - login to cumin host and run the decom cookbook: cookbook sre.hosts.decommission <host fqdn> -t <phab task>. This does: bootloader wipe, host power down, netbox update to decommissioning status, puppet node clean, puppet node deactivate, debmonitor removal, and run homer.
  • - remove all remaining puppet references and all host entries in the puppet repo
  • - reassign task from service owner to no owner and ensure the site project (ops-sitename depending on site of server) is assigned.

End service owner steps / Begin DC-Ops team steps:

  • - system disks removed (by onsite)
  • - determine system age, under 5 years are reclaimed to spare, over 5 years are decommissioned.
  • - IF DECOM: system unracked and decommissioned (by onsite), update netbox with result and set state to offline
  • - IF DECOM: mgmt dns entries removed.
  • - IF RECLAIM: set netbox state to 'inventory' and hostname to asset tag
testreduce1001

Steps for service owner:

  • - all system services confirmed offline from production use
  • - set all icinga checks to maint mode/disabled while reclaim/decommmission takes place. (likely done by script)
  • - remove system from all lvs/pybal active configuration
  • - any service group puppet/hiera/dsh config removed
  • - login to cumin host and run the decom cookbook: cookbook sre.hosts.decommission <host fqdn> -t <phab task>. This does: bootloader wipe, host power down, netbox update to decommissioning status, puppet node clean, puppet node deactivate, debmonitor removal, and run homer.
  • - remove all remaining puppet references and all host entries in the puppet repo
  • - reassign task from service owner to no owner and ensure the site project (ops-sitename depending on site of server) is assigned.

End service owner steps / Begin DC-Ops team steps:

  • - system disks removed (by onsite)
  • - determine system age, under 5 years are reclaimed to spare, over 5 years are decommissioned.
  • - IF DECOM: system unracked and decommissioned (by onsite), update netbox with result and set state to offline
  • - IF DECOM: mgmt dns entries removed.
  • - IF RECLAIM: set netbox state to 'inventory' and hostname to asset tag

Event Timeline

jijiki added a subscriber: Muehlenhoff.

@Muehlenhoff following up on our discussion, and unless there are no objections, we could repurpose this machine for the ganeti cluster on eqiad. It was purchased on 2024-06-26.

@Muehlenhoff following up on our discussion, and unless there are no objections, we could repurpose this machine for the ganeti cluster on eqiad. It was purchased on 2024-06-26.

Sounds good to me!

@Muehlenhoff following up on our discussion, and unless there are no objections, we could repurpose this machine for the ganeti cluster on eqiad. It was purchased on 2024-06-26.

@wiki_willy Would that be okay? How do we track this?

Procedure-wise Effie could simply complete the decom steps until "remove all remaining puppet references and all host entries in the puppet repo" and we create a task with DC ops to re-label the server and re-provision under the new name?

@MoritzMuehlenhoff if current location location is ok. Could it just be renamed and reimaged by service owner?

@MoritzMuehlenhoff - sure, no problem. Repurposing it works for us as well. As long as we have all the updates in Phabricator, I think that's all we really need for tracking. Procedural wise (to John's point above), treating it as a rename/reimage would work for us too.

@Muehlenhoff following up on our discussion, and unless there are no objections, we could repurpose this machine for the ganeti cluster on eqiad. It was purchased on 2024-06-26.

@wiki_willy Would that be okay? How do we track this?

Procedure-wise Effie could simply complete the decom steps until "remove all remaining puppet references and all host entries in the puppet repo" and we create a task with DC ops to re-label the server and re-provision under the new name?

jijiki changed the task status from Open to Stalled.Jun 10 2026, 2:18 PM
jijiki added a subscriber: Jgiannelos.

Hey folks, the last bits are being wrapped up by @Jgiannelos. I am happy to decom the server when we are good to go, and hand it over to DC-ops. Sounds good?

I think this train is the last week we run both envs for our weekly testing (both internal parsoidtest and mw-parsoid). I have one more thing to check for our regression testing and then we are good to decommission.

I think this train is the last week we run both envs for our weekly testing (both internal parsoidtest and mw-parsoid). I have one more thing to check for our regression testing and then we are good to decommission.

Is this completed?

Porting here my comment from children task: given this service is part of Bullseye hosts, we'd need an ETA for full decommissioning. Thanks

Porting here my comment from children task: given this service is part of Bullseye hosts, we'd need an ETA for full decommissioning. Thanks

Due to T428909: Investigate X-Wikimedia-Debug routing for api.wikimedia.org and T431838: mw-parsoid endpoints are returning 404 for the rest.php endpoints used by rt-testing, folks had to use their old infra last week. We will try to wrap up this week.

As discussed with Effie yesterday, ETA is end of this week (Jul 24).

We just finished one more round of rt-testing after switching over to the old stack (T428909) and I've given a heads up to the team to backup any scripts on homedirs, so we should be good to decommision parsoidtestXXXX and testreduceXXXX

jijiki renamed this task from decommission parsoidtest1001.eqiad.wmnet to Reclaim parsoidtest1001.eqiad.wmnet.Tue, Jul 21, 10:32 AM
jijiki changed the task status from Stalled to In Progress.
jijiki updated the task description. (Show Details)

We just finished one more round of rt-testing after switching over to the old stack (T428909) and I've given a heads up to the team to backup any scripts on homedirs, so we should be good to decommision parsoidtestXXXX and testreduceXXXX

I will decomm both hosts next week, thank you!

jijiki renamed this task from Reclaim parsoidtest1001.eqiad.wmnet to Reclaim parsoidtest1001 and testreduce1001.Tue, Jul 21, 10:49 AM

Change #1320683 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] site.pp: retire parsoidtest and testreduce

https://gerrit.wikimedia.org/r/1320683

Change #1320683 merged by Effie Mouzeli:

[operations/puppet@production] site.pp: retire parsoidtest and testreduce

https://gerrit.wikimedia.org/r/1320683