Page MenuHomePhabricator

Reclaim aqs100[123]
Closed, ResolvedPublic

Description

Hello DC-Ops,

the old AQS cluster composed by aqs100[123] has been decomissioned by Analytics and the hosts have now role spare as indicated in the docs. It would be great to return this hardware to the spare pool for further usage.

Note: the Services team is interested to use some hardware for Restbase testing, just mentioning in case you'd want to follow up with them for a later re-assignment.

Thanks!

reclaim to spares steps

All steps are documented on the Server Lifecycle page.

aqs1001:

  • - disable all service level checks on hosts
  • - ensure system is no longer in any service pool
  • - remove from all dsh node lists
  • - remove all host hiera data
  • - disable puppet after role::spare run (have to do the next few steps immediately following puppet halt)
  • - system shutdown
  • - remove all puppet references
  • - puppet node clean/deactivate
  • - puppet run to update icinga
  • - salt key removal
  • - system network port disabled
  • - system disks wiped
  • - upgrade disk controllers from H310 to H710
  • - update racktables to list as asset tag name
  • - assign task back to @RobH for immediate allocation on T136340

aqs1002:

  • - disable all service level checks on hosts
  • - ensure system is no longer in any service pool
  • - remove from all dsh node lists
  • - remove all host hiera data
  • - disable puppet after role::spare run (have to do the next few steps immediately following puppet halt)
  • - system shutdown
  • - remove all puppet references
  • - puppet node clean/deactivate
  • - puppet run to update icinga
  • - salt key removal
  • - system network port disabled
  • - system disks wiped
  • - upgrade disk controllers from H310 to H710
  • - update racktables to list as asset tag name
  • - assign task back to @RobH for immediate allocation on T136340

aqs1003:

  • - disable all service level checks on hosts
  • - ensure system is no longer in any service pool
  • - remove from all dsh node lists
  • - remove all host hiera data
  • - disable puppet after role::spare run (have to do the next few steps immediately following puppet halt)
  • - system shutdown
  • - remove all puppet references
  • - puppet node clean/deactivate
  • - puppet run to update icinga
  • - salt key removal
  • - system network port disabled
  • - system disks wiped
  • - upgrade disk controllers from H310 to H710
  • - update racktables to list as asset tag name
  • - assign task back to @RobH for immediate allocation on T136340

Event Timeline

elukey triaged this task as Medium priority.Oct 12 2016, 11:58 AM

These nodes have half the RAM of the proposed AMS nodes, and (I just learned), have H310 raid controllers, which are apparently notorious for "extremely poor performance".

These nodes have half the RAM of the proposed AMS nodes, and (I just learned), have H310 raid controllers, which are apparently notorious for "extremely poor performance".

So, to follow up here:

  • Pros: alleviates an additional burden on @mark while improving turn-around time in the event of hardware issues, and avoids potential PII issues with test data
  • Cons: Has half the memory of the proposed AMS nodes, and very poor disk performance

As much as I'd like to avoid the pitfalls (for everyone) of the ams-based machines, aqs100[1-3] seem sufficiently problematic that we run the risk of failing at our primary objective (to create a test environment with similar performance characteristics to production). I think my preference would be to stick to plan A.

Indeed, @Eevans . Getting the nodes just for the sake of it doesn't seem prudent, especially since it would not allow us to synthesise production load no more than we are capable of doing so now.

reading "have role spare" and then "return to spare pool", doesn't that mean it's already in the spare pool?

reading "have role spare" and then "return to spare pool", doesn't that mean it's already in the spare pool?

Hey Daniel, I didn't know if there was paperwork between the "role spare" and the "spare pool" assignment :)

reading "have role spare" and then "return to spare pool", doesn't that mean it's already in the spare pool?

No, it doesn't. "role spare" is used for the time until they're returned to spares/shutdown (since that can take quite some time, especially in esams).

RobH added subscribers: Cmjohnson, RobH.

These need to be reclaimed to spare (back to asset tag use only, no hostname) and then moved into use for T136340.

Please upgrade the controllers from H310 to H710 controllers.

Assigning to @Cmjohnson (after I remove all the puppet/salt stuff) so he can reclaim to spare (with disk wipes and such.) then please upgrade the controller and let me know!

RobH added a project: ops-eqiad.

Adding ops-eqiad tag since now its pending onsite actions.

just chatted with chris, disks still need to be wiped. assigning back to him until that is completed

faidon mentioned this in Unknown Object (Task).Nov 3 2016, 3:24 PM

@Cmjohnson: Are these R720xd's able to be retrofitted with SFF SSDs without issue? Please advise.

If so, we'll need to order some Intel S3610 SSDs to place into these for their work with services (linked task T136340)

Sure, they can be retrofitted I have plenty of LFF/SFF adapters .....but are you sure we want to do that?....they're out of warranty and aqs1003.

Oh, these are LFF bay systems? (I didn't think the R720xd could retrofit SFF into LFF bays on the hot swap chassis.) At LFF, its only 12 bays correct?

My understanding is services wants these since they are powerful enough for their testing systems. Since it won't be live traffic, but a test bed, the lack of warranty is less of an issue.

If we can swap in SFF disks into these without an issue, I'll price out the options for SSDs on a procurement task.

Updated to my comment about the ssd's fitting for the aqs systems. I attached the ssd with the SFF adapter to the disk caddy for the R720XD and the ssds with that SFF will not work.

The setup task of these systems in their reclaimed role has been setup. I had this task open as a reminder until I completed that, so resolving.

aqs100[1-3] are still listed in site.pp

Indeed, a quick grep shows no other entries. Removed.