Page MenuHomePhabricator

eqiad row A&B host migration details request for Service Ops
Open, Stalled, MediumPublic

Description

Summary:

Hosts: 112 total hosts, see google sheet

Details:
Please note we're currently obtaining service group and individual server depool directions for the eqiad row A/B network migration. This migration will follow the same general cadence of the eqiad C/D migration, which will be detailed below.

All hosts are listed on the google sheet, and we ask you update each hosts line for your group with either full directions for the migration, or a note that the migration must be scheduled in advance. If the migration must be scheduled, please detail if you (the service owner or team member) will be in attendance for the migration, or if the directions and date/time will be coordiated to DC Ops.

Current Row A & B Hosts Google Sheet - Please update this sheet: https://docs.google.com/spreadsheets/d/1jUbkqTwWPYq-eIdKI7M56y7VvzaWTPH4c2BPMfoGeBk/edit?usp=sharing

Past C & D Hosts Google Sheet for past examples - do not update this sheet: https://docs.google.com/spreadsheets/d/13ow4JxrsQdz8KSsdBBNwvlrAuGKo8OHWcnR4RhXTYc0/edit?usp=sharing

Projected Target Start Date: After the eqiad to codfw failover in September 2026.

The service owners for a given host will decide if the host can be migrated by DC ops directly or should have a scheduled window for the host's migration.

DC Operations direct migration:

  • Service Owners must update full [de|re]pool directions on the google sheet.
  • DC Ops will follow directions to move hosts at the date/time of DC ops choosing (post projected start date listed above.)
  • Unless otherwise noted, service owner not required to attend actual migration.
  • Assumed host must have cookbook sre.hosts.downtime, in addition to any further instructions.
  • This can include special notes. Examples: 'only 1 cp host depooled in text or upload at a time', 'only one dbproxy host per day', 'Can be done anytime - just needs cookbook sre.hosts.downtime'
  • Please provide a point of contact on your team per host.

Scheduled migration:

  • Service Owners must update full DC Operations side directions to the google sheet and propose a date/time window for DC ops review.
  • Please note if the migration must be attended by a member of the service owner team.
  • Please provide a point of contact on your team per host.

Event Timeline

RobH assigned this task to Kappakayala.
RobH triaged this task as Medium priority.

Kavitha,

Please review the task description and assign this within your team as needed for details on the hosts listed. Thanks in advance!

Kavitha,

Please review the task description and assign this within your team as needed for details on the hosts listed. Thanks in advance!

I did a first pass in the doc for Row A & B Hosts, we should do another one sometime in September

@jasmine_ as this is tied to DC Switchover, can I assign it to you to keep it top of mind and coordinate in September?

jijiki changed the task status from Open to Stalled.Tue, Jul 21, 1:35 PM
jijiki moved this task from Inbox to Scheduled (this Q) on the ServiceOps board.

Clarification Questions:

"Depool using k8s" This doesn't include the full directions and cookbook commands, can this be provided?

I've also broken it back up from 1 giant cell to 1 cell per line, as the document is resorted, we cannot have merged cells.

Can someone detail here exactly what 'Depool using k8s' entails, all the commands?

@jasmine_ as this is tied to DC Switchover, can I assign it to you to keep it top of mind and coordinate in September?

Sure, sounds good~

Clarification Questions:

"Depool using k8s" This doesn't include the full directions and cookbook commands, can this be provided?

I've also broken it back up from 1 giant cell to 1 cell per line, as the document is resorted, we cannot have merged cells.

Can someone detail here exactly what 'Depool using k8s' entails, all the commands?

Thanks for asking! - I've edited the rows to note depooling with the sre.k8s.pool-depool-node cookbook, which will handle the k8s drain and cordon logic for these nodes. Feel free to let me know if any further specification would be helpful here.

Thank you for the clarification! This seems straightforward enough. Once we get to the actual migration dates we'll try one or two and if no problems continuie normally. If any issues, we'll reach back out to you via irc and/or this phab task.