See parent and grand-parent tasks T426197: codfw: pod AB switches upgrade (2026)
This task is to schedule the software upgrade of rack B7 top of rack switch scheduled for Tuesday July 21st at 14:00 UTC with an expected network connectivity loss of ~20min
https://wikitech.wikimedia.org/wiki/Network_leaf_maintenance
Depool needed
wikikube-ctrl2001: depool using k8s
wikikube-worker2139: depool using k8s
wikikube-worker2140: depool using k8s
wikikube-worker2157: depool using k8s
wikikube-worker2284: depool using k8s
wikikube-worker2285: depool using k8s
ml-serve2009: depool using k8s
ml-staging2003: depool using k8s
kubestage2003: depool using k8s
pc2017: skipping host (manual depool needed)
db2229: skipping host (manual depool needed) T430964 - depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
es2046: skipping host (manual depool needed)
cirrussearch2079: depool using local_command depool
cirrussearch2080: depool using local_command depool
db2228: depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
db2242: depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
ganeti2032: skipping host (Use sre.ganeti.drain-node, primaries need to be failed-over too)
ganeti2033: Depooled (special-case two node cluster with routed Ganeti)
ganeti2049: skipping host (Use sre.ganeti.drain-node, primaries need to be failed-over too)
mc-gp2005: service ops are prepping a patch
Depool not needed
logging-hd2004: skipping host (No cookbook, no depool needed but there's a switch we can flip to mitigate the churn caused when the cluster detects a down node)
logstash2036: skipping host (No cookbook, no depool needed but there's a switch we can flip to mitigate the churn caused when the cluster detects a down node)
mc2046: skipping host (no depool needed)
ms-be2085: skipping host (Can't be depooled, need to go down one at a time with special care)
sretest2001: Couldn't get or parse depool Hiera key
db2230: nothing apart from downtime needed - testing host.
backup2012: Should be ok for a small downtime (repo backups)
cloudbackup2003: Should be ok for a small interruption (cc cloud-admin)
alert2002: Most likely not an issue - all clients are aware of both alertmanager instances.
dbproxy2006: not in use - otherwise we could make dns patch to change cname
Per team grouping
observability : alert2002, logging-hd2004, logstash2036
Data-Persistence : backup2012, db2228, db2229, db2230, db2242, dbproxy2006, es2046, ms-be2085, pc2017
Data-Platform-SRE : cirrussearch2079, cirrussearch2080
WMCS: cloudbackup2003
Infrastructure-Foundations : ganeti2032, ganeti2033, ganeti2049, sretest2001
ServiceOps : kubestage2003, mc2046, mc-gp2005, wikikube-ctrl2001, wikikube-worker2139, wikikube-worker2140, wikikube-worker2157, wikikube-worker2284, wikikube-worker2285
Machine-Learning-Team : ml-serve2009, ml-staging2003