Page MenuHomePhabricator

codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC
Closed, ResolvedPublic

Description

See parent and grand-parent tasks T426197: codfw: pod AB switches upgrade (2026)

This task is to schedule the software upgrade of rack B7 top of rack switch scheduled for Tuesday July 21st at 14:00 UTC with an expected network connectivity loss of ~20min

https://wikitech.wikimedia.org/wiki/Network_leaf_maintenance

Depool needed
wikikube-ctrl2001: depool using k8s
wikikube-worker2139: depool using k8s
wikikube-worker2140: depool using k8s
wikikube-worker2157: depool using k8s
wikikube-worker2284: depool using k8s
wikikube-worker2285: depool using k8s
ml-serve2009: depool using k8s
ml-staging2003: depool using k8s
kubestage2003: depool using k8s
pc2017: skipping host (manual depool needed)
db2229: skipping host (manual depool needed) T430964 - depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
es2046: skipping host (manual depool needed)
cirrussearch2079: depool using local_command depool
cirrussearch2080: depool using local_command depool
db2228: depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
db2242: depool using cookbook sre.mysql.depool -r 'codfw rack B7 depool for maintenance' {name}
ganeti2032: skipping host (Use sre.ganeti.drain-node, primaries need to be failed-over too)
ganeti2033: Depooled (special-case two node cluster with routed Ganeti)
ganeti2049: skipping host (Use sre.ganeti.drain-node, primaries need to be failed-over too)
mc-gp2005: service ops are prepping a patch

Depool not needed
logging-hd2004: skipping host (No cookbook, no depool needed but there's a switch we can flip to mitigate the churn caused when the cluster detects a down node)
logstash2036: skipping host (No cookbook, no depool needed but there's a switch we can flip to mitigate the churn caused when the cluster detects a down node)
mc2046: skipping host (no depool needed)
ms-be2085: skipping host (Can't be depooled, need to go down one at a time with special care)
sretest2001: Couldn't get or parse depool Hiera key
db2230: nothing apart from downtime needed - testing host.
backup2012: Should be ok for a small downtime (repo backups)
cloudbackup2003: Should be ok for a small interruption (cc cloud-admin)
alert2002: Most likely not an issue - all clients are aware of both alertmanager instances.
dbproxy2006: not in use - otherwise we could make dns patch to change cname

Per team grouping
observability : alert2002, logging-hd2004, logstash2036
Data-Persistence : backup2012, db2228, db2229, db2230, db2242, dbproxy2006, es2046, ms-be2085, pc2017
Data-Platform-SRE : cirrussearch2079, cirrussearch2080
WMCS: cloudbackup2003
Infrastructure-Foundations : ganeti2032, ganeti2033, ganeti2049, sretest2001
ServiceOps : kubestage2003, mc2046, mc-gp2005, wikikube-ctrl2001, wikikube-worker2139, wikikube-worker2140, wikikube-worker2157, wikikube-worker2284, wikikube-worker2285
Machine-Learning-Team : ml-serve2009, ml-staging2003

Details

Due Date
Tue, Jul 21, 1:00 PM
Related Changes in Gerrit:

Event Timeline

jcrespo added a subscriber: CWilliams-WMF.

backup2012: it is repo backups, should be ok for a small downtime, at most 1 backup run will fail, it will be retried one hour later.

cmooney triaged this task as Medium priority.Jul 6 2026, 2:28 PM

Mentioned in SAL (#wikimedia-operations) [2026-07-09T08:26:37Z] <moritzm> failover Ganeti master in codfw to ganeti2048 T430928

Maintenance done, last repool in progress.

cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: testvm2007.codfw.wmnet

  • testvm2007.codfw.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found Ganeti VM
    • VM shutdown
    • Started forced sync of VMs in Ganeti cluster codfw02 to Netbox
    • Removed from DebMonitor
    • Removed from Puppet server and PuppetDB
    • VM removed
    • Started forced sync of VMs in Ganeti cluster codfw02 to Netbox

cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: testvm2008.wikimedia.org

  • testvm2008.wikimedia.org (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found Ganeti VM
    • VM shutdown
    • Started forced sync of VMs in Ganeti cluster codfw02 to Netbox
    • Removed from DebMonitor
    • Removed from Puppet server and PuppetDB
    • VM removed
    • Started forced sync of VMs in Ganeti cluster codfw02 to Netbox

Mentioned in SAL (#wikimedia-operations) [2026-07-09T10:21:42Z] <moritzm> failover Ganeti master in codfw/routed to ganeti2034 T430928

Draining ganeti2033.codfw.wmnet of running VMs

cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: testvm2005.codfw.wmnet

  • testvm2005.codfw.wmnet (PASS)
    • Downtimed host on Icinga/Alertmanager
    • Found Ganeti VM
    • VM shutdown
    • Started forced sync of VMs in Ganeti cluster codfw_test to Netbox
    • Removed from DebMonitor
    • Removed from Puppet server and PuppetDB
    • VM removed
    • Started forced sync of VMs in Ganeti cluster codfw_test to Netbox
hnowlan subscribed.

alert2002 *shouldn't* cause many problems as all clients are aware of both alert hosts. However, this *is* an alerting host so we're not 100% certain that there won't be some flapping or noise unfortunately.

cmooney set Due Date to Tue, Jul 21, 1:00 PM.Thu, Jul 16, 12:25 PM
cmooney renamed this task from codfw: rack B7 maintenance to codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC.Thu, Jul 16, 4:33 PM

Change #1313118 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] mcrouter_wancache: remove mc-gp2005 for rack maintenance

https://gerrit.wikimedia.org/r/1313118

Change #1313118 merged by Effie Mouzeli:

[operations/puppet@production] mcrouter_wancache: remove mc-gp2005 for rack maintenance

https://gerrit.wikimedia.org/r/1313118

Draining ganeti2032.codfw.wmnet of running VMs

Icinga downtime and Alertmanager silence (ID=da979790-60cc-4098-88a0-763d9ca77d1c) set by cmooney@cumin1003 for 1:00:00 on 29 host(s) and their services with reason: lsw1-b7-codfw JunOS upgrade

backup2012.codfw.wmnet,cirrussearch[2079-2080].codfw.wmnet,cloudbackup2003.codfw.wmnet,db[2228-2230,2242].codfw.wmnet,dbproxy2006.codfw.wmnet,es2046.codfw.wmnet,ganeti[2032-2033,2049].codfw.wmnet,kubestage2003.codfw.wmnet,logging-hd2004.codfw.wmnet,logstash2036.codfw.wmnet,mc2046.codfw.wmnet,mc-gp2005.codfw.wmnet,ml-serve2009.codfw.wmnet,ml-staging2003.codfw.wmnet,ms-be2085.codfw.wmnet,pc2017.codfw.wmnet,sretest2001.codfw.wmnet,wikikube-ctrl2001.codfw.wmnet,wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet

Icinga downtime and Alertmanager silence (ID=e7152add-70e3-4a71-9065-f1276f12abdb) set by cmooney@cumin1003 for 1:00:00 on 5 host(s) and their services with reason: lsw1-b7-codfw JunOS upgrade

lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt

Mentioned in SAL (#wikimedia-operations) [2026-07-21T14:23:31Z] <topranks> reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) T430928

Icinga downtime and Alertmanager silence (ID=988f9bff-f6d2-4af1-9bf4-182d9352f225) set by cmooney@cumin1003 for 1:00:00 on 2 host(s) and their services with reason: lsw1-b7-codfw JunOS upgrade

ssw1-a[1,8]-codfw

Switch is back up after upgrade and things look good at first glance.

All seems good, hosts have been repooled.

Icinga downtime and Alertmanager silence (ID=7f1c1ea0-38e4-46b3-89a9-b606ab8ac69b) set by cmooney@cumin1003 for 1:00:00 on 30 host(s) and their services with reason: lsw1-b8-codfw JunOS upgrade

apus-be2004.codfw.wmnet,apus-fe2003.codfw.wmnet,db[2163-2164,2185,2189,2249].codfw.wmnet,dns2004.wikimedia.org,es2052.codfw.wmnet,ganeti2050.codfw.wmnet,ganeti-jumbo2001.codfw.wmnet,gitlab-runner2002.codfw.wmnet,kafka-main2007.codfw.wmnet,kubestage2002.codfw.wmnet,restbase[2029-2030].codfw.wmnet,thanos-be2007.codfw.wmnet,wdqs2007.codfw.wmnet,wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet

Icinga downtime and Alertmanager silence (ID=e840132a-4d0f-4713-b838-f9316d37ac55) set by cmooney@cumin1003 for 1:00:00 on 5 host(s) and their services with reason: lsw1-b8-codfw JunOS upgrade

lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw