Page MenuHomePhabricator

Machine Learning: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets
Closed, ResolvedPublic

Description

eqiad rows C and D have been migrated to new Nokia switches, and more importantly to the new network design.

You can find all the information on https://wikitech.wikimedia.org/wiki/Vlan_migration

All new servers are by default getting in those new vlans, but to not have to wait for a full 5+ years server refresh cycle, we're now asking service owners to re-image their existing baremetal servers using the --move-vlan cookbook parameter, at their own pace/convenience. This will change the server's IP.
There will of course be some special cases (like Ganeti, or DBs) and that's ok to not covert 100% of the servers, but the most we can get, the better.

Please contact netops for any help.

cumin1003:~$ sudo cumin 'A:owner-machine-learning and P{P:netbox::host%location ~ "[C|D].*eqiad"} and P{F:fqdn ~ ".wmnet$"} and not A:vms and not P{F:netmask = "255.255.255.0"}'
5 hosts will be targeted:
dse-k8s-worker[1003-1004].eqiad.wmnet,ml-cache1002.eqiad.wmnet,ml-serve[1003-1004].eqiad.wmnet
  • dse-k8s-worker[1003-1004].eqiad.wmnet
  • ml-cache1002.eqiad.wmnet
  • ml-serve[1003-1004].eqiad.wmnet

Event Timeline

Actually dse-k8s-worker were already taken care of in T421714: Data platform: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets so we should remove them from here, @bking do you confirm?

klausman subscribed.

ml-cache1002 has been decom'd (cookbook running atm)

I'll take care of the ml-serve machines this week.

klausman changed the task status from Open to In Progress.Jul 9 2026, 7:58 AM
klausman claimed this task.

Roll-reimage of nodes in ml-serve-eqiad cluster started by klausman:

  • ml-serve1003.eqiad.wmnet

Cookbook cookbooks.sre.k8s.roll-reimage-nodes -- rolling reimage on P{ml-serve1003.eqiad.wmnet} and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) completed:

  • ml-serve1003 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607090839_klausman_3087583_ml-serve1003.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Host ml-serve1003.eqiad.wmnet pooled in ml-serve-eqiad

Cookbook cookbooks.sre.hosts.reimage was started by klausman@cumin1003 for host ml-serve1003.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by klausman@cumin1003 for host ml-serve1003.eqiad.wmnet with OS bookworm completed:

  • ml-serve1003 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607090933_klausman_3103779_ml-serve1003.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage was started by klausman@cumin1003 for host ml-serve1004.eqiad.wmnet with OS bookworm

old:
ml-serve1003.eqiad.wmnet has address 10.64.32.81
ml-serve1003.eqiad.wmnet has IPv6 address 2620:0:861:103:10:64:32:81
ml-serve1004.eqiad.wmnet has address 10.64.48.50
ml-serve1004.eqiad.wmnet has IPv6 address 2620:0:861:107:10:64:48:50

new:
ml-serve1003.eqiad.wmnet has address 10.64.141.8
ml-serve1003.eqiad.wmnet has IPv6 address 2620:0:861:113:10:64:141:8
ml-serve1004.eqiad.wmnet has address 10.64.189.7
ml-serve1004.eqiad.wmnet has IPv6 address 2620:0:861:144:10:64:189:7

Cookbook cookbooks.sre.hosts.reimage started by klausman@cumin1003 for host ml-serve1004.eqiad.wmnet with OS bookworm completed:

  • ml-serve1004 (PASS)
    • Downtimed on Icinga/Alertmanager
    • Disabled Puppet
    • Host successfully migrated to the new VLAN
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202607091018_klausman_3142547_ml-serve1004.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

All ML worder nodes done.
ML cache is already half-decommed and will either be full decom'd or reused.
DSE hosts were already done as discussion above (confirmed with BKing).