Page MenuHomePhabricator

2 VMs for mw-experimental
Closed, ResolvedPublic

Description

Site/Location:eqiad, codfw
Number of systems: 2
Service: mw-experimental
Networking Requirements: internal IP
Processor Requirements: 4
Memory: 16G
Disks: 200G
Wikitech: https://wikitech.wikimedia.org/wiki/Mw-experimental (TBA)
ServiceOps would like to request 2 VMs (1 in each DC), which will be used as kubernetes workers, dedicated to run mw-experimental pods.

  • wikikube-worker-exp1001 (row B)
  • wikikube-worker-exp2001 (row C)

Event Timeline

Looks fine, please use codfw/row C and eqiad eqiad/row B

Change #1159502 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] site.pp: add wikikube-worker-exp(1001|2001)

https://gerrit.wikimedia.org/r/1159502

Change #1159518 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] site.pp: make wikikube-worker-exp* k8s workers

https://gerrit.wikimedia.org/r/1159518

Change #1159502 merged by Effie Mouzeli:

[operations/puppet@production] site.pp: add wikikube-worker-exp(1001|2001)

https://gerrit.wikimedia.org/r/1159502

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1002 for host wikikube-worker-exp1001.eqiad.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1002 for host wikikube-worker-exp1001.eqiad.wmnet with OS bookworm completed:

  • wikikube-worker-exp1001 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via gnt-instance
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Set boot media to disk
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202506171306_jiji_3425476_wikikube-worker-exp1001.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Mentioned in SAL (#wikimedia-operations) [2025-06-17T13:54:22Z] <jiji@cumin1002> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Adding wikikube-worker-exp1001 - jiji@cumin1002 - T397051"

Mentioned in SAL (#wikimedia-operations) [2025-06-17T13:54:43Z] <jiji@cumin1002> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Adding wikikube-worker-exp1001 - jiji@cumin1002 - T397051"

Change #1159518 merged by Effie Mouzeli:

[operations/puppet@production] site.pp: make wikikube-worker-exp1001 a k8s worker

https://gerrit.wikimedia.org/r/1159518

Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1002 for host wikikube-worker-exp2001.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1002 for host wikikube-worker-exp2001.codfw.wmnet with OS bookworm completed:

  • wikikube-worker-exp2001 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via gnt-instance
    • Host up (Debian installer)
    • Add puppet_version metadata (7) to Debian installer
    • Set boot media to disk
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202506171726_jiji_3563660_wikikube-worker-exp2001.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Mentioned in SAL (#wikimedia-operations) [2025-06-17T19:26:26Z] <jiji@cumin1002> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Adding wikikube-worker-exp2001 - jiji@cumin1002 - T397051"

Mentioned in SAL (#wikimedia-operations) [2025-06-17T19:26:34Z] <jiji@cumin1002> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Adding wikikube-worker-exp2001 - jiji@cumin1002 - T397051"

Change #1160238 had a related patch set uploaded (by Effie Mouzeli; author: Effie Mouzeli):

[operations/puppet@production] site.pp: make wikikube-worker-exp2001 a k8s worker

https://gerrit.wikimedia.org/r/1160238

jijiki triaged this task as Medium priority.
jijiki updated the task description. (Show Details)

Change #1160238 merged by Effie Mouzeli:

[operations/puppet@production] site.pp: make wikikube-worker-exp2001 a k8s worker

https://gerrit.wikimedia.org/r/1160238