Page MenuHomePhabricator

Q3:rack/setup/install pc202[1-4]
Closed, ResolvedPublic

Description

This task will track the racking, setup, and OS installation of pc202[1-4]

Hostname / Racking / Installation Details

Hostnames: pc2021 pc2022 pc2023 pc2024
Racking Proposal: One per row
Networking Setup: # of Connections:1 Speed:1G/10G (we can do both, prefer 10G, but we currently use 1G). - VLAN:Private
OS Distro: Debian Trixie
Boot Method: UEFI
Sub-team Technical Contact: @Marostegui

Per host setup checklist

Each host should have its own setup checklist copied and pasted into the list below.

pc2021
  • Receive in system on procurement task T417069 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook
pc2022
  • Receive in system on procurement task T417069 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook
pc2023
  • Receive in system on procurement task T417069 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook
pc2024
  • Receive in system on procurement task T417069 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook

Event Timeline

RobH mentioned this in Unknown Object (Task).
RobH added a parent task: Unknown Object (Task).
RobH moved this task from Backlog to Racking Tasks on the ops-codfw board.
RobH unsubscribed.

Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations) and add the new servers to preseed.yml for partition info.

If possible, please reference this task number in your patch set, so it is clear when complete. Once complete, just un-assign yourself (leaving no assignee) for this task and once the hardware arrives on-site engineerss will claim this task for racking and setup. Please don't re-subscribe me to this task unless there is a direct question for me.

Thank you!

Change #1247868 had a related patch set uploaded (by Marostegui; author: Marostegui):

[operations/puppet@production] installserver: Add pc2021-pc2024

https://gerrit.wikimedia.org/r/1247868

Change #1247868 merged by Marostegui:

[operations/puppet@production] installserver: Add pc2021-pc2024

https://gerrit.wikimedia.org/r/1247868

Change #1247869 had a related patch set uploaded (by Marostegui; author: Marostegui):

[operations/puppet@production] site.pp: Add pc202[1-4]

https://gerrit.wikimedia.org/r/1247869

Change #1247869 merged by Marostegui:

[operations/puppet@production] site.pp: Add pc202[1-4]

https://gerrit.wikimedia.org/r/1247869

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2021.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2022.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2023.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2024.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2021.codfw.wmnet with OS trixie completed:

  • pc2021 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604172139_jhancock_2138162_pc2021.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2023.codfw.wmnet with OS trixie completed:

  • pc2023 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604172142_jhancock_2138679_pc2023.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2022.codfw.wmnet with OS trixie completed:

  • pc2022 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604172147_jhancock_2138345_pc2022.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2024.codfw.wmnet with OS trixie completed:

  • pc2024 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604172159_jhancock_2138794_pc2024.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully
Jhancock.wm updated the task description. (Show Details)

@Marostegui these are complete!

@Jhancock.wm unfortunately pc2022, pc2023 and pc2024 have the wrong RAID. They should have RAID10 but they have RAID 0
pc2021 is correct.

See pc2021 config:

root@pc2021:~# perccli64 /c0 show all | grep -E "Size|PD|RAID10|RAID0"
PD Firmware Download in progress = No
Support PD Firmware Download = Yes
Support SystemPD = No
Supported PD Operations :
NVRAM Size = 128KB
Flash Size = 16MB
On Board Memory Size = 8192MB
CacheVault Flash Size = 21.650 GB
Current Size of CacheCade (GB) = 0
Current Size of FW Cache (MB) = 6678
ECC Bucket Size = 255
Maintain PD Fail History = Off
Strip Size = 256 KB
Direct PD Mapping = No
Maintain PD Fail History = No
RAID Level Supported = RAID0, RAID1, RAID5, RAID6, RAID10(2 or more drives per span), RAID50, RAID60
Enable SystemPD = No
Max Data Transfer Size = 2048 sectors
Max Configurable CacheCade Size(GB) = 0
Min Strip Size = 64 KB
Max Strip Size = 1.000 MB
DG Arr Row EID:Slot DID Type   State BT     Size PDC  PI SED DS3  FSpace TR
 0 -   -   -        -   RAID10 Optl  N  8.729 TB dflt N  N   dflt N      N
PDC=PD Cache|PI=Protection Info|SED=Self Encrypting Drive|Frgn=Foreign
DG/VD TYPE   State Access Consist Cache Cac sCC     Size Name
0/239 RAID10 Optl  RW     Yes     RWBD  -   OFF 8.729 TB
PD LIST :
EID:Slt DID State DG     Size Intf Med SED PI SeSz Model                      Sp Type
SeSz=Sector Size|Sp=Spun|U=Up|D=Down|T=Transition|F=Foreign
EID State Slots PD PS Fans TSs Alms SIM Port# ProdID VendorSpecific
EID=Enclosure Device ID | PD=Physical drive count | PS=Power Supply count

And pc2022 and the other two:

root@pc2022:~# perccli64 /c0 show all | grep -E "Size|PD|RAID"
PD Firmware Download in progress = No
Support PD Firmware Download = Yes
Current Personality = RAID-Mode
Support Odd & Even Drive count in RAID1E = No
Support SystemPD = No
Supported PD Operations :
NVRAM Size = 128KB
Flash Size = 16MB
On Board Memory Size = 8192MB
CacheVault Flash Size = 21.650 GB
Current Size of CacheCade (GB) = 0
Current Size of FW Cache (MB) = 6678
ECC Bucket Size = 255
Maintain PD Fail History = Off
Strip Size = 256 KB
Direct PD Mapping = No
Maintain PD Fail History = No
BreakMirror RAID Support = No
RAID Level Supported = RAID0, RAID1, RAID5, RAID6, RAID10(2 or more drives per span), RAID50, RAID60
Enable SystemPD = No
Max Data Transfer Size = 2048 sectors
Max Configurable CacheCade Size(GB) = 0
Read cache bypass enabled for Parity RAID LDs = Yes
Min Strip Size = 64 KB
Max Strip Size = 1.000 MB
DG Arr Row EID:Slot DID Type  State BT      Size PDC  PI SED DS3  FSpace TR
 0 -   -   -        -   RAID0 Optl  N  17.458 TB dflt N  N   dflt N      N
 0 0   -   -        -   RAID0 Optl  N  17.458 TB dflt N  N   dflt N      N
PDC=PD Cache|PI=Protection Info|SED=Self Encrypting Drive|Frgn=Foreign
DG/VD TYPE  State Access Consist Cache Cac sCC      Size Name
0/239 RAID0 Optl  RW     Yes     RWBD  -   OFF 17.458 TB
PD LIST :
EID:Slt DID State DG     Size Intf Med SED PI SeSz Model                      Sp Type
SeSz=Sector Size|Sp=Spun|U=Up|D=Down|T=Transition|F=Foreign
EID State Slots PD PS Fans TSs Alms SIM Port# ProdID VendorSpecific
EID=Enclosure Device ID | PD=Physical drive count | PS=Power Supply count

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2022.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2023.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host pc2024.codfw.wmnet with OS trixie

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2022.codfw.wmnet with OS trixie completed:

  • pc2022 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604202314_jhancock_774094_pc2022.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2023.codfw.wmnet with OS trixie completed:

  • pc2023 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604202319_jhancock_774293_pc2023.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host pc2024.codfw.wmnet with OS trixie completed:

  • pc2024 (WARN)
    • Downtimed on Icinga/Alertmanager
    • Unable to disable Puppet, the host may have been unreachable
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced UEFI HTTP Boot for next reboot
    • Host rebooted via Redfish
    • Host up (Debian installer)
    • Host up (new fresh trixie OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • Removed previous downtime on Alertmanager (old OS)
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202604202324_jhancock_774433_pc2024.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB

@Marostegui fixed it. but please reopen the ticket if anything seems off.