Page MenuHomePhabricator

Q#:rack/setup/install (2) cloudbackup hosts
Closed, ResolvedPublic

Description

This task will track the racking, setup, and OS installation of two new hosts to replace cloudbackup200[1,2]

Hostname / Racking / Installation Details

Hostnames: cloudbackup2003.codfw.wmnet and cloudbackup2004.codfw.wmnet
Racking Proposal: These are not part of the codfw1dev cluster, so they should not be placed in one of the existing codfw1dev racks (which have switches set up specifically for that.) They're not an HA pair, so it's fine if they're in the same rack elsewhere.
Networking Setup: 1 10G connection each. Private vlan (.codfw.wmnet), yes to AAAA records.
Partitioning/Raid: Hardware raid 1 for the two smaller OS drives, Hardware raid6 collecting all the remaining drives into one volume.
OS Distro: Bookworm
Sub-team Technical Contact: @Andrew

Per host setup checklist

Each host should have its own setup checklist copied and pasted into the list below.

cloudbackup2003:
  • Receive in system on procurement task T356088 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook
cloudbackup2004:
  • Receive in system on procurement task T356088 & in Coupa
  • Rack system with proposed racking plan (see above) & update Netbox (include all system info plus location, state of planned)
  • Run the Provision a server's network attributes Netbox script - Note that you must run the DNS and Provision cookbook after completing this step
  • Immediately run the sre.dns.netbox cookbook
  • Immediately run the sre.hosts.provision cookbook
  • Run the sre.hardware.upgrade-firmware cookbook
  • Update the operations/puppet repo - this should include updates to preseed.yaml, and site.pp with roles defined by service group: https://wikitech.wikimedia.org/wiki/SRE/Dc-operations
  • Run the sre.hosts.reimage cookbook

Event Timeline

RobH renamed this task from Q#:rack/setup/install X to Q#:rack/setup/install (2) cloudbackup hosts.Jan 30 2024, 8:33 PM
RobH assigned this task to Andrew.
RobH moved this task from Backlog to Racking Tasks on the ops-codfw board.

@Andrew: I've assigned this task to you for you to populate the racking details, additionally please add the servers to the site.pp file with the insetup role and their partition info.

Once done, you can unassign this from yourself, it is already in the proper queue and column for when the hosts arrive.

RobH mentioned this in Unknown Object (Task).Jan 30 2024, 8:34 PM
RobH unsubscribed.

@Andrew we've received these servers. Could you update this ticket with racking requirements and names of the servers?

Andrew updated the task description. (Show Details)

Sorry for the slow response! I hope I've now included all that you need.

@Andrew thanks for the update. Can I bug you to update the site.pp as well? Thanks!

Change #1013559 had a related patch set uploaded (by Andrew Bogott; author: Andrew Bogott):

[operations/puppet@production] site.pp: add insetup entries for new cloudbackup200[34] hosts

https://gerrit.wikimedia.org/r/1013559

Change #1013559 merged by Andrew Bogott:

[operations/puppet@production] site.pp: add insetup entries for new cloudbackup200[34] hosts

https://gerrit.wikimedia.org/r/1013559

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2004.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host cloudbackup2004.codfw.wmnet with OS bookworm completed:

  • cloudbackup2004 (PASS)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Downtimed the new host on Icinga/Alertmanager
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202403221814_jhancock_3473792_cloudbackup2004.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is optimal
    • Icinga downtime removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully

cloudbackup2003 has os after a few attempts. had to delete and redo the virtual disks twice before it took. but now it is having issues with the puppet server. will try again later.

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm executed with errors:

  • cloudbackup2003 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Generated Puppet certificate
    • The reimage failed, see the cookbook logs for the details,You can also try typing "install-console" cloudbackup2003.codfw.wmnet to get a root shellbut depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm executed with errors:

  • cloudbackup2003 (FAIL)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • The reimage failed, see the cookbook logs for the details,You can also try typing "install-console" cloudbackup2003.codfw.wmnet to get a root shellbut depending on the failure this may not work.

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

@Jhancock.wm this is what 2003 is showing on console

┌───────────────────────┤ [!!] Partition disks ├────────────────────────┐
   │                                                                       │
   │                 Failed to partition the selected disk                 │
   │ This probably happened because the selected disk or free space is too │
   │ small to be automatically partitioned.                                │
   │                                                                       │
   │     <Go Back>

I suspect that on re-run partman isn't properly zeroing out the partition table before it starts so we're accumulating cruft. I'd try visiting the CLI and deleting everything manually before rerunning.

Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm

Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host cloudbackup2003.codfw.wmnet with OS bookworm completed:

  • cloudbackup2003 (WARN)
    • Removed from Puppet and PuppetDB if present and deleted any certificates
    • Removed from Debmonitor if present
    • Forced PXE for next reboot
    • Host rebooted via IPMI
    • Host up (Debian installer)
    • Add puppet_version metadata to Debian installer
    • Checked BIOS boot parameters are back to normal
    • Host up (new fresh bookworm OS)
    • Generated Puppet certificate
    • Signed new Puppet certificate
    • Run Puppet in NOOP mode to populate exported resources in PuppetDB
    • Found Nagios_host resource for this host in PuppetDB
    • Unable to downtime the new host on Icinga/Alertmanager, the sre.hosts.downtime cookbook returned 99
    • First Puppet run completed and logged in /var/log/spicerack/sre/hosts/reimage/202403281626_jhancock_1120412_cloudbackup2003.out
    • configmaster.wikimedia.org updated with the host new SSH public key for wmf-update-known-hosts-production
    • Rebooted
    • Automatic Puppet run was successful
    • Forced a re-check of all Icinga services for the host
    • Icinga status is not optimal, downtime not removed
    • Updated Netbox data from PuppetDB
    • Updated Netbox status planned -> active
    • The sre.puppet.sync-netbox-hiera cookbook was run successfully
    • Cleared switch DHCP cache and MAC table for the host IP and MAC (EVPN Switch)
Jhancock.wm updated the task description. (Show Details)

issue fixed and ready to go @Andrew

Andrew claimed this task.

Change #1015526 had a related patch set uploaded (by Andrew Bogott; author: Andrew Bogott):

[operations/puppet@production] cloudbackup: put new cloudbackup hosts into service with minimal test load

https://gerrit.wikimedia.org/r/1015526

Change #1015526 merged by Andrew Bogott:

[operations/puppet@production] cloudbackup: put new cloudbackup hosts into service with minimal test load

https://gerrit.wikimedia.org/r/1015526

Change #1015535 had a related patch set uploaded (by Andrew Bogott; author: Andrew Bogott):

[operations/puppet@production] backy2: apply fix-backy2-crypto-imports on Bookworm

https://gerrit.wikimedia.org/r/1015535

Change #1015535 merged by Andrew Bogott:

[operations/puppet@production] backy2: apply fix-backy2-crypto-imports on Bookworm

https://gerrit.wikimedia.org/r/1015535

These are now set up and should start running a few backup jobs over the weekend. I need to check back and make sure things are actually working.

Change #1017915 had a related patch set uploaded (by Andrew Bogott; author: Andrew Bogott):

[operations/puppet@production] cinder backups: move real backup workloads to 200[34]

https://gerrit.wikimedia.org/r/1017915

Change #1017916 had a related patch set uploaded (by Andrew Bogott; author: Andrew Bogott):

[operations/puppet@production] Prepare cloudbackup200[12] for decom

https://gerrit.wikimedia.org/r/1017916

Change #1017915 merged by Andrew Bogott:

[operations/puppet@production] cinder backups: move real backup workloads to 200[34]

https://gerrit.wikimedia.org/r/1017915

Change #1017916 merged by Andrew Bogott:

[operations/puppet@production] Prepare cloudbackup200[12] for decom

https://gerrit.wikimedia.org/r/1017916

@Jhancock.wm anything else left to be done on this task?

These are now in service and working fine.