Page MenuHomePhabricator

Requesting access to Jenkins server for PWangai-WMF
Closed, ResolvedPublicRequest

Description

Requestor provided information and prerequisites

Complete ALL items below as the individual person who is requesting access:

  • Wikimedia developer account username: pwangai
  • Email address: pwangai@wikimedia.org
  • SSH public key (must be a separate key from Wikimedia cloud SSH access): ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAINRw4zb+BA6iB6CO3PvbJbKLQDzPKEUfEP55jxr7u+/n
  • Requested group membership:
  • Reason for access: I need access the Jenkins server (contint1003.wikimedia.org) to be able to collect castor metrics
  • Name of approving party (manager for WMF/WMDE staff): @AMarkossyan-WMF
  • Ensure you have signed the L3 Wikimedia Server Access Responsibilities document: signed on Wed, Jul 29, 10:39 PM
  • Please coordinate obtaining a comment of approval on this task from the approving party.

SRE Clinic Duty Confirmation Checklist for Access Requests

This checklist should be used on all access requests to ensure that all steps are covered, including expansion to existing access. Please double check the step has been completed before checking it off.

This section is to be confirmed and completed by a member of the SRE team.

  • - User has signed the L3 Acknowledgement of Wikimedia Server Access Responsibilities Document.
  • - User has a valid NDA on file with WMF legal. (All WMF Staff/Contractor hiring are covered by NDA. Other users can be validated via the NDA tracking sheet)
  • - User has provided the following: developer account username, email address, and full reasoning for access (including what commands and/or tasks they expect to perform)
  • - User has provided a public SSH key. This ssh key pair should only be used for WMF cluster access, and not shared with any other service (this includes not sharing with WMCS access, no shared keys.)
  • - The provided SSH key has been confirmed out of band and is verified not being used in WMCS.
  • - access request (or expansion) has sign off of WMF sponsor/manager (sponsor for volunteers, manager for wmf staff)
  • - access request (or expansion) has sign off of group approver indicated by the approval field in data.yaml

For additional details regarding access request requirements, please see https://wikitech.wikimedia.org/wiki/Requesting_shell_access

Event Timeline

Hi there! This is not a super common type of access request, so let me start off with some general comments.

Jenkins just recently moved to this new dedicated server role for just jenkins itself, separate from the rest of CI servers.

Assuming it's really about this in the stricter sense, "where the jenkins service runs", then this is a also a relatively new puppet role.

I am saying this because we give access based on roles, rather than individual server names (and groups rather than individual user names directly).

One option here is for you to become member of one of the existing release engineering team groups that already have access to the jenkins server role. But that also comes with other access you do not need necessarily.

So the other option is that we create a new group for this type of access.

Would you say this is more like a temporary one-off or more like an ongoing need for a specific type of access. Is it likely other people would also need this in the future? Ideally would there be more than 1 person with this access?

Let me tag release engineering team to discuss.

Hi releng team,

so the actual request here is "collect castor metrics". Given our recent split of the CI server roles let's confirm it really means the dedicated jenkins servers and not the "regular CI servers".

Furthermore, need your input if you generally approve this type of access and if so.. whether it makes the most sense we create a new group for this type of request or, alternatively, add the user into one of your existing team groups.

The status quo is:

Members of:

- contint-users
- contint-admins
- contint-roots
- contint-docker

have access to jenkins servers but ALSO to the main CI servers.

The difference between the groups is "how much sudo" they have. From full root to some commands to shell without elevated privileges.

So it depends what commands are actually need to fulfill the task of collecting castor metrics.

If no sudo is needed for it we should probably just add them to "contint-users" and be done with it.

Dzahn changed the task status from Open to In Progress.Thu, Jul 30, 8:53 PM

I would say this may not be a temporary one-off. Right now as a team we are formulating a KR revolving around castor improvements, and based on the current work, a need to access generated XML log files is necessary. Moving forward, we will keep observing castor performance, so continued access will be needed. Last quarter we were working on CI improvements and had several patches that needed deployment. We had help from Antoine from release engineering to do deployments.

As the test platform team, we will keep doing a lot of work on CI improvements, and this often requires deployments. After some team discussions with our manager, we concluded we need to learn how to do deployments ourselves in case we do not have someone from release engineering to lend us a hand. Access to CI is not necessary for now in my case, but at some point permissions may be requested depending on the projects our team undertakes in future. One of my team members (Peter Hedenskog) already has access due to his work on performance testing, and we have been depending on him to get access to this data. @AMarkossyan-WMF can give more context in case I left something out.

Noting here that I've added this to the RelEng team discussion agenda for Weds this week; we'll get back to you.

Thank you, Brennen!

On my end, I know that this requires my approval, so here you go: 👍

In general, as this is blocking us a little bit, we would ideally prefer to take the path of least resitance to be able to start collecting the data we need for now.

Fundamentally though Peter is right that our team is likely to continue working with CI and related topics a lot and go deeper over time, so I think that eventually having a dedicated group makes a lot of sense.

I strongly recommend to define the proper solution right now and avoid a diverting "path of least resistence". (Also there is no "just add this user quickly to this host name" fix here).

All experience shows that temporary solutions have a habit of becoming permanent and delaying it just means repeating this discussion later or never.

What can be done to speed this up drastically though is to:

  • define what shell command it is they need to be running

If this does not involve sudo we can add the user to "contint-users" and be done.

@Dzahn This will not involve sudo, I will be running a few scripts under my username to summarize the log data we need.

@pwangai Cool! in that case I think the simplest path forward is we just add people to "contint-users" which gives unprivileged shell access to:

  • contint servers ("main" & Jenkins servers)
  • doc machines (backends of doc.wikimedia.org)

I also recommend this to releng and they have this on their agenda for tomorrow. So we will get this sorted out soon. Cheers

Cool! in that case I think the simplest path forward is we just add people to "contint-users"

Discussed with team in RelEng weekly, consensus is this is fine, please go ahead.

Change #1321601 had a related patch set uploaded (by Dzahn; author: Dzahn):

[operations/puppet@production] admin: upgrade pwangai from ldap_only to contint-users shell

https://gerrit.wikimedia.org/r/1321601

Dzahn added a subscriber: dancy.

@dancy You are an approver for the contint-users group. Approved?

@dancy You are an approver for the contint-users group. Approved?

Approved.

Change #1321601 merged by Dzahn:

[operations/puppet@production] admin: upgrade pwangai from ldap_only to contint-users shell

https://gerrit.wikimedia.org/r/1321601

thanks all!

@pwangai Alright, here you go! you have access now. Your user exists on these hosts:

  • bast1004.wikimedia.org in the eqiad data center in Virginia, United States
  • bast2003.wikimedia.org in the codfw data center in Texas, United States
  • bast3007.wikimedia.org in the esams data center in Amsterdam, The Netherlands
  • bast4006.wikimedia.org in the ulsfo data center in San Francisco, United States
  • bast5005.wikimedia.org in the eqsin data center in Singapore
  • bast6003.wikimedia.org in the drmrs data center in Marseille, France
  • bast7002.wikimedia.org in the magru data center in São Paulo, Brazil

These are the bastion hosts you connect to from external to jump to the private hosts behind them.

See Wikitech:Bastion for more details. We recommend using the one closest to your physical location but you can use any of them.

Here are some docs on setting up your ssh config to jump via the bastions to other machines.

  • contint1002.wikimedia.org the main CI server (website, zuul-server, zuul-merger, proxy to jenkins, ..) in the eqiad data center
  • contint2002.wikimedia.org the main CI server (website, zuul-server, zuul-merger, proxy to jenkins, ..) in the codfw data center

Only one of them is the "active" or main one at a given time. Which one that is can be determined by looking up the host name contint.wikimedia.org.

contint.wikimedia.org is an alias for contint1002.wikimedia.org tells us contint1002 is currently the main server.

  • contint1003.wikimedia.org the dedicated Jenkins server in the eqiad data center
  • contint2003.wikimedia.org the dedicated Jenkins server in the eqiad data center

Here as well, one of them is the currently active server. jenkins.discovery.wmnet is an alias for contint1003.wikimedia.org tells us that contint1003 is currently the main jenkins server.

So putting these together I recommend:

  • pick any one of the bastion hosts and verify you can SSH to it directly
  • once that works, setup your SSH config to jump via this bastion host to contint1003.wikimedia.org
  • once you are on the shell of contint1003 you are where the current jenkins is running.

Let us know if this works for you. Cheers

Dzahn claimed this task.

claiming it's resolved - always ok to reopen if there are problems