Page MenuHomePhabricator

beta-scap-sync-world failing due to ssh failure for mwdeploy@deployment-mwmaint03.deployment-prep.eqiad1.wikimedia.cloud
Closed, ResolvedPublicBUG REPORT

Description

Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: Accepted key RSA SHA256:nUV3qf86EbG/cslV8H2DkV2upw8CGoIgqYYH2UPK7QE found at /etc/ssh/userkeys/mwdeploy:1
Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: Postponed publickey for mwdeploy from 172.16.1.63 port 56670 ssh2 [preauth]
Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: Accepted key RSA SHA256:nUV3qf86EbG/cslV8H2DkV2upw8CGoIgqYYH2UPK7QE found at /etc/ssh/userkeys/mwdeploy:1
Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: pam_sss(sshd:account): Access denied for user mwdeploy: 4 (System error)
Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: Failed publickey for mwdeploy from 172.16.1.63 port 56670 ssh2: RSA SHA256:nUV3qf86EbG/cslV8H2DkV2upw8CGoIgqYYH2UPK7QE
Oct  2 19:47:32 deployment-mwmaint03 sshd[1156848]: fatal: Access denied for user mwdeploy by PAM account configuration [preauth]

Event Timeline

Mentioned in SAL (#wikimedia-releng) [2024-10-02T20:01:35Z] <bd808> Rebooting deployment-mwmaint03.deployment-prep.eqiad1.wikimedia.cloud (T376336)

bd808 claimed this task.
bd808 added a subscriber: hashar.

The reboot fixed the problem:

[20:08]  <wmf-insecte> Yippee, build fixed!
[20:08]  <wmf-insecte> Project beta-scap-sync-world build #174652: FIXED in 3 min 5 sec: https://integration.wikimedia.org/ci/job/beta-scap-sync-world/174652/

My theory here is that sssd got jammed up somehow, likely as a result of a network blip, and that broke the ssh key validation process. My network blip theory is because of @hashar reporting new transient failures at T374830: Various CI jobs running in the integration Cloud VPS project failing due to transient DNS lookup failures, often for our own hosts such as gerrit.wikimedia.org around the same time that the sync job started failing.