Page MenuHomePhabricator

Figure out what to do with an-worker1189
Closed, ResolvedPublic

Description

Per this Slack conversation with @Scott_French , an-worker1189, which is currently associated with Puppet role analytics_cluster::hadoop::worker is not in Puppet, but seems to be creating TCP connections to confd ports. I'm not familiar with the hadoop workers, but @RKemper might remember that this host had RAID issues in the past.

Creating this ticket to:

  • Check the current status of an-worker1189. Is it currently part of the Hadoop cluster? Answer: yes
  • Decide as a team what do to with this host (retire, reimage, repurpose?)

Event Timeline

bking changed the task status from Open to In Progress.Wed, Aug 5, 7:41 PM
bking claimed this task.
bking triaged this task as Medium priority.
bking updated the task description. (Show Details)

I checked an-worker1189's status using this Wikitech guide , the Hadoop web UI shows it is an active namenode (see attached)

Screenshot 2026-08-05 at 19.42.53.png (2,362×1,304 px, 377 KB)
.

So it's likely we'll just need to re-enroll this host into Puppet. I'll give that a shot now.

Mentioned in SAL (#wikimedia-operations) [2026-08-05T19:51:17Z] <inflatador> [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet T434142

Yes, it looks like Puppet was disabled for a server reboot and then never got re-enabled:

The last Puppet run was at Thu Jul 16 19:04:21 UTC 2026 (28840 minutes ago). Puppet is disabled. Reboot.

I manually re-enrolled the host in Puppet. So to answer the original question, an-worker1189 is an active production server that fell out of Puppet. That's been fixed, so I'm closing out this ticket.

If you are interested in the work to get this host onto Bookworm or later, see T406746 .