Page MenuHomePhabricator

NodeBGPSessionStatusNotEstablished
Closed, ResolvedPublic

Description

Common information

  • alertname: NodeBGPSessionStatusNotEstablished
  • instance: wikikube-worker1036:0
  • job: gnmi
  • network_instance_name: DEFAULT
  • peer_as: 64601
  • peer_type: EXTERNAL
  • prometheus: ops
  • protocol_identifier: BGP
  • protocol_name: DEFAULT
  • remote_instance: cr2-eqiad:9804
  • severity: task
  • site: eqiad
  • source: prometheus
  • team: serviceops

Firing alerts



Details

Event Timeline

Restricted Application added a subscriber: Aklapper. · View Herald Transcript

So, wikikube-worker1036 underwent a VLAN move earlier today that would have switched it from peering with the CR to peering with its ToR switch (lsw1-c6-eqiad).

That operation seems to have wrapped up around 14:15 UTC per https://sal.toolforge.org/log/XyruPJ8B8tZ8Ohr0rdxp.

Now, presumably some time prior to that we should expect to see the new session successfully established, and indeed that's what we see:

swfrench@wikikube-worker1036:~$ sudo calicoctl node status
Calico process is running.

IPv4 BGP status
+--------------+---------------+-------+----------+-------------+
| PEER ADDRESS |   PEER TYPE   | STATE |  SINCE   |    INFO     |
+--------------+---------------+-------+----------+-------------+
| 10.64.173.1  | node specific | up    | 13:15:55 | Established |
+--------------+---------------+-------+----------+-------------+

IPv6 BGP status
+-------------------+---------------+-------+----------+-------------+
|   PEER ADDRESS    |   PEER TYPE   | STATE |  SINCE   |    INFO     |
+-------------------+---------------+-------+----------+-------------+
| 2620:0:861:133::1 | node specific | up    | 13:15:55 | Established |
+-------------------+---------------+-------+----------+-------------+

i.e.., according to Calico, it's had a BGP session established with its ToR switch since at least ~ 13:16 today.

Now, it took me a little while to wrap my mind around the fact that the remote_instance:gnmi_bgp_neighbor_session_state recording rule upon which the alert is based "inverts" the way I normally think of the instance label.

What I mean is, Calico's BIRD instance is of course not the gNMI target here - i.e., the underlying metrics are collected from the network devices - but this recording rule rewrites instance as the peer, while remote_instance is the network device from which the observation was made. Meaning, this alert is telling us the cr2-eqiad does not have an established BGP session with wikikube-worker1036, even though it's configured to.

I see nothing in the SAL about when the homer commit to drop the old session configuration from the CR happened as part of the VLAN move, and I don't have network device access to check the config history directly.

@Blake - Do you happen to recall when you ran homer? It looks like cr2-eqiad was attempting to reconnect until a bit before 16:00.

Now, what I don't understand is why it took until 14:59 for this to be opened, even though the session dropped around 12:35, and the alert trigger duration is 10m.


Edit: To clarify one point, the reason I'm asking about the timing of the homer commit is because if it _did_ happen significantly earlier, then maybe there's something wonky going on with the gNMI exporter when a session disappears (I don't think that's the case here, since we can pretty clearly see state-flapping, but I want to rule it out).

Scott_French moved this task from Inbox to Needs Info / Blocked on the ServiceOps board.

Ah, and now I just come across T430290: Silence NodeBGPSessionStatusNotEstablished during reimages when searching for prior instances of this. I'll follow up there on the question of what needs bracketed by the silence - i.e., NodeBGPSessionStatusNotEstablished should be silenced during a reimage, but in the specific case of a VLAN move, the silence should be there until the previous-peer devices have their configs updated.

I suspect I failed to run homer during the reimage of wikikube-worker1036. Looking at my shell history on cumin1003, it looks like I ran homer for wikikube-worker1037 (homer lsw1-d8-eqiad* commit 'T421711'), but I do not see a matching command for wikikube-worker1036 (and would expect homer lsw1-c6-eqiad* commit 'T421711'). It appears that I didn't !log this in either case, and will ensure to do that in the future. It also appears that I missed running homer for the cr in both cases, even though I can see it asking me to quite clearly in the cookbook logs:

2026-07-07 13:18:39,629 blake 2192371 [INFO] Please run the following homer commands
2026-07-07 13:18:39,630 blake 2192371 [INFO] ----------------------------------------------------
2026-07-07 13:18:39,630 blake 2192371 [INFO] # Long, can be run in parallel while finishing the cookbook
2026-07-07 13:18:39,630 blake 2192371 [INFO] !log homer cr*eqiad* commit 'T421711'
2026-07-07 13:18:39,631 blake 2192371 [INFO] homer cr*eqiad* commit 'T421711'
2026-07-07 13:18:39,631 blake 2192371 [INFO] # Before continuing
2026-07-07 13:18:39,631 blake 2192371 [INFO] !log homer lsw1-c6-eqiad* commit 'T421711'
2026-07-07 13:18:39,632 blake 2192371 [INFO] homer lsw1-c6-eqiad* commit 'T421711'
2026-07-07 13:18:39,632 blake 2192371 [INFO] ----------------------------------------------------

My apologies. I checked homer for diffs in eqiad this morning, and there appear not to be any, so I don't think there's anything left to fix at the moment. I'm going to try to run the procedure once more today, and ensure that the homer updates are run correctly.

As for why this alert fired when it did, I think it's just 10m after the downtime for wikikube-worker1036 expired:
2026-07-07 12:49:21,405 blake 2192371 [INFO] START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage

I think that 2h is coming out of sre.hosts.reimage here.

I'm not sure if that's what we want? It's possible that the renumber-node cookbook ought to set a longer downtime (or request a longer downtime when it calls out to sre.hosts.reimage here. We could set a 12 or 24h downtime in order to allow for operational friction around the reimage, but I'm not sure what's appropriate.

@Scott_French, any thoughts?

In pool-depool-node, the Icinga downtime appears to be 4h:

2026-07-07 12:23:04,457 blake 2192371 [INFO] Scheduling downtime for 4:00:00 on Icinga server alert1002.wikimedia.org for hosts: wikikube-worker1036
2026-07-07 12:23:04,458 blake 2192371 [INFO] Executing commands [cumin.transports.Command('bash -c \'echo -n "[1783426984] SCHEDULE_HOST_DOWNTIME;wikikube-worker1036;1783426984;1783441384;1;0;14400;blake@cumin1003;Host reimage" > /var/lib/icinga/rw/icinga.cmd \'')] on '1' hosts: alert1002.wikimedia.org

So does the alertmanager silence:

2026-07-07 12:23:17,615 blake 2192371 [INFO] Created silence ID 2bf414c4-6dd3-4d40-a676-c89a6344628a for 4:00:00

pool-depool-node also cleans up its downtime when it repools the node, which is why the alert fired when the sre.hosts.reimage silence expired, and not when this silence was removed.

It's a little confusing that we have overlapping silences here - pool-depool-node created a 4h silence that it expects to tidy up later, and sre.hosts.reimage creates a 2h silence that it does not tidy up, so we end up in a situation where we're either downtimed for 2h, or until the node is repooled, whichever happens later.

It seems like the problem is that, because running homer is currently external to the cookbook, and dependent on separate human action, that we'd need to make a guess when creating the original downtime with pool-depool-node, and hope that the human running the cookbook is done running homer before the silence expires.

My preference would be to have pool-depool-node create a longer silence which it should clean up, and (if possible) run homer inline with the renumber-node cookbook execution. That way, we have a silence that's bounded to the duration of the maintenance event, and don't rely on external parallel execution of another process.

Another way to make this slightly safer would be to have the cookbook check whether NodeBGPSessionStatusNotEstablished is firing for the node we're renumbering in the cookbook, and if it is, to not allow the cookbook user to progress beyond the prompt homer execution stage. I think that would give us a slightly stronger guarantee that homer has been run before proceeding to remove the silence. This feels more complicated to me, though, and it would be more convenient to just run homer inline.

Arzhel mentioned that it would be acceptable to make an inline call to run_homer, so I'll start work on that.

Change #1308595 had a related patch set uploaded (by Blake; author: Blake):

[operations/cookbooks@master] renumber-node: Add an option to run homer inline.

https://gerrit.wikimedia.org/r/1308595

I think the reason I skipped the later homer run for wikikube-worker1036 is because I saw a homer run earlier in the output, and assumed that that was the run the cookbook was asking me about. sre.hosts.reimage calls sre.hosts.move-vlan, which calls sre.network.configure-switch-interfaces, which actively runs homer and outputs to the cookbook log, but only for the TOR switch, and not for the core router.

Change #1308595 merged by jenkins-bot:

[operations/cookbooks@master] renumber-node: Add an option to run homer inline.

https://gerrit.wikimedia.org/r/1308595

I ran a renumber-node with the patch, and it looks like the 4h silence created by pool-depool-node is long enough to cover the reimage and the homer runs afterwards, so I don't think that duration necessarily needs to change.

Scott_French assigned this task to Blake.

Many thanks for the follow-ups, @Blake!

Ah, I didn't realize there was the 2h "hanging" silence, and indeed that would explain the timing.

In any case, I would agree that running homer inline (rather than prompting the operator to do it in the background, eventually) seems like a solid option, and I see you've landed that. Checking NodeBGPSessionStatusNotEstablished as an exit condition for the homer-prompt step is a nice fallback approach, but I guess that's now superseded by your patch.

(dons triage hat) Since you've already done the work here, I'm optimistically assigning to you and resolving. Feel free to reopen if you think there's additional work here to harden the router config cleanup story.