Page MenuHomePhabricator

lsw1-a8-codfw: fpc0 PFE Statistics received unknown trigger (type Semaphore, id 0)
Closed, ResolvedPublic

Description

Noticed that the logs on that host are flooded with this one message:
fpc0 PFE Statistics received unknown trigger (type Semaphore, id 0)
at about 300k logs/h.

It's less visible on the cli as they're aggregated with:

Jul  2 11:30:22  lsw1-a8-codfw fpc0 PFE Statistics received unknown trigger (type Semaphore, id 0)
Jul  2 11:30:25  lsw1-a8-codfw last message repeated 273 times
Jul  2 11:31:22  lsw1-a8-codfw last message repeated 4692 times

But much more visible on the Opensearch collector.

No clear pointers on a quick web search. Either we upgrade, or open a JTAC, or reboot.

Event Timeline

Restricted Application added a subscriber: Aklapper. · View Herald Transcript

the upside is that there are only 3 hosts on that switch at the moment:

  • db2146
  • wikikube-worker2046
  • wikikube-worker2042

So better to upgrade sooner than latter.

My vote would be to try a reboot first. We've 49 EVPN switches running 22.2R3.15, and we only have this issue on one of them.

Having a mixture of versions in a cluster is likely not going to be an issue, but it's best practice to keep things standard. So I figure if we upgrade this one we're kicking off the project to upgrade them all. We wouldn't have to complete that at an insane pace, but given our current workload it might be best to leave it for now.

If a reboot doesn't work we can review again?

Sounds good!
@jijiki @Ladsgroup @Marostegui
Would it be possible to sync up to depool those 3 hosts for a switch reboot?

db2146 is a normal s1 replica. I can depool the db at any time (you can depool it yourself too if you want to). Let me know when do you want to do the maint work.

Sweet, what about 12:00UTC on Monday 7th ?

sounds good. I try to be around but if I couldn't for any reason, depool it yourself (https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting#Depooling_a_replica)

Mentioned in SAL (#wikimedia-operations) [2025-07-07T11:54:57Z] <ladsgroup@cumin1002> dbctl commit (dc=all): 'Depool db2146 T398433', diff saved to https://phabricator.wikimedia.org/P78771 and previous config saved to /var/cache/conftool/dbconfig/20250707-115457-ladsgroup.json

Mentioned in SAL (#wikimedia-operations) [2025-07-07T11:59:02Z] <ladsgroup@cumin1002> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2146.codfw.wmnet with reason: Just in case (T398433)

Cookbook cookbooks.sre.k8s.pool-depool-node started by akosiaris@cumin1003 depool for host wikikube-worker2042.codfw.wmnet completed:

  • wikikube-worker2042.codfw.wmnet (PASS)
    • Host wikikube-worker2042.codfw.wmnet depooled from wikikube-codfw

Cookbook cookbooks.sre.k8s.pool-depool-node started by akosiaris@cumin1003 depool for host wikikube-worker2046.codfw.wmnet completed:

  • wikikube-worker2046.codfw.wmnet (PASS)
    • Host wikikube-worker2046.codfw.wmnet depooled from wikikube-codfw

Sweet, what about 12:00UTC on Monday 7th ?

wikikube-worker204[26] have been drained and depooled.

ayounsi claimed this task.

Screenshot From 2025-07-07 14-24-43.png (955×280 px, 37 KB)

Much better.

Thanks for the depool, you can repool the hosts.

Cookbook cookbooks.sre.k8s.pool-depool-node started by akosiaris@cumin1003 pool for host wikikube-worker2046.codfw.wmnet completed:

  • wikikube-worker2046.codfw.wmnet (PASS)
    • Host wikikube-worker2046.codfw.wmnet pooled in wikikube-codfw

Cookbook cookbooks.sre.k8s.pool-depool-node started by akosiaris@cumin1003 pool for host wikikube-worker2042.codfw.wmnet completed:

  • wikikube-worker2042.codfw.wmnet (PASS)
    • Host wikikube-worker2042.codfw.wmnet pooled in wikikube-codfw