User Details
- User Since
- May 30 2017, 5:25 PM (480 w, 4 d)
- Availability
- Busy Busy until Sep 2.
- IRC Nick
- herron
- LDAP User
- Herron
- MediaWiki User
- Unknown
May 21 2026
May 18 2026
I don't think so
May 13 2026
May 12 2026
May 8 2026
I've spent some time experimenting with https://github.com/mahendrapaipuri/grafana-dashboard-reporter-app which looks promising for our use case. This is a Grafana plugin that creates an api endpoint that rendres a dashboard as a pdf report based upon the query string presented to the api. In my view this can do the heavy lifting of generating a visual report, and allows us to do so without straying very far from our existing tooling. We can report on existing dashboards, or create new streamlined report dashboards to be captured on the reporting interval
May 6 2026
All kafka-logging hosts have been upgraded to trixie
May 5 2026
May 4 2026
May 1 2026
Thanks for looking @cmooney @elukey! I think we can certainly rule out a network issue. Having a look this morning with fresh eyes, 2005 does appear to be working fine, but very lightly loaded. As in no load at all. I expected 2005 to resync and pull data after reimage, but the host is only responsible for topics/partitions that are presently 0B in size and so, although working, the /srv filesystem is suspiciously empty (2.1M used).
Apr 30 2026
Today I reimaged kafka-logging2005 with --move-vlan and afterwards the node is having trouble rejoining the cluster. I'm seeing errors like this like this logged by other cluster nodes when it attempts to bring the replica up to date:
Apr 28 2026
All kafka-logging brokers have been upgraded to 3.7
Apr 27 2026
Apr 23 2026
Cookbook worked well!
Apr 20 2026
Apr 17 2026
@brouberol, do the steps in the description look ok to you, anything I missed?
Apr 15 2026
Apr 14 2026
Apr 13 2026
setting low priority as we've just upgrade kafkamon hosts to trixie, so we'll have a Debian upgrade cycle worth of time to perform the move
Apr 10 2026
Low priority question, but since things may have changed since the setup of the current approach I'll ask -- Is smart-data-dump still the better approach for this type of monitoring compared to something along the lines of https://github.com/prometheus-community/smartctl_exporter?
Apr 7 2026
At a high level the alerting system is coupled to prometheus metrics. For the purposes of generating metrics from log events we support what's called the promethues-es-exporter which executes queries against logs (opensearch) and transforms the results into metrics which can then be used to generate alerts.
Apr 4 2026
Apr 3 2026
Apr 1 2026
Mar 30 2026
Mar 25 2026
That would be nice! Off hand there's a name convention on the Kafka side as well where clusters are named like logging-eqiad, main-eqiad. I'm not sure if it'd be a problem to change the assumed prefix for the hostname while keeping the same cluster name inside Kafka. Good question for the Kafka SIG!
It could be a different thing. AFAIK sre.hosts.reboot-single sets downtime itself as well, maybe an edge case that happens when downtimes are doubled?
Mar 19 2026
The sloth onboarding backlog is empty!
We chatted about this a bit at the o11y team meeting this week and consensus was that we're looking ok capacity wise, but would like to explore the potential of an SSD backed bucket. I'll give @tappof an opportunity to chime in as well. In the mean time @MatthewVernon what's your initial impression of the SSD idea?