User Details
- User Since
- Jul 23 2024, 9:16 AM (106 w, 6 d)
- Availability
- Available
- IRC Nick
- tappof
- LDAP User
- Tiziano Fogli
- MediaWiki User
- Tiziano Fogli [ Global Accounts ]
Thu, Jul 23
The Debian package can be downloaded from the artifacts of the latest successfully executed pipeline (https://gitlab.wikimedia.org/repos/sre/grafana-image-renderer/-/jobs/905936).
Wed, Jul 15
Tue, Jul 14
Jul 3 2026
The alerts have been acknowledged (and then silenced) in Alertmanager: 38e40dbc-0d5b-4267-ad1f-338d2e651a23 (https://github.com/prometheus/alertmanager/issues/226).
These alerts are now fully based on Prometheus, although their definitions still reside in Puppet leveraging prometheus::alert::rule (T370153: Move kafka-mirror Prometheus-based alerts from Icinga to alerts.git).
They've been migrated using prometheus::alert::rule, as the migration to alerts.git turned out to be too complex to manage and maintain.
Jul 1 2026
Thanks!
Jun 29 2026
Taking a look at the IRC logs, I can see the notification from June 23:
#wikimedia-operations/2026-06-23.log:815:[14:49:12] <icinga-wm> PROBLEM - Host dbproxy1028 is DOWN: PING CRITICAL - Packet loss = 100%
Jun 25 2026
Jun 24 2026
Jun 16 2026
Yeah, you're correct. While taking a look, I didn't notice the "no" here (I only focused on "off" and "disabled"... my bad):
case "no", "off", "disabled": return 0, true
so it's exactly the scenario where the metric is only exported when something interesting is happening.
Jun 11 2026
I took a brief look at the code, and I don't think (although I may be wrong) that the mysql_slave_status_using_gtid metric is really reliable.
Jun 10 2026
Ok, the AlertLintProblem alert here is making you aware that you have configured an alert on some series that were generated by an exporter and are no longer present. In other words, it is telling you that your original rule (MySQLReplicaNotUsingGTID) may no longer be effective. The instantiated alert here is not MySQLReplicaNotUsingGTID but AlertLintProblem (related to the MySQLReplicaNotUsingGTID rule definition).
Yes. Resolved.
You can filter alerts of interest using team=data-persistence. The AlertLintProblem check groups together different linting issues, which are described in detail in the alert description itself.
Jun 9 2026
Jun 8 2026
Confirmed: it is related to Kafka (https://gerrit.wikimedia.org/r/q/repo:operations/deployment-charts+mergedbefore:2026-03-25+mergedafter:2026-03-22+kafka, @elukey ). There are many per-topic metrics that effectively double the number of scraped series:
Jun 5 2026
Jun 4 2026
May 29 2026
May 21 2026
May 20 2026
Thresholds have been updated, and notifications will now happen on IRC.
Thresholds have been updated, and notifications will now happen on IRC.
Thresholds have been updated, and notifications will now happen on IRC.
The Sloth dashboard is already using rows for SLO services. Nested rows, which would allow forecasting panels to be computed on demand, will be available starting with Grafana 13. In the meantime, I’ve removed the forecasting widget.
Waiting for feedback about performance.
Logs have been shipped to OpenSearch through a Logstash pipeline. A dashboard has been created. The documentation has been updated.
May 19 2026
Relying on Envoy seems to be the right approach to avoid metric duplication. I'm moving the task to our radar. Feel free to reach out if you need anything else.
May 14 2026
The proposed solutions seem fine to the o11y team. If you want to go ahead on your own, you're welcome to do so. Let us know if you need anything from us.
May 11 2026
I believe you can go ahead with the plan you shared. I checked the configuration, and the other components in place are the NSCA daemon itself (which is in charge of observability) and settings related to notifications that do not affect your migration.
Thank you.
May 6 2026
May 5 2026
Checking hosts... Error: 'asw2-ulsfo' is not a valid parent for host 'cp4038' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10713)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4040' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10747)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4042' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10781)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4044' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10815)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4046' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10849)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4048' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10883)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4050' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10917)! Error: 'asw2-ulsfo' is not a valid parent for host 'cp4052' (file '/etc/icinga/objects/puppet_hosts.cfg', line 10951)! Error: 'asw2-ulsfo' is not a valid parent for host 'dns4004' (file '/etc/icinga/objects/puppet_hosts.cfg', line 16630)! Error: 'asw2-ulsfo' is not a valid parent for host 'ganeti4006' (file '/etc/icinga/objects/puppet_hosts.cfg', line 20499)! Error: 'asw2-ulsfo' is not a valid parent for host 'ganeti4008' (file '/etc/icinga/objects/puppet_hosts.cfg', line 20533)! Error: 'asw2-ulsfo' is not a valid parent for host 'lvs4008' (file '/etc/icinga/objects/puppet_hosts.cfg', line 23611)! Error: 'asw2-ulsfo' is not a valid parent for host 'lvs4010' (file '/etc/icinga/objects/puppet_hosts.cfg', line 23645)! Error: 'asw2-ulsfo' is not a valid parent for host 'mr1-ulsfo' (file '/etc/nagios/nagios_host.cfg', line 978)!
Apr 30 2026
Apr 29 2026
Apr 28 2026
Apr 27 2026
root@titan2001:/srv/rewrite# thanos tools bucket mark --id=01KJFRKS6R61D6CADD4KGABTYR --marker=no-downsample-mark.json --details="downsampling this block would cause overlaps" --objstore.config-file=/etc/thanos-store@main/objstore.yaml --remove ts=2026-04-27T06:37:09.624961469Z caller=factory.go:54 level=info msg="loading bucket configuration" ts=2026-04-27T06:37:09.891558511Z caller=block.go:456 level=info msg="mark has been removed from the block" block=01KJFRKS6R61D6CADD4KGABTYR ts=2026-04-27T06:37:09.891591663Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-downsample-mark.json IDs=01KJFRKS6R61D6CADD4KGABTYR ts=2026-04-27T06:37:09.891620788Z caller=main.go:174 level=info msg=exiting
Apr 23 2026
root@titan2001:/srv/rewrite/tsdb/thanos/01KFFPFB9562T35FWHEZ6G10NG# thanos tools bucket mark --id=01KPRDEHBVNZQ8P99YD26E3ZWW --marker=deletion-mark.json --details="manual deletion" --objstore.config-file=/etc/thanos-store@main/objstore.yaml
ts=2026-04-23T10:06:14.366864436Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-04-23T10:06:14.797160137Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KPRDEHBVNZQ8P99YD26E3ZWW
ts=2026-04-23T10:06:14.797194266Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KPRDEHBVNZQ8P99YD26E3ZWW
ts=2026-04-23T10:06:14.7972193Z caller=main.go:174 level=info msg=exiting
Apr 21 2026
root@titan2001:/srv/rewrite# thanos tools bucket mark --id=01KPPJZ4WB223H5X3AMPAN1ZQM --marker=deletion-mark.json --details="manual deletion" --objstore.config-file=/etc/thanos-store@main/objstore.yaml ts=2026-04-21T10:03:55.637776792Z caller=factory.go:54 level=info msg="loading bucket configuration" ts=2026-04-21T10:03:56.532181856Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KPPJZ4WB223H5X3AMPAN1ZQM ts=2026-04-21T10:03:56.532217508Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KPPJZ4WB223H5X3AMPAN1ZQM ts=2026-04-21T10:03:56.532257098Z caller=main.go:174 level=info msg=exiting
Apr 20 2026
root@titan2001:/srv/rewrite# thanos tools bucket mark --id=01KPEAMZWYDACKP8STG0NY6M0W --marker=deletion-mark.json --details="manual deletion" --objstore.config-file=/etc/thanos-store@main/objstore.yaml ts=2026-04-20T08:08:37.087157771Z caller=factory.go:54 level=info msg="loading bucket configuration" ts=2026-04-20T08:08:37.556039027Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KPEAMZWYDACKP8STG0NY6M0W ts=2026-04-20T08:08:37.556077272Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KPEAMZWYDACKP8STG0NY6M0W ts=2026-04-20T08:08:37.556108937Z caller=main.go:174 level=info msg=exiting
Apr 15 2026
The multi-instance Thanos compactor has been deployed: Prometheus instances are assigned to compactor instances on the titan hosts via the prometheus::instances Hiera variable.
Apr 10 2026
Apr 9 2026
root@titan2001:/srv/rewrite# cat /tmp/tbd | awk '{print $2}' | xargs -I % thanos tools bucket mark --id=% --marker=deletion-mark.json --details="manual deletion" --objstore.config-file=/etc/thanos-store@main/objstore.yaml
ts=2026-04-09T07:31:33.471510733Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-04-09T07:31:33.842489347Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KNB5NY6PMVSNTAAZHX6VE5CQ
ts=2026-04-09T07:31:33.842528357Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KNB5NY6PMVSNTAAZHX6VE5CQ
ts=2026-04-09T07:31:33.842585231Z caller=main.go:174 level=info msg=exitingApr 8 2026
This has been done on purpose (see modules/profile/manifests/installserver/proxy.pp:65). I believe another approach, to check for a 200 OK, could have been to allow Prometheus hosts, via an ACL, to query the :8080/squid-internal-mgr/info endpoint.
Apr 1 2026
Mar 31 2026
Yes, as reported by volans, this has been done on purpose. The current maximum range length is 365 days, plus an additional 10 days to allow comparison over a 10-day window across one year. Unfortunately, there’s no way to tune the parameter on a per-query basis.
Anyway, the suggestion of adding a second query to the panel (or a second panel) with the offset query is a valid one.
If needed, I think we can add a few hours to the limit to reach a window of 1 year and 1 month.
Mar 30 2026
Mar 27 2026
topk(1000,
count by (metric_name) (
label_replace({__name__=~".+", job="k8s-pods"}, "metric_name", "$1", "__name__", "(.+)")
)
-
(
count by (metric_name) (
label_replace({__name__=~".+", job="k8s-pods"} offset 10d, "metric_name", "$1", "__name__", "(.+)")
)
or
(0 * count by (metric_name) (
label_replace({__name__=~".+", job="k8s-pods"}, "metric_name", "$1", "__name__", "(.+)")
))
)
)The ext alert is likely related to the DC switchover.
Related to the DC switchover. This will be resolved once the seasonality approach has enough data to correctly compute the standard pattern.
Mar 26 2026
root@titan2001:/srv/rewrite# journalctl -u thanos-compact | grep halt | tail -n 1 | sed -nr 's/^.*\[(.*)\].*$/\1/p' | tr -s ' ' '\n' | awk -F '/' '{print $NF}' | xargs -I % thanos tools bucket --objstore.config-file /etc/thanos-compact/objstore.yaml mark --id=% --marker=no-compact-mark.json --details="compactor halted due to size"
ts=2026-03-26T22:36:19.520898876Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T22:36:20.136700616Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KJ1B31QZ82XFTBRW8XM5B10F
ts=2026-03-26T22:36:20.136731522Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KJ1B31QZ82XFTBRW8XM5B10F
ts=2026-03-26T22:36:20.136763999Z caller=main.go:174 level=info msg=exiting
ts=2026-03-26T22:36:20.161640077Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T22:36:20.782733716Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KK5A391P07EHN0QRWMPY9MMG
ts=2026-03-26T22:36:20.782769235Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KK5A391P07EHN0QRWMPY9MMG
ts=2026-03-26T22:36:20.782808922Z caller=main.go:174 level=info msg=exitingroot@titan2001:/srv/rewrite# journalctl -u thanos-compact | grep halt | tail -n 1 | sed -nr 's/^.*\[(.*)\].*$/\1/p' | tr -s ' ' '\n' | awk -F '/' '{print $NF}' | xargs -I % thanos tools bucket --objstore.config-file /etc/thanos-compact/objstore.yaml mark --id=% --marker=no-compact-mark.json --details="compactor halted due to size"
ts=2026-03-26T20:40:55.969210977Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T20:40:56.573109549Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KK4MX9QZMCR6JHKV25C8TCFN
ts=2026-03-26T20:40:56.573155708Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KK4MX9QZMCR6JHKV25C8TCFN
ts=2026-03-26T20:40:56.57319678Z caller=main.go:174 level=info msg=exiting
ts=2026-03-26T20:40:56.60021435Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T20:40:57.320865233Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KK4XGGQCB44P5D5XTCN8YXGY
ts=2026-03-26T20:40:57.32092733Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KK4XGGQCB44P5D5XTCN8YXGY
ts=2026-03-26T20:40:57.320999165Z caller=main.go:174 level=info msg=exitingroot@titan2001:/srv/rewrite# journalctl -u thanos-compact | grep halt | tail -n 1 | sed -nr 's/^.*\[(.*)\].*$/\1/p' | tr -s ' ' '\n' | awk -F '/' '{print $NF}' | xargs -I % thanos tools bucket --objstore.config-file /etc/thanos-compact/objstore.yaml mark --id=% --marker=no-compact-mark.json --details="compactor halted due to size"
ts=2026-03-26T17:34:10.457880981Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T17:34:11.069489832Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KJ2W68ZNPN16MXB7D283WS31
ts=2026-03-26T17:34:11.069538396Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KJ2W68ZNPN16MXB7D283WS31
ts=2026-03-26T17:34:11.069588421Z caller=main.go:174 level=info msg=exiting
ts=2026-03-26T17:34:11.121816863Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T17:34:11.826060123Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KK0E1DGZWDEF0JJ1YF9T8SXB
ts=2026-03-26T17:34:11.826095643Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KK0E1DGZWDEF0JJ1YF9T8SXB
ts=2026-03-26T17:34:11.826132829Z caller=main.go:174 level=info msg=exiting
ts=2026-03-26T17:34:11.852982441Z caller=factory.go:54 level=info msg="loading bucket configuration"
ts=2026-03-26T17:34:12.491859477Z caller=block.go:406 level=info msg="block has been marked for no compaction" block=01KK4EHGC2VAASWWWS6KJ5ACHT
ts=2026-03-26T17:34:12.491908721Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=no-compact-mark.json IDs=01KK4EHGC2VAASWWWS6KJ5ACHT
ts=2026-03-26T17:34:12.49194897Z caller=main.go:174 level=info msg=exitingMar 25 2026
ts=2026-03-25T20:32:04.70702737Z caller=factory.go:54 level=info msg="loading bucket configuration" ts=2026-03-25T20:32:05.423033143Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KDG1XXFWVF3WK02GNR8QF44Z ts=2026-03-25T20:32:05.423071023Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KDG1XXFWVF3WK02GNR8QF44Z ts=2026-03-25T20:32:05.423098889Z caller=main.go:174 level=info msg=exiting ts=2026-03-25T20:32:05.450567006Z caller=factory.go:54 level=info msg="loading bucket configuration" ts=2026-03-25T20:32:06.036279314Z caller=block.go:203 level=info msg="block has been marked for deletion" block=01KE1TX1X3A0S0WXS3MY0TX42C ts=2026-03-25T20:32:06.03630991Z caller=tools_bucket.go:1134 level=info msg="marking done" marker=deletion-mark.json IDs=01KE1TX1X3A0S0WXS3MY0TX42C ts=2026-03-25T20:32:06.03635356Z caller=main.go:174 level=info msg=exiting