Page MenuHomePhabricator

colewhite (cwhite)
User

Projects (17)

Today

  • No visible events.

Tomorrow

  • No visible events.

Sunday

  • No visible events.

User Details

User Since
Aug 21 2018, 6:05 PM (416 w, 2 d)
Availability
Available
LDAP User
Cwhite
MediaWiki User
CWhite (WMF) [ Global Accounts ]

Recent Activity

Wed, Aug 12

colewhite created T434690: Create a tool that will remove files when disk exceeds a certain threshold.
Wed, Aug 12, 4:15 PM · Observability-Logging

Wed, Aug 5

colewhite renamed T434114: Scap support of http basic auth for Logstash queries from Scap support http basic auth for Logstash queries to Scap support of http basic auth for Logstash queries.
Wed, Aug 5, 4:37 PM · Observability-Logging, Scap
colewhite added a subtask for T350516: Enable OpenSearch security plugin - Beta Logs: T434114: Scap support of http basic auth for Logstash queries.
Wed, Aug 5, 4:37 PM · SRE Observability
colewhite added a parent task for T434114: Scap support of http basic auth for Logstash queries: T350516: Enable OpenSearch security plugin - Beta Logs.
Wed, Aug 5, 4:37 PM · Observability-Logging, Scap
colewhite created T434114: Scap support of http basic auth for Logstash queries.
Wed, Aug 5, 4:36 PM · Observability-Logging, Scap
colewhite added a comment to T433684: Create a template-radar pipeline.

From the Observability side, this is pretty easy to accomplish. I couldn't tell you how this can be accomplished on the producer side, though.

Wed, Aug 5, 2:26 PM · ServiceOps, observability

Tue, Aug 4

colewhite updated the task description for T350516: Enable OpenSearch security plugin - Beta Logs.
Tue, Aug 4, 8:58 PM · SRE Observability

Mon, Aug 3

colewhite closed T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet) as Resolved.

Seems the memory upgrade has had a positive effect. Optimistically resolving!

Mon, Aug 3, 10:08 PM · Observability-Metrics, SRE, Wikimedia-production-error
colewhite added a comment to T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format.

99.5% of these logs are warnings: "no index mapper found for field: [_type] returning default postings format".

Luckily, these logs only appear in cloudelastic (a non-production environment), so this issue should not be a blocker for production.

Mon, Aug 3, 10:05 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review, Observability-Logging

Thu, Jul 30

colewhite closed T432254: scap on deployment-deploy04 failing to reach logging-logstash-04.logging.eqiad1.wikimedia.cloud as Resolved.

I've not seen any more OOMs since lowering Xmx on the logstash hosts. Optimistically resolving, but do let me know if we see a recurrence!

Thu, Jul 30, 9:52 PM · Observability-Logging, Beta-Cluster-Infrastructure

Mon, Jul 27

colewhite added a comment to T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet).

Memory was upgraded on arclamp1001 today and I'm seeing used memory max out at around 32GB. I'll keep an eye on it.

Mon, Jul 27, 10:55 PM · Observability-Metrics, SRE, Wikimedia-production-error
colewhite added a comment to T433284: arclamp1001 ram upgrade.

Feel free to offline the hosts whenever you're ready. There's no paging associated and all I would do to depool is add host and services downtime in Icinga.

Mon, Jul 27, 5:26 PM · DC-Ops, ops-eqiad, Observability-Metrics, SRE
colewhite added a comment to T433285: arclamp2001 ram upgrade.

Feel free to offline the hosts whenever you're ready. There's no paging associated and all I would do to depool is add host and services downtime in Icinga.

Mon, Jul 27, 5:26 PM · DC-Ops, Observability-Metrics, SRE, ops-codfw

Thu, Jul 23

colewhite added a comment to T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet).

The restart doesn't appear to have alleviated the problem - I see more redis connection exceptions since the restart.

Thu, Jul 23, 12:27 AM · Observability-Metrics, SRE, Wikimedia-production-error

Wed, Jul 22

colewhite edited projects for T432881: Fields that only differ in capitalization override each other (?) in Logstash, added: WikimediaCustomizations; removed Wikimedia-Logstash.

Hey! Thanks for reaching out!

Wed, Jul 22, 5:38 PM · WikimediaCustomizations, Observability-Logging
colewhite claimed T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet).
Wed, Jul 22, 5:02 PM · Observability-Metrics, SRE, Wikimedia-production-error
colewhite edited projects for T432730: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet), added: Observability-Metrics; removed observability.

The redis logs don't have anything really interesting in them. I restarted redis on arclamp1001 as a first step to see if these errors clear up. I will circle back on the logs later this week to see if this issue persists.

Wed, Jul 22, 5:01 PM · Observability-Metrics, SRE, Wikimedia-production-error
colewhite assigned T432444: Provision kafka-logging100[6-8] and kafka-logging200[6-8] to tappof.

Planning to do after Aug 12.

Wed, Jul 22, 2:09 PM · Observability-Logging

Tue, Jul 21

colewhite added a comment to T343787: Find a replacement for the unmaintained eventrouter.

@colewhite @hnowlan is there a corresponding task on o11y side which we should link this to for future reference?

Tue, Jul 21, 3:35 PM · Observability-Logging, ServiceOps, ServiceOps-good-first-task, Technical-Debt, Prod-Kubernetes, Kubernetes
colewhite added a parent task for T343787: Find a replacement for the unmaintained eventrouter: T333731: Investigate pre-kafka log agent replacement for rsyslog.
Tue, Jul 21, 3:34 PM · Observability-Logging, ServiceOps, ServiceOps-good-first-task, Technical-Debt, Prod-Kubernetes, Kubernetes
colewhite added a subtask for T333731: Investigate pre-kafka log agent replacement for rsyslog: T343787: Find a replacement for the unmaintained eventrouter.
Tue, Jul 21, 3:34 PM · Observability-Logging

Mon, Jul 20

colewhite claimed T432254: scap on deployment-deploy04 failing to reach logging-logstash-04.logging.eqiad1.wikimedia.cloud.

I suspect the network unavailability as the cause because the victim of the oom-killer was Logstash and not OpenSearch.

Mon, Jul 20, 11:33 PM · Observability-Logging, Beta-Cluster-Infrastructure

Jul 10 2026

colewhite removed a project from T431826: July 2026 Bookworm/Trixie reboots: Data Platform SRE: SRE Observability.
Jul 10 2026, 4:49 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Vuln-VulnComponent, SecTeam-Processed, Infrastructure Security, SRE, Security

Jun 30 2026

colewhite closed T430739: beta logs MediaWiki errors dashboard shows few events after June 24, 2026 as Resolved.

Optimistically resolving as I see logs on the dashboard again.

Jun 30 2026, 9:20 PM · observability

Jun 25 2026

colewhite added a comment to T343787: Find a replacement for the unmaintained eventrouter.

This happens to be an ask we had not yet incorporated into our pre-kafka logging daemon evaluation. I have added it to the evaluation matrix.

Jun 25 2026, 8:46 PM · Observability-Logging, ServiceOps, ServiceOps-good-first-task, Technical-Debt, Prod-Kubernetes, Kubernetes

Jun 11 2026

colewhite added a project to T428832: db1262 crashed: ops-eqiad.
SeqNumber       = 574
Message ID      = CPU0000
Category        = System
AgentID         = iDRAC
Severity        = Information
Timestamp       = 2026-06-11 00:34:15
Message         = Internal error has occurred check for additional logs.
FQDD            = iDRAC.Embedded.1
Jun 11 2026, 12:49 AM · SRE, DC-Ops, ops-eqiad, Data-Persistence, DBA

Jun 10 2026

colewhite renamed T428813: Create an event.duration transformation to do time unit conversions for ECS events from Create a duration transformation to do time unit conversions for ECS messages to Create an event.duration transformation to do time unit conversions for ECS events.
Jun 10 2026, 8:51 PM · Observability-Logging
colewhite created T428813: Create an event.duration transformation to do time unit conversions for ECS events.
Jun 10 2026, 8:49 PM · Observability-Logging
colewhite added a comment to T421996: Create an automation against the logs.

Do you (or anyone reading this!) know if there's a way to have it send a periodic Slack summary of those aggregated messages, as a low-friction way to stay informed without active monitoring?

Jun 10 2026, 8:25 PM · SRE Observability

Jun 4 2026

colewhite added a comment to T425528: Rework ACLs on Kafka 3.x clusters.

@colewhite @tappof @andrea.denisse Hi! I have to add some ACLs to both Kafka logging clusters, I am going to add some rationale here and you can tell me what you think about it :)
...
We do have the same rules in the other clusters, so they are safe to be applied. With your permission, I'd add them one cluster at the time to keep things consistent across clusters. If any impact is registered, we can quickly rollback.

Let me know your thoughts :)

Jun 4 2026, 11:18 PM · Kafka-Infrastructure, SRE
colewhite added a comment to T422056: wdqs: database node logs should be pushed to logstash.

It this a hard limit or could we lift it? If wanted to target higher qps (say, > 1k qps), would logstash/opensearch be a suitable target?

Ideally we should be able to support that volume, but not at this time: T390215: Logstash is still overwhelmed (since March 2025) We're still working through capacity issues since the k8s migration and ECS adoption has been slow.

Jun 4 2026, 1:57 PM · Wikidata Platform Team (Sprint 08 (2026/08/04)), SRE Observability, OKR-Work
colewhite closed T427748: Degraded RAID on centrallog1002 as Resolved.

Raid is rebuilt. Thanks for the hardware and prompt response!

Jun 4 2026, 1:49 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad

Jun 3 2026

colewhite added a comment to T422056: wdqs: database node logs should be pushed to logstash.

For counting logs, I did a dummy query in OpenSearch with HOSTNAME : wdqs* (https://logstash.wikimedia.org/goto/28f8acb76ce48b797b395786dd2618f5). But this consistently reports an order of magnitude less (query) logs

Jun 3 2026, 10:43 PM · Wikidata Platform Team (Sprint 08 (2026/08/04)), SRE Observability, OKR-Work

Jun 2 2026

colewhite moved T427748: Degraded RAID on centrallog1002 from Radar to Watching on the Observability-Logging board.
Jun 2 2026, 10:32 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad
colewhite changed the status of T427748: Degraded RAID on centrallog1002 from Open to In Progress.
Jun 2 2026, 9:58 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad
colewhite added a comment to T427748: Degraded RAID on centrallog1002.

@colewhite can this be swapped at any time would you be able to rebuild after swapping?

Jun 2 2026, 1:54 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad

Jun 1 2026

colewhite added a comment to T390215: Logstash is still overwhelmed (since March 2025).

@colewhite I was thinking that we could also lower down the retention of the istio access logs to say one or two weeks, if it is 30 days or more now. We really don't care about history at the moment, a couple of weeks or even one should be ok (2 better of course). What do you think?

Jun 1 2026, 11:47 PM · SRE Observability (FY2025/2026-Q1), Patch-For-Review, Observability-Logging
colewhite moved T427748: Degraded RAID on centrallog1002 from Inbox to Radar on the Observability-Logging board.
Jun 1 2026, 11:43 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad
colewhite added a project to T427748: Degraded RAID on centrallog1002: Observability-Logging.
Jun 1 2026, 11:42 PM · Observability-Logging, DC-Ops, SRE, ops-eqiad

May 28 2026

colewhite added a comment to T327033: Special:ExtensionDistributor no longer shows popular extensions list.

I have 2 questions;
(1) are we still able to record metrics to the graphite store somehow (Cc @Krinkle / @colewhite)? If so, how come?

May 28 2026, 5:42 PM · ExtensionDistributor
colewhite added a comment to T427306: Migrate Relforge clusters from OpenSearch 1.x->2x.

We've gotten away with this because we haven't formed a net-new cluster since the previous config was deprecated, but if we won't be able to create a net-new cluster until this is addressed. CCing @colewhite and observability as this affects the observability cluster as well.

Cole/0lly, if you are not comfortable deploying these changes on y'all's clusters, we can also use a systemd override to add values as environment variables:

May 28 2026, 12:20 AM · Discovery-Search (2026.06.01 - 2026.07.03), Data-Platform-SRE (2026-06-05 - 2026-06-26)

May 21 2026

colewhite claimed T426368: Upgrade XHGui from 0.21.3 to latest (0.23.6).
May 21 2026, 2:00 PM · Observability-Metrics, WikimediaDebug

May 19 2026

colewhite updated the task description for T424204: profile/module violations in use of profile::pki::get_cert().
May 19 2026, 9:53 PM · SRE Observability (FY2025/2026-Q4), Data-Persistence, ServiceOps, Infrastructure-Foundations
colewhite closed T415289: logging-hd nodes evict-rejoin troubles as Invalid.

Resolution was found in T422048: SLOBudgetBurn

May 19 2026, 9:22 PM · Observability-Logging
colewhite removed a subtask for T390215: Logstash is still overwhelmed (since March 2025): T415289: logging-hd nodes evict-rejoin troubles.
May 19 2026, 9:21 PM · SRE Observability (FY2025/2026-Q1), Patch-For-Review, Observability-Logging
colewhite removed a parent task for T415289: logging-hd nodes evict-rejoin troubles: T390215: Logstash is still overwhelmed (since March 2025).
May 19 2026, 9:21 PM · Observability-Logging
colewhite renamed T390215: Logstash is still overwhelmed (since March 2025) from Logstash is overwhelmed (March 2025) to Logstash is still overwhelmed (since March 2025).
May 19 2026, 4:14 PM · SRE Observability (FY2025/2026-Q1), Patch-For-Review, Observability-Logging

Apr 28 2026

colewhite updated the task description for T424673: Migrate o11y Envoy TLS proxy services to the 2026 discovery intermediate.
Apr 28 2026, 4:45 PM · observability

Apr 27 2026

colewhite added a comment to T424574: arclamp gzip logs are not gzipped.

The file on the arclamp server is 83M:

-rw-r--r-- 1 xenon xenon   83M Apr 22 23:59 2026-04-22.excimer-wall.all.log.gz

When downloading it, Chrome reports the file size as 7.6G. Examining the file locally:

$ ls -lah | grep excimer
-rw-r--r--. 1 cwhite cwhite 7.6G Apr 27 23:12 2026-04-22.excimer-wall.all.log.gz
$ file 2026-04-22.excimer-wall.all.log.gz 
2026-04-22.excimer-wall.all.log.gz: ASCII text, with very long lines (1961)
Apr 27 2026, 11:39 PM · Arc-Lamp

Apr 22 2026

colewhite closed T422048: SLOBudgetBurn as Resolved.

Haven't had another event since. Optimistically resolving.

Apr 22 2026, 2:49 AM · observability

Apr 17 2026

colewhite updated the task description for T350516: Enable OpenSearch security plugin - Beta Logs.
Apr 17 2026, 11:06 PM · SRE Observability
colewhite updated the task description for T350516: Enable OpenSearch security plugin - Beta Logs.
Apr 17 2026, 10:36 PM · SRE Observability

Apr 16 2026

colewhite added a comment to T423327: Explore options for OpenSearch 2.x/3.x plugin packaging and distribution.

Copying my comment here too, now that I've seen this task:

Something to consider:
Apr 16 2026, 4:23 PM · Patch-For-Review, Data-Platform-SRE (2026-03-27 - 2026-04-17)

Apr 15 2026

colewhite added a comment to T422068: Consider "inner" and "outer" ssh keys to reduce taps through the day.

I wonder if ssh certificates could be of use? I imagine a "tap to access bastion, tap to sign short-lived certificate" then on subsequent logins maybe just one tap to make the first jump?

Apr 15 2026, 10:24 PM · SRE, Infrastructure Security
colewhite removed a project from T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format: observability.
Apr 15 2026, 2:10 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Patch-For-Review, Observability-Logging

Apr 14 2026

colewhite closed T267664: Enhance smart_data_dump to support gathering metrics from both raid and standalone disks, a subtask of T267135: smart-data-dump should fail loudly when it can't gather metrics, as Resolved.
Apr 14 2026, 11:29 PM · Observability-Alerting, SRE
colewhite closed T267664: Enhance smart_data_dump to support gathering metrics from both raid and standalone disks as Resolved.

Change is deployed and seems to pass the smoke test. Optimistically closing.

Apr 14 2026, 11:29 PM · SRE Observability

Apr 8 2026

colewhite added a comment to T422048: SLOBudgetBurn.

I overwrote that sector in the hope that the disk will reallocate it and OpenSearch correct any data discrepancies when it gets around to it.

Apr 8 2026, 5:08 PM · observability
colewhite added a comment to T422048: SLOBudgetBurn.

I'm thinking disk read errors:

2026-04-07T04:56:16.083470+00:00 logging-hd2001 kernel: [1672676.995139] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083496+00:00 logging-hd2001 kernel: [1672676.995151] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083498+00:00 logging-hd2001 kernel: [1672676.995155] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083499+00:00 logging-hd2001 kernel: [1672676.995164] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083500+00:00 logging-hd2001 kernel: [1672676.995169] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083502+00:00 logging-hd2001 kernel: [1672676.995175] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:16.083503+00:00 logging-hd2001 kernel: [1672676.995195] sd 0:0:1:0: [sdb] tag#6784 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=2s
2026-04-07T04:56:16.083504+00:00 logging-hd2001 kernel: [1672676.995204] sd 0:0:1:0: [sdb] tag#6784 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:16.083506+00:00 logging-hd2001 kernel: [1672676.995209] sd 0:0:1:0: [sdb] tag#6784 Add. Sense: Read retries exhausted
2026-04-07T04:56:16.083507+00:00 logging-hd2001 kernel: [1672676.995214] sd 0:0:1:0: [sdb] tag#6784 CDB: Read(16) 88 00 00 00 00 01 1e e7 30 00 00 00 04 00 00 00
2026-04-07T04:56:16.083508+00:00 logging-hd2001 kernel: [1672676.995218] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x80700 phys_seg 5 prio class 2
2026-04-07T04:56:19.210392+00:00 logging-hd2001 kernel: [1672680.153496] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:19.210409+00:00 logging-hd2001 kernel: [1672680.153501] sd 0:0:1:0: [sdb] tag#6815 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=2s
2026-04-07T04:56:19.210410+00:00 logging-hd2001 kernel: [1672680.153504] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:19.210411+00:00 logging-hd2001 kernel: [1672680.153513] sd 0:0:1:0: [sdb] tag#6815 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:19.210413+00:00 logging-hd2001 kernel: [1672680.153518] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:19.210414+00:00 logging-hd2001 kernel: [1672680.153521] sd 0:0:1:0: [sdb] tag#6815 Add. Sense: Read retries exhausted
2026-04-07T04:56:19.210437+00:00 logging-hd2001 kernel: [1672680.153525] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:19.210438+00:00 logging-hd2001 kernel: [1672680.153528] sd 0:0:1:0: [sdb] tag#6815 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:19.210440+00:00 logging-hd2001 kernel: [1672680.153533] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:22.318883+00:00 logging-hd2001 kernel: [1672683.261973] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:22.318906+00:00 logging-hd2001 kernel: [1672683.261977] sd 0:0:1:0: [sdb] tag#6817 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:22.318909+00:00 logging-hd2001 kernel: [1672683.261981] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:22.318911+00:00 logging-hd2001 kernel: [1672683.261991] sd 0:0:1:0: [sdb] tag#6817 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:22.318912+00:00 logging-hd2001 kernel: [1672683.261997] sd 0:0:1:0: [sdb] tag#6817 Add. Sense: Read retries exhausted
2026-04-07T04:56:22.318914+00:00 logging-hd2001 kernel: [1672683.262003] sd 0:0:1:0: [sdb] tag#6817 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:22.318915+00:00 logging-hd2001 kernel: [1672683.262006] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:25.443753+00:00 logging-hd2001 kernel: [1672686.386824] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:25.443778+00:00 logging-hd2001 kernel: [1672686.386848] sd 0:0:1:0: [sdb] tag#6826 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:25.443781+00:00 logging-hd2001 kernel: [1672686.386861] sd 0:0:1:0: [sdb] tag#6826 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:25.443782+00:00 logging-hd2001 kernel: [1672686.386867] sd 0:0:1:0: [sdb] tag#6826 Add. Sense: Read retries exhausted
2026-04-07T04:56:25.443784+00:00 logging-hd2001 kernel: [1672686.386873] sd 0:0:1:0: [sdb] tag#6826 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:25.443785+00:00 logging-hd2001 kernel: [1672686.386877] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:28.518821+00:00 logging-hd2001 kernel: [1672689.461882] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:28.518837+00:00 logging-hd2001 kernel: [1672689.461908] sd 0:0:1:0: [sdb] tag#6846 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:28.518839+00:00 logging-hd2001 kernel: [1672689.461920] sd 0:0:1:0: [sdb] tag#6846 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:28.518841+00:00 logging-hd2001 kernel: [1672689.461926] sd 0:0:1:0: [sdb] tag#6846 Add. Sense: Read retries exhausted
2026-04-07T04:56:28.518842+00:00 logging-hd2001 kernel: [1672689.461931] sd 0:0:1:0: [sdb] tag#6846 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:28.518844+00:00 logging-hd2001 kernel: [1672689.461935] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:31.610758+00:00 logging-hd2001 kernel: [1672692.553802] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:31.610769+00:00 logging-hd2001 kernel: [1672692.553806] sd 0:0:1:0: [sdb] tag#6788 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:31.610772+00:00 logging-hd2001 kernel: [1672692.553816] sd 0:0:1:0: [sdb] tag#6788 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:31.610773+00:00 logging-hd2001 kernel: [1672692.553823] sd 0:0:1:0: [sdb] tag#6788 Add. Sense: Read retries exhausted
2026-04-07T04:56:31.610774+00:00 logging-hd2001 kernel: [1672692.553829] sd 0:0:1:0: [sdb] tag#6788 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:31.610776+00:00 logging-hd2001 kernel: [1672692.553832] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:34.819055+00:00 logging-hd2001 kernel: [1672695.762075] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:34.819080+00:00 logging-hd2001 kernel: [1672695.762081] sd 0:0:1:0: [sdb] tag#6789 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:34.819084+00:00 logging-hd2001 kernel: [1672695.762095] sd 0:0:1:0: [sdb] tag#6789 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:34.819086+00:00 logging-hd2001 kernel: [1672695.762100] sd 0:0:1:0: [sdb] tag#6789 Add. Sense: Read retries exhausted
2026-04-07T04:56:34.819087+00:00 logging-hd2001 kernel: [1672695.762106] sd 0:0:1:0: [sdb] tag#6789 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:34.819088+00:00 logging-hd2001 kernel: [1672695.762110] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:37.894085+00:00 logging-hd2001 kernel: [1672698.837089] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:37.894109+00:00 logging-hd2001 kernel: [1672698.837113] sd 0:0:1:0: [sdb] tag#6818 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:37.894112+00:00 logging-hd2001 kernel: [1672698.837126] sd 0:0:1:0: [sdb] tag#6818 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:37.894113+00:00 logging-hd2001 kernel: [1672698.837132] sd 0:0:1:0: [sdb] tag#6818 Add. Sense: Read retries exhausted
2026-04-07T04:56:37.894115+00:00 logging-hd2001 kernel: [1672698.837138] sd 0:0:1:0: [sdb] tag#6818 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:37.894116+00:00 logging-hd2001 kernel: [1672698.837141] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:41.027614+00:00 logging-hd2001 kernel: [1672701.970600] mpt3sas_cm0: log_info(0x31080000): originator(PL), code(0x08), sub_code(0x0000)
2026-04-07T04:56:41.027638+00:00 logging-hd2001 kernel: [1672701.970624] sd 0:0:1:0: [sdb] tag#6829 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=3s
2026-04-07T04:56:41.027640+00:00 logging-hd2001 kernel: [1672701.970636] sd 0:0:1:0: [sdb] tag#6829 Sense Key : Medium Error [current] [descriptor] 
2026-04-07T04:56:41.027642+00:00 logging-hd2001 kernel: [1672701.970642] sd 0:0:1:0: [sdb] tag#6829 Add. Sense: Read retries exhausted
2026-04-07T04:56:41.027643+00:00 logging-hd2001 kernel: [1672701.970647] sd 0:0:1:0: [sdb] tag#6829 CDB: Read(16) 88 00 00 00 00 01 1e e7 33 80 00 00 00 08 00 00
2026-04-07T04:56:41.027645+00:00 logging-hd2001 kernel: [1672701.970651] critical medium error, dev sdb, sector 4813435780 op 0x0:(READ) flags 0x0 phys_seg 1 prio class 2
2026-04-07T04:56:42.866662+00:00 logging-hd2001 opensearch[378008]: fatal error in thread [opensearch[logging-hd2001-production-elk7-codfw][generic][T#22]], exiting
2026-04-07T04:56:42.867561+00:00 logging-hd2001 opensearch[378008]: java.lang.InternalError: a fault occurred in a recent unsafe memory access operation in compiled Java code
2026-04-07T04:56:42.867655+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.store.BufferedChecksumIndexInput.readBytes(BufferedChecksumIndexInput.java:46)
2026-04-07T04:56:42.902773+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.store.DataInput.readBytes(DataInput.java:72)
2026-04-07T04:56:42.902903+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.store.ChecksumIndexInput.skipByReading(ChecksumIndexInput.java:79)
2026-04-07T04:56:42.937754+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.store.ChecksumIndexInput.seek(ChecksumIndexInput.java:64)
2026-04-07T04:56:42.937876+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.CodecUtil.checksumEntireFile(CodecUtil.java:618)
2026-04-07T04:56:42.937946+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.lucene90.Lucene90PostingsReader.checkIntegrity(Lucene90PostingsReader.java:2049)
2026-04-07T04:56:42.938023+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.lucene90.blocktree.Lucene90BlockTreeTermsReader.checkIntegrity(Lucene90BlockTreeTermsReader.java:330)
2026-04-07T04:56:42.938114+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.perfield.PerFieldPostingsFormat$FieldsReader.checkIntegrity(PerFieldPostingsFormat.java:370)
2026-04-07T04:56:42.938261+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.perfield.PerFieldMergeState$FilterFieldsProducer.checkIntegrity(PerFieldMergeState.java:296)
2026-04-07T04:56:42.938384+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.FieldsConsumer.merge(FieldsConsumer.java:83)
2026-04-07T04:56:42.938459+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.codecs.perfield.PerFieldPostingsFormat$FieldsWriter.merge(PerFieldPostingsFormat.java:205)
2026-04-07T04:56:42.938531+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.SegmentMerger.mergeTerms(SegmentMerger.java:209)
2026-04-07T04:56:42.938623+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.SegmentMerger.mergeWithLogging(SegmentMerger.java:298)
2026-04-07T04:56:42.938727+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.SegmentMerger.merge(SegmentMerger.java:137)
2026-04-07T04:56:42.938841+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.IndexWriter.mergeMiddle(IndexWriter.java:5140)
2026-04-07T04:56:42.946807+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.IndexWriter.merge(IndexWriter.java:4680)
2026-04-07T04:56:42.946917+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.IndexWriter$IndexWriterMergeSource.merge(IndexWriter.java:6432)
2026-04-07T04:56:42.947003+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.ConcurrentMergeScheduler.doMerge(ConcurrentMergeScheduler.java:639)
2026-04-07T04:56:42.947078+00:00 logging-hd2001 opensearch[378008]: #011at org.opensearch.index.engine.OpenSearchConcurrentMergeScheduler.doMerge(OpenSearchConcurrentMergeScheduler.java:120)
2026-04-07T04:56:43.004903+00:00 logging-hd2001 opensearch[378008]: #011at org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:700)
2026-04-07T04:56:44.157299+00:00 logging-hd2001 systemd[1]: opensearch_2@production-elk7-codfw.service: Main process exited, code=exited, status=128/n/a
2026-04-07T04:56:44.174262+00:00 logging-hd2001 systemd[1]: opensearch_2@production-elk7-codfw.service: Failed with result 'exit-code'.
2026-04-07T04:56:44.186310+00:00 logging-hd2001 systemd[1]: opensearch_2@production-elk7-codfw.service: Consumed 6h 53min 12.273s CPU time.
Apr 8 2026, 2:23 PM · observability

Apr 3 2026

colewhite added a comment to T390215: Logstash is still overwhelmed (since March 2025).

The istio-system namespace is logging ~980 events/sec. Many are just istio-ingressgateway for authority:page-analytics.discovery.wmnet (~833 events/sec).

Apr 3 2026, 3:23 AM · SRE Observability (FY2025/2026-Q1), Patch-For-Review, Observability-Logging

Mar 25 2026

colewhite added a comment to T418929: Q4:rack/setup/install kafka-logging100[6-8].

I'd like to propose a rename to logging-kafka so that these hosts follow the other logging-* hosts indicating its role in the larger cluster.

Mar 25 2026, 3:41 PM · observability, SRE, ops-eqiad, DC-Ops
colewhite added a comment to T418931: Q3:rack/setup/install kafka-logging200[6-8].

I'd like to propose a rename to logging-kafka so that these hosts follow the other logging-* hosts indicating its role in the larger cluster.

Mar 25 2026, 3:41 PM · observability, SRE, ops-codfw, DC-Ops

Mar 16 2026

colewhite updated the task description for T420158: Eqiad: lsw1-c2-eqiad BGP maintenance/ Tuesday 17th at 9:30 CDT.
Mar 16 2026, 10:45 PM · Data-Platform-SRE (2026-03-06 - 2026-03-27), ServiceOps, netops, Infrastructure-Foundations, SRE
colewhite updated subscribers of T413127: Directory Listing and Download from Object Storage.

2M files can be too much for swift for one container and generally the backends. I have some suggestions:

  • Maybe drop all hourly graphs and logs after three months?
  • We could also drop all non "all" graphs and logs after a year too?

I think dropping those would save a lot of space without dropping too much useful stuff.

Mar 16 2026, 8:51 PM · MediaWiki-Core-Platform-Team (Radar), Data-Persistence, Arc-Lamp
colewhite added a comment to T420034: deployment-kafka-logging01 is down for maintenance because Trixie is not yet well supported.

@colewhite I see logstash consumer groups connecting, when you have a moment could you verify if everything works for beta logs?

Mar 16 2026, 4:56 PM · Beta-Cluster-Infrastructure

Mar 13 2026

colewhite edited projects for T417001: Upgrade Observability Kafka-logging hosts to trixie, added: Observability-Logging; removed SRE Observability.
Mar 13 2026, 5:17 PM · Observability-Logging
colewhite added a comment to T416384: Reduce logstash logs from machine learning infra.

I think we are out of the woods, we have around ~20k/minute logs now mostly coming from the inference-service pods (kserve-container). I think that we could definitely improve things there:

  1. The logs are not emitted using json or ECS so in case of errors, like Python stacktraces, we get one log for each line. It is a waste on the logstash side, but also not really great for human readers that need to investigate an outage the day afterwards. If the logs are not on the pods because of rotation, getting a complete stacktrace from logstash is really really tedious.
  1. We have a mixture of kserve traces, unicorn access logs, latency timings etc.. Do we need all of them?
Mar 13 2026, 3:42 PM · Machine-Learning-Team

Mar 12 2026

colewhite closed T414098: Move https://status.wikimedia.org/ away from rackspace, a subtask of T376400: Redesign wikitech-static, as Resolved.
Mar 12 2026, 5:37 PM · Patch-For-Review, serviceops-radar, SRE-Unowned, SRE, wikitech.wikimedia.org
colewhite closed T414098: Move https://status.wikimedia.org/ away from rackspace as Resolved.

Last thing the old rackspace host handles is [301] http[s]://wikimediastatus.net -> https://www.wikimediastatus.net

Mar 12 2026, 5:37 PM · Patch-For-Review, Collaboration-Services, SRE Observability, cloud-services-team
colewhite closed T414098: Move https://status.wikimedia.org/ away from rackspace, a subtask of T408704: offline rackspace wikitech-static, online aws wikitech-static, as Resolved.
Mar 12 2026, 5:37 PM · Infrastructure-Foundations
colewhite added a parent task for T419887: Move wikimediastatus.net 301 to ncredir: T408704: offline rackspace wikitech-static, online aws wikitech-static.
Mar 12 2026, 5:36 PM · Patch-For-Review, Traffic, SRE Observability
colewhite added a subtask for T408704: offline rackspace wikitech-static, online aws wikitech-static: T419887: Move wikimediastatus.net 301 to ncredir.
Mar 12 2026, 5:36 PM · Infrastructure-Foundations
colewhite created T419887: Move wikimediastatus.net 301 to ncredir.
Mar 12 2026, 5:36 PM · Patch-For-Review, Traffic, SRE Observability
colewhite reopened T414098: Move https://status.wikimedia.org/ away from rackspace, a subtask of T376400: Redesign wikitech-static, as Open.
Mar 12 2026, 5:32 PM · Patch-For-Review, serviceops-radar, SRE-Unowned, SRE, wikitech.wikimedia.org
colewhite reopened T414098: Move https://status.wikimedia.org/ away from rackspace as "Open".

Last thing the old rackspace host handles is [301] http[s]://wikimediastatus.net -> https://www.wikimediastatus.net

Mar 12 2026, 5:32 PM · Patch-For-Review, Collaboration-Services, SRE Observability, cloud-services-team
colewhite reopened T414098: Move https://status.wikimedia.org/ away from rackspace, a subtask of T408704: offline rackspace wikitech-static, online aws wikitech-static, as Open.
Mar 12 2026, 5:32 PM · Infrastructure-Foundations

Mar 11 2026

colewhite added a comment to T413127: Directory Listing and Download from Object Storage.

Can you give me an idea of number & size of objects, and what sort of bandwidth you expect this to need, please?

The files I'd like to serve can be found in swift at https://ms-fe.svc.eqiad.wmnet/v1/AUTH_performance/arclamp-(logs|svgs)-(hourly|daily)
Some napkin math and a 3-year retention period (currently configured) yields 2,023,560 files. and ~634GB of data. Files range from a few hundred bytes to a little over a hundred megabytes.

Mar 11 2026, 8:36 PM · MediaWiki-Core-Platform-Team (Radar), Data-Persistence, Arc-Lamp

Mar 6 2026

colewhite added a project to T418612: Audit mwlog storage and retention: Observability-Logging.

Poolcounter logs are greatly reduced post-deploy.

$ zcat poolcounter.log-20260304*.gz | wc -l
844298029
$ zcat poolcounter.log-20260305*.gz | wc -l
73026
Mar 6 2026, 12:57 AM · Observability-Logging, SRE Observability (FY2025/2026-Q3)

Mar 2 2026

colewhite updated the task description for T418772: Eqiad: lsw1-d7-eqiad BGP maintenance.
Mar 2 2026, 9:17 PM · Prod-Kubernetes, ServiceOps, netops, Infrastructure-Foundations, SRE

Feb 24 2026

colewhite added a project to T249663: write some recording rules for queries used in the appserver RED k8s dashboard: Observability-Metrics.
Feb 24 2026, 10:24 PM · Observability-Metrics, SRE Observability (FY2025/2026-Q3), Prod-Kubernetes, ServiceOps, SRE
colewhite added a comment to T343020: Converting MediaWiki Metrics to StatsLib.

However it has now received code review and is just waiting on response from the submitter, so it probably no longer needs to be in that column.

Feb 24 2026, 10:19 PM · SRE Observability (FY2025/2026-Q1), Essential-Work, MW-1.44-notes (1.44.0-wmf.28; 2025-05-06), Patch-For-Review, Observability-Metrics
colewhite added a project to T416863: csp-report-only topic has exploded in size: Observability-Logging.
Feb 24 2026, 10:53 AM · SecTeam-Processed, Security-Team, Observability-Logging, SRE Observability (FY2025/2026-Q3), ContentSecurityPolicy

Feb 23 2026

colewhite closed T418063: wikimediastatus.net has expired certificate as Resolved.

As part of T376400: Redesign wikitech-static, the backup wikitech moved away from the legacy wikitech-static host. Part of that process was pointing wikitech-static.wikimedia.org away from the legacy host to its new home. This made it impossible for certbot to complete its certificate renewal for a host that it no longer served. This renewal discontinuity interrupted wikimediastatus.net and status.wikimedia.org renewals.

Feb 23 2026, 8:18 PM · Incident Tooling, Traffic

Jan 28 2026

colewhite created T415784: GitLab Explore projects "Maximum Page Reached".
Jan 28 2026, 1:31 PM · GitLab (Upstream pit of despair 🕳️)

Jan 26 2026

colewhite closed T409363: Setup service name for Beta Cluster access to logstash service in logging project as Resolved.
Jan 26 2026, 10:02 AM · Scap, Observability-Logging, Beta-Cluster-Infrastructure

Jan 22 2026

colewhite added a subtask for T390215: Logstash is still overwhelmed (since March 2025): T415289: logging-hd nodes evict-rejoin troubles.
Jan 22 2026, 4:59 PM · SRE Observability (FY2025/2026-Q1), Patch-For-Review, Observability-Logging
colewhite added a parent task for T415289: logging-hd nodes evict-rejoin troubles: T390215: Logstash is still overwhelmed (since March 2025).
Jan 22 2026, 4:59 PM · Observability-Logging
colewhite created T415289: logging-hd nodes evict-rejoin troubles.
Jan 22 2026, 4:59 PM · Observability-Logging
colewhite closed T409339: Beta Cluster MediaWiki updates require logging-logstash-02.logging.eqiad1.wikimedia.cloud to allow access to port 9200 by `scap` as Resolved.

I've added monitoring and alerting to the new cname record. Considering this done!

Jan 22 2026, 4:58 PM · Observability-Logging, Beta-Cluster-Infrastructure, Scap
colewhite updated subscribers of T415270: librenms.syslog table is 800GB.
Jan 22 2026, 3:46 PM · observability, netops, Infrastructure-Foundations, DBA

Jan 21 2026

colewhite closed T414670: Change units for "network utilization" on "host overview" dashboard to bits/sec as Resolved.

Was bold and made this change. Will direct here for further discussion if there are other opinions.

Jan 21 2026, 9:36 PM · Observability-Metrics, SRE Observability (FY2025/2026-Q3)

Jan 15 2026

colewhite added a comment to T414670: Change units for "network utilization" on "host overview" dashboard to bits/sec.

I'm +1 for this change. Usually, I'm looking at the network panels to see where utilization is relative to the link speed. It'd be nice to not have to manually calculate it.

Jan 15 2026, 10:42 PM · Observability-Metrics, SRE Observability (FY2025/2026-Q3)
colewhite added a comment to T414648: Request to increase quotas for logging project.

Thank you!!

Jan 15 2026, 4:32 PM · Observability-Logging, Cloud-VPS (Quota-requests)

Jan 14 2026

colewhite added a project to T414648: Request to increase quotas for logging project: Observability-Logging.
Jan 14 2026, 11:45 PM · Observability-Logging, Cloud-VPS (Quota-requests)
colewhite created T414648: Request to increase quotas for logging project.
Jan 14 2026, 11:45 PM · Observability-Logging, Cloud-VPS (Quota-requests)
colewhite added a comment to T414501: Improve tooling for long-running Thanos queries.

For context, the outage was caused by saturated nics on the titan hosts.

Jan 14 2026, 8:36 PM · SRE Observability
colewhite added a comment to T414607: Improve Grafana scalability.

For visibility, the outage today was a "grafana consumed all the memory" condition. https://grafana-next.wikimedia.org which links to the read-only backup in the standby DC remained available and was used in to diagnose the primary grafana host.

Jan 14 2026, 8:31 PM · SRE Observability

Jan 12 2026

colewhite added a comment to T414098: Move https://status.wikimedia.org/ away from rackspace.

We (observability) asked about this in the all-SRE meeting and most preferred to 301 status.wm.o -> wikimediastatus.net on the grounds that there are many Wikimedia properties using third-party tools and hosting with different privacy policies. Another option presented was to host a small static page on miscweb.

Jan 12 2026, 9:29 PM · Patch-For-Review, Collaboration-Services, SRE Observability, cloud-services-team

Jan 6 2026

colewhite added a comment to T413842: Icinga config can break on ensure=>absent in nagios_common::check_command::config.

Puppet is running on icinga again after that last patch.

Jan 6 2026, 2:09 AM · Observability-Alerting
colewhite renamed T413842: Icinga config can break on ensure=>absent in nagios_common::check_command::config from Icinga config broken due to ensure absent to Icinga config can break on ensure=>absent in nagios_common::check_command::config.
Jan 6 2026, 2:08 AM · Observability-Alerting