Page MenuHomePhabricator

BWojtowicz-WMF (bwojtowicz)
User

Today

  • No visible events.

Tomorrow

  • No visible events.

Saturday

  • No visible events.

User Details

User Since
May 6 2025, 11:26 AM (66 w, 2 d)
Availability
Available
LDAP User
Bartosz Wójtowicz
MediaWiki User
BWojtowicz-WMF [ Global Accounts ]

Recent Activity

Today

BWojtowicz-WMF added a comment to T431017: Pods evicted on ml-serve-eqiad due to disk pressure.

FYI we have added ephemeral-storage requests/limits to all predictor containers that we are serving on the LiftWing cluster and everything is now redeployed with storage declared on staging and prod clusters. It was tracked here: https://phabricator.wikimedia.org/T431089.

Thu, Aug 13, 10:15 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF moved T434763: Resolve issues with python-mwapi dependency on LiftWing from Unsorted/Needs Triage to Ready to Go on the Machine-Learning-Team (Q1 FY2026-27) board.
Thu, Aug 13, 10:05 AM · Essential-Work, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T433319: Update Revise Tone pipelines to incorporate additional languages and update topic filters.

We have now disabled the topic filtering alltogether.

Thu, Aug 13, 9:55 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T407843: Introduce re-try mechanisms for MW API requests in LiftWing models.

This task should be picked up only after we resolve the work tracked in here https://phabricator.wikimedia.org/T434763 since it relies heavily on the underlying package we use for MW API calls.

Thu, Aug 13, 9:48 AM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF created T434763: Resolve issues with python-mwapi dependency on LiftWing.
Thu, Aug 13, 9:47 AM · Essential-Work, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T432128: Create dataset of article titles for TTS V1.

The second version of the dataset is now created and available here:

Thu, Aug 13, 8:13 AM · Machine-Learning-Team (Q1 FY2026-27)

Yesterday

BWojtowicz-WMF added a comment to T433319: Update Revise Tone pipelines to incorporate additional languages and update topic filters.

We've done a first bump of allowed topics to 30 out of 64 topics covering Culture, History_and_Society topics plus STEM.Technology and there are no problems so far.
We'll let it run for a day to make sure everything is working smoothly and we'll enable remaining topics if there will be no problems.

Wed, Aug 12, 12:49 PM · Machine-Learning-Team (Q1 FY2026-27)

Tue, Aug 11

BWojtowicz-WMF moved T433319: Update Revise Tone pipelines to incorporate additional languages and update topic filters from Ready to Go to In Progress on the Machine-Learning-Team (Q1 FY2026-27) board.
Tue, Aug 11, 2:36 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF claimed T433319: Update Revise Tone pipelines to incorporate additional languages and update topic filters.
Tue, Aug 11, 1:38 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF created T434534: Add ephemeral-storage to KServe sidecar containers and enforce it on the chart level.
Tue, Aug 11, 1:08 PM · Essential-Work, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF claimed T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction.
Tue, Aug 11, 12:51 PM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF moved T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Tue, Aug 11, 12:51 PM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF closed T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction, a subtask of T433975: [draft] Q1 FY2026-27 LLM Platform Milestone 2, as Resolved.
Tue, Aug 11, 12:51 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction as Resolved.
Tue, Aug 11, 12:51 PM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF added a comment to T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction.

I've also updated non-KServe deployments to include ephemeral-storage limits/requests and re-deployed them across staging and prod clusters.

Tue, Aug 11, 12:51 PM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF added a comment to T432128: Create dataset of article titles for TTS V1.

Thank you for the review!

Tue, Aug 11, 8:31 AM · Machine-Learning-Team (Q1 FY2026-27)

Fri, Aug 7

BWojtowicz-WMF added a comment to T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction.

All KServe deployments on staging and production are now re-deployed with ephemeral-storage request and limits declared. During the work on this task I also noticed that we have some deployments not declaring memory and cpu requests/limits either, which can be tackled in a follow up ticket. Once this will be done, we could add requirements for all cpu/memory/storage be required on the chart level.

Fri, Aug 7, 8:29 AM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work

Wed, Aug 5

BWojtowicz-WMF moved T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction from Ready to Go to In Progress on the Machine-Learning-Team (Q1 FY2026-27) board.
Wed, Aug 5, 10:48 AM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF moved T432933: gpt-oss-safeguard-20b fails with 500 when output cannot be parsed as Harmony messages from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Wed, Aug 5, 10:43 AM · Patch-For-Review, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T432933: gpt-oss-safeguard-20b fails with 500 when output cannot be parsed as Harmony messages as Resolved.
Wed, Aug 5, 10:43 AM · Patch-For-Review, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T432933: gpt-oss-safeguard-20b fails with 500 when output cannot be parsed as Harmony messages.

The gpt-oss-safeguard-20b service was hardened such that if the model output cannot be parsed as Harmony messages, was truncated at max_tokens, or contains no final channel (degenerate generation on pathological input, or reasoning loops exhausting the token budget), the server returns a readable error response instead of an HTTP 500 or a bogus verdict.

Wed, Aug 5, 10:43 AM · Patch-For-Review, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF created T434059: Enable Structured Outputs for LLMs.
Wed, Aug 5, 10:40 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T432128: Create dataset of article titles for TTS V1.

I've performed first full run of dataset generation and I'm attaching the resulting dataset csv file here (columns: Topic, Article title, Article URL, Type, Edit rate in the last 7 days).

Wed, Aug 5, 8:29 AM · Machine-Learning-Team (Q1 FY2026-27)

Thu, Jul 30

BWojtowicz-WMF created T433580: Research current best approaches for hosting LLM models.
Thu, Jul 30, 8:42 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T429540: Hydrate mobile-html response with Topics metadata.

@BWojtowicz-WMF , should this be using the new Linked Artifacts Cache topics endpoint?

Thu, Jul 30, 8:34 AM · App Experience (AppEx Sprint [FY2627 Jul 29 - Aug 11])

Mon, Jul 27

BWojtowicz-WMF added a comment to T432128: Create dataset of article titles for TTS V1.

I've started working on the script to generate the dataset and I've come across a few questions / things to be confirmed:

Mon, Jul 27, 12:54 PM · Machine-Learning-Team (Q1 FY2026-27)

Thu, Jul 23

BWojtowicz-WMF created T432933: gpt-oss-safeguard-20b fails with 500 when output cannot be parsed as Harmony messages.
Thu, Jul 23, 8:30 AM · Patch-For-Review, Machine-Learning-Team (Q1 FY2026-27)

Tue, Jul 21

BWojtowicz-WMF claimed T432128: Create dataset of article titles for TTS V1.
Tue, Jul 21, 2:19 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF moved T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Tue, Jul 21, 12:30 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x as Resolved.
Tue, Jul 21, 12:29 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x.

Closing this for now. The underlying issue of corrupted responses on distributed inference is not fully solved, but we agreed to not do distributed inference on partitioned MI300X GPUs in line with AMD docs.

Tue, Jul 21, 12:23 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF moved T431978: Enable FP8 KV Cache for Qwen36-27b from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Tue, Jul 21, 12:15 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T431978: Enable FP8 KV Cache for Qwen36-27b as Resolved.
Tue, Jul 21, 12:15 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T431978: Enable FP8 KV Cache for Qwen36-27b, a subtask of T431846: Q1 FY2026-27 Milestone 1: Deliver a V0 LLM platform for Wikimania, as Resolved.
Tue, Jul 21, 12:15 PM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T403254: Article topic cache backfilling.

The idea of warming up LAC this way via MW JobQueue sounds great to me!

Tue, Jul 21, 12:09 PM · Machine-Learning-Team

Mon, Jul 20

BWojtowicz-WMF added a comment to T403254: Article topic cache backfilling.

We can re-use this task to decide if we want to do backfilling and how to perform it if yes :)

Mon, Jul 20, 8:59 AM · Machine-Learning-Team
BWojtowicz-WMF renamed T403254: Article topic cache backfilling from Article topic cache backfilling using article_topic hive table to Article topic cache backfilling.
Mon, Jul 20, 8:38 AM · Machine-Learning-Team

Jul 13 2026

BWojtowicz-WMF created T431978: Enable FP8 KV Cache for Qwen36-27b.
Jul 13 2026, 8:56 AM · Machine-Learning-Team (Q1 FY2026-27)

Jul 10 2026

BWojtowicz-WMF moved T425680: Host Qwen 3.6-27B as an inference service from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Jul 10 2026, 9:32 AM · Traffic, Lift-Wing, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF closed T425680: Host Qwen 3.6-27B as an inference service as Resolved.
Jul 10 2026, 9:32 AM · Traffic, Lift-Wing, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T425680: Host Qwen 3.6-27B as an inference service.

The model is now deployed in the llm namespace and is exposed via REST Gateway with LiftWingLLM rate limits. Additionally, I've added it to LiftWing studio so everyone can play around with it via UI.
Model can be queried like below:

curl -s https://api.wikimedia.org/service/lw/inference/v1/models/llm-qwen36-27b/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"llm-qwen36-27b","messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":200}' | jq .
Jul 10 2026, 9:30 AM · Traffic, Lift-Wing, Machine-Learning-Team (Q1 FY2026-27)

Jul 9 2026

BWojtowicz-WMF added a comment to T425680: Host Qwen 3.6-27B as an inference service.

I removed the DTYPE=float16 and exchanged it with DTYPE=auto env var and service still responds happily, which confirms that the FP8 quantization works.

Jul 9 2026, 1:29 PM · Traffic, Lift-Wing, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T431089: Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction.

I've done first batch of updates targeting our LLM models, each of them now is now defining both requests and limits for ephemeral storage. All models are redeployed.

Jul 9 2026, 12:52 PM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work

Jul 8 2026

BWojtowicz-WMF added a comment to T431580: httpbb tests from `test_article-descriptions.yaml` fail.

Regarding timeouts, our service definitely takes _a lot_ of time to respond (with success!), which results in the 10 second timeout:

INFO:root:Opening a new Asyncio session for restgateway.
2026-07-08 15:34:46.402 kserve.trace requestId: a6e43b78-520a-4e53-9e6d-a2f428de93f3, preprocess_ms: 311.614513397, explain_ms: 0, predict_ms: 11384.052991867, postprocess_ms: 0.011920929
2026-07-08 15:34:46.402 uvicorn.access INFO:     127.0.0.6:0 1 - "POST /v1/models/article-descriptions%3Apredict HTTP/1.1" 200 OK
2026-07-08 15:34:46.403 kserve.trace kserve.io.kserve.protocol.rest.v1_endpoints.predict: 11.697105884552002
2026-07-08 15:34:46.403 kserve.trace kserve.io.kserve.protocol.rest.v1_endpoints.predict: 22.731085999999777
2026-07-08 15:34:49.880 uvicorn.access INFO:     127.0.0.6:43489 1 - "GET /metrics HTTP/1.1" 200 OK
2026-07-08 15:34:49.880 kserve.trace kserve.io.kserve.protocol.rest.server.metrics_handler: 0.0014925003051757812
2026-07-08 15:34:49.881 kserve.trace kserve.io.kserve.protocol.rest.server.metrics_handler: 0.001480999999330379
2026-07-08 15:35:02.738 uvicorn.access INFO:     127.0.0.6:60479 1 - "GET /metrics HTTP/1.1" 200 OK
2026-07-08 15:35:02.738 kserve.trace kserve.io.kserve.protocol.rest.server.metrics_handler: 0.001093149185180664
2026-07-08 15:35:02.738 kserve.trace kserve.io.kserve.protocol.rest.server.metrics_handler: 0.0010870000005525071
INFO:root:Opening a new Asyncio session for restgateway.
2026-07-08 15:35:27.365 kserve.trace requestId: b0ae7ad0-1133-4d81-8e1c-a6f65d322cd3, preprocess_ms: 242.444753647, explain_ms: 0, predict_ms: 11316.508769989, postprocess_ms: 0.016212463
Jul 8 2026, 3:43 PM · Machine-Learning-Team
BWojtowicz-WMF added a comment to T425680: Host Qwen 3.6-27B as an inference service.

We've noticed similar issue with corrupt responses when deploying qwen3-14b model on 2x 24GB VRAM partitions of MI300X - https://phabricator.wikimedia.org/T431268.

Jul 8 2026, 11:59 AM · Traffic, Lift-Wing, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T424923: Support GPU Partition specification in deployments.

Thank you for your inputs @achou @isarantopoulos! It seems we are all agreeing on Option 1, which means tainting all current GPU nodes.

Jul 8 2026, 11:33 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Essential-Work, Machine-Learning-Team (Q1 FY2026-27)

Jul 6 2026

BWojtowicz-WMF added a comment to T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x.

Moved the experimental qwen3-14b deployment to a single 192GB GPU on ml-serve1012 and everything works fine. Same change for llm-qwen3-14b is on the way.

Jul 6 2026, 11:06 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x.

One potential issue I've noticed in startup logs is this:

(EngineCore_DP0 pid=186) [aiter] WARNING: NUMA balancing is enabled, which may cause errors. It is recommended to disable NUMA balancing by running "sudo sh -c 'echo 0 > /proc/sys/kernel/numa_balancing'" for more details: https://rocm.docs.amd.com/en/latest/how-to/system-optimization/mi300x.html#disable-numa-auto-balancing
(EngineCore_DP0 pid=186) [2026-07-06 09:02:14] WARNING core.py:454: WARNING: NUMA balancing is enabled, which may cause errors. It is recommended to disable NUMA balancing by running "sudo sh -c 'echo 0 > /proc/sys/kernel/numa_balancing'" for more details: https://rocm.docs.amd.com/en/latest/how-to/system-optimization/mi300x.html#disable-numa-auto-balancing
Jul 6 2026, 10:37 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF claimed T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x.
Jul 6 2026, 10:29 AM · Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF created T431268: Qwen3-14b responses are corrupted when running on 2x24GB partitions of MI300x.
Jul 6 2026, 10:28 AM · Machine-Learning-Team (Q1 FY2026-27)

Jul 2 2026

BWojtowicz-WMF closed T426749: Access control for LiftWing LLM services exposed to external clients through REST Gateway as Resolved.
Jul 2 2026, 11:53 AM · Machine-Learning-Team, ServiceOps-SharedInfra, ServiceOps
BWojtowicz-WMF added a comment to T426749: Access control for LiftWing LLM services exposed to external clients through REST Gateway.

The changes to the gateway (https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1305621) are merged and deployed. We've settled on this rate limiting policy:

"LiftWingLLM":
  shadow_mode: false
  limits:
    "*": # strict shared limit for public access (anon*, unauthed*, authed-*)
      HOUR: 100
    "known-network": # network under our control: WMCS / Toolforge (x-trusted-request: A)
      HOUR: 9999999
    "known-client": # network associated with a known client (x-trusted-request: B)
      HOUR: 9999999
    "approved-bot": # community approved bot, based on jwt auth
      HOUR: 9999999
Jul 2 2026, 11:52 AM · Machine-Learning-Team, ServiceOps-SharedInfra, ServiceOps

Jun 25 2026

BWojtowicz-WMF added a comment to T426749: Access control for LiftWing LLM services exposed to external clients through REST Gateway.

I've drafted the REST Gateway changes in here https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1305621.
I went for the llm-* path matching routing to llm namespace. So all services matching the llm-* regex inside llm namespace would be automatically exposed with the LiftWingLLM rate limiting policy. I included conservative 100 requests per hour limit for public traffic according to the discussion above.

Jun 25 2026, 10:46 AM · Machine-Learning-Team, ServiceOps-SharedInfra, ServiceOps

Jun 24 2026

BWojtowicz-WMF updated subscribers of T426749: Access control for LiftWing LLM services exposed to external clients through REST Gateway.

I've looked into the current rate limiting setup for LiftWing and have questions and ideas on how we could approach it. We'd love to push this work forwards so I'd love to hear inputs from ServiceOps and MW-Interfaces.
Pinging @daniel and @Clement_Goubert I got a suggestion to include you here :)

Jun 24 2026, 9:14 AM · Machine-Learning-Team, ServiceOps-SharedInfra, ServiceOps

Jun 22 2026

BWojtowicz-WMF added a comment to T424923: Support GPU Partition specification in deployments.

First, let me give some context on options I'll share below - taints/tolerations and nodeAffinity do opposite jobs: a taint repels everything that does not explicitly tolerate it, while nodeAffinity only attracts a pod when it is present, but does not do any "blocking" during scheduling. The practical consequence is that if we want hard guarantees about which GPUs we land on, it has to come from taints as affinities cannot prevent accidental schedules.

Jun 22 2026, 9:10 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Essential-Work, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T424923: Support GPU Partition specification in deployments.

Current state.

Jun 22 2026, 8:39 AM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Essential-Work, Machine-Learning-Team (Q1 FY2026-27)

Jun 17 2026

BWojtowicz-WMF moved T428127: Improve and optimize Article Topics model to better handle queries with `revision_id` from In Progress to Done on the Machine-Learning-Team (Q1 FY2026-27) board.
Jun 17 2026, 8:00 AM · Machine-Learning-Team, OKR-Work
BWojtowicz-WMF closed T428127: Improve and optimize Article Topics model to better handle queries with `revision_id`, a subtask of T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review, as Resolved.
Jun 17 2026, 8:00 AM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal
BWojtowicz-WMF closed T428127: Improve and optimize Article Topics model to better handle queries with `revision_id` as Resolved.
Jun 17 2026, 8:00 AM · Machine-Learning-Team, OKR-Work
BWojtowicz-WMF added a comment to T428127: Improve and optimize Article Topics model to better handle queries with `revision_id`.

Update
Two patches landed to improve our MW API call queries within Article Topics

  • First change, which was related to passing invalid revision_id usually linked to deleted/renamed/moved pages. This previously was throwing 500 error without explanation, but now throws a nice 400 with explanation.
  • Second change added retries on MW API calls, which failed for transient reasons. We are doing a lot of MW API calls for single request and this should allow us to have lower error rate for Article Topic requests.
Jun 17 2026, 7:59 AM · Machine-Learning-Team, OKR-Work

Jun 5 2026

BWojtowicz-WMF added a comment to T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review.

Status Update

Jun 5 2026, 10:05 AM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal

Jun 4 2026

BWojtowicz-WMF closed T418493: Integrate Article Topic model with the new caching service, a subtask of T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review, as Resolved.
Jun 4 2026, 7:50 AM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal
BWojtowicz-WMF closed T418493: Integrate Article Topic model with the new caching service as Resolved.
Jun 4 2026, 7:50 AM · Machine-Learning-Team, OKR-Work
BWojtowicz-WMF added a comment to T418493: Integrate Article Topic model with the new caching service.

We've turned on the integration on production as well 🎉
Example of how to query Hoarde is below, it is only reachable internally at the moment so you need to be e.g. on a stat host.

# Production:
curl -D - https://linked-artifacts.discovery.wmnet:30443/revisions/v1/article_topics/enwiki/39755715/1235690033; echo
# Staging:
curl -D - https://linked-artifacts.k8s-staging.discovery.wmnet:30443/revisions/v1/article_topics/enwiki/39755715/1235690033; echo
Jun 4 2026, 7:47 AM · Machine-Learning-Team, OKR-Work
BWojtowicz-WMF created T428127: Improve and optimize Article Topics model to better handle queries with `revision_id`.
Jun 4 2026, 7:47 AM · Machine-Learning-Team, OKR-Work
BWojtowicz-WMF added a comment to T418493: Integrate Article Topic model with the new caching service.

I did some more digging on the 500s.
Big chunk of those errors stems from the fact that I was using ~year old revision_id, which later broke our MW API calls (using action=parse&oldid=... call), because some of those pages have been deleted/suppressed/renamed. Such calls make MW API return an error ( nusuchrevid / missingtitle ), which we currently surface as generic 500.
Good news is that in practice we would be usually using fresh revision_ids (e.g. by page_change events), which would mean that error rate would be lower. However, using old revision_ids is still a legitimate use-case, which we should be able to handle.

Jun 4 2026, 7:22 AM · Machine-Learning-Team, OKR-Work

Jun 3 2026

BWojtowicz-WMF added a comment to T418493: Integrate Article Topic model with the new caching service.

Status update

Jun 3 2026, 8:46 AM · Machine-Learning-Team, OKR-Work

May 29 2026

BWojtowicz-WMF added a comment to T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review.

Status Update

May 29 2026, 6:42 AM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal

May 26 2026

BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

We confirmed that gRPC endpoints works via standard 30443 port on production server without the LVS changes. Glad that we found that out, thank you @JMeybohm and @elukey!

May 26 2026, 1:57 PM · Machine-Learning-Team, ServiceOps, Traffic
BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

@elukey
Thanks, I indeed missed it! Initially I thought that 2nd gateway might be a way to overcome the Knatives auto-managed Gateway single-port limitation, but now I see that it is not the case.
I indeed just tested the script above, but pointing at the 30443 port directly instead of the new 30051 port and it succeeded, which proves that the LVS changes were indeed not needed(?).

May 26 2026, 7:05 AM · Machine-Learning-Team, ServiceOps, Traffic
BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

Thanks to the changes to LVS, I was successful with testing the gRPC connection on staging with the script below! 🎉
We will now do the gRPC integration with Hoarde on staging and if everything goes smooth, we can merge similar LVS changes to production. Thank you for all the work here!

May 26 2026, 6:37 AM · Machine-Learning-Team, ServiceOps, Traffic

May 6 2026

BWojtowicz-WMF added a comment to T419288: Q4 FY2025-26 Goal: Text-to-Speech V0 Prototype.

I'm sharing a draft of solution architecture we discussed in our ML Team Meeting for this problem. The approach is similar to the prototype solution developed by Kevin.
It could accommodate both batch jobs and real-time updates. End users would communicate with the backend via /generate endpoint to enqueue new TTS jobs or via /audio backend to retrieve path to the generated TTS files.

May 6 2026, 1:40 PM · Machine-Learning-Team, Patch-For-Review, Goal, Research

Apr 30 2026

BWojtowicz-WMF claimed T424923: Support GPU Partition specification in deployments.
Apr 30 2026, 2:55 PM · Data-Platform-SRE (2026-08-07 - 2026-08-28), Essential-Work, Machine-Learning-Team (Q1 FY2026-27)
BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

Happy update!

Apr 30 2026, 10:00 AM · Machine-Learning-Team, ServiceOps, Traffic

Apr 29 2026

BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

Update on current state of things.

Apr 29 2026, 2:59 PM · Machine-Learning-Team, ServiceOps, Traffic

Apr 28 2026

BWojtowicz-WMF added a comment to T424049: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC).

Small status update from debugging efforts.

Apr 28 2026, 12:16 PM · Machine-Learning-Team, ServiceOps, Traffic

Apr 15 2026

BWojtowicz-WMF added a comment to T421903: Investigate enabling gRPC in LiftWing model servers.

To speak on enabling gRPC for ISVC, our plan would be to use the Kserve's V2 Inference Protocol, which supports both gRPC and HTTP/REST interfaces. Currently, all our services were built with V1 protocol in mind, which only supports HTTP/REST interface.

Apr 15 2026, 7:33 AM · Machine-Learning-Team, Lift-Wing

Mar 27 2026

BWojtowicz-WMF added a comment to T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review.

Weekly Update

Mar 27 2026, 1:22 PM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal
BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

What are you doing with the threshold argument? Are you late filtering the response from the inference service, or invoking the service with a threshold as the constraint? If the latter, is there any reason you couldn't late filter a cached response (i.e. is the cached response somehow constrained to a limited set of thresholds)?

Mar 27 2026, 1:11 PM · User-Eevans, Data-Persistence

Mar 26 2026

BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

Okay, I've done a few not too technical sketches trying to visualize the issue we're facing.

Mar 26 2026, 2:14 PM · User-Eevans, Data-Persistence

Mar 25 2026

BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

@Joe Hoarde's HTTP API only exposes wiki_id/page_id/revision_id parameters, which would cover the use-case for the Mobile Apps team. However, our service also exposes additional parameters (e.g. page_title, threshold) that some users rely on. On top of that, exposing HTTP API is extremely useful for us for development/debugging.
I think those would not be as problematic if we were building a new service with hoarde in mind from the beginning, however we're trying to integrate caching into existing services.

Mar 25 2026, 3:02 PM · User-Eevans, Data-Persistence
BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

I want to share a small update from our side on where we are.

Mar 25 2026, 12:38 PM · User-Eevans, Data-Persistence

Mar 24 2026

BWojtowicz-WMF added a comment to T420931: Load test current state of the Article Topic service.

I see the regime with >10s p99 latencies, however it happened during the night and not during running those tests. It seems to me that the Grafana numbers aligns well with the reported latencies above see:

  1. page_id + lang requests: https://grafana.wikimedia.org/goto/cfgzhd4aveg3kf?orgId=1
  2. page_title + lang requests: https://grafana.wikimedia.org/goto/ffgzhfegn63uoc?orgId=1
  3. page_id + lang + revision_id requests: https://grafana.wikimedia.org/goto/bfgzhglvqmpdse?orgId=1
Mar 24 2026, 1:06 PM · OKR-Work, Machine-Learning-Team
BWojtowicz-WMF added a comment to T416475: Unify and improve load testing strategy for inference services.

When investigating T420931, I found that my custom async load test script achieves >300 RPS against the same service with 5 replicas, whereas the locust test against 1 replica reports only ~0.67 RPS. The discrepancy comes down to the Locust configuration:

Mar 24 2026, 9:36 AM · Machine-Learning-Team (Q1 FY2026-27), Essential-Work
BWojtowicz-WMF added a comment to T420931: Load test current state of the Article Topic service.

I'm sharing load test numbers tested against production deployment on eqiad using internal endpoint. I've made sure the responses return valid predictions and I ran the load test after a few hours of cooldown to make results are not skewed by caching on the MWAPI side.

Mar 24 2026, 6:54 AM · OKR-Work, Machine-Learning-Team

Mar 23 2026

BWojtowicz-WMF added a comment to T420931: Load test current state of the Article Topic service.

@Isaac The details of the cache and how exactly will it be implemented to Article Topics is still not fully decided. Current approaches we explored would work with page_id, whereas page_title requests would not go through cache. This ticket does not take cache into consideration, but we're verifying how fast can we get without cache. As a bonus, I can also check the page_title variant in this ticket so we'll have more context on it :)

Mar 23 2026, 2:36 PM · OKR-Work, Machine-Learning-Team
BWojtowicz-WMF created T420931: Load test current state of the Article Topic service.
Mar 23 2026, 2:06 PM · OKR-Work, Machine-Learning-Team

Mar 17 2026

BWojtowicz-WMF added a comment to T418832: Deploy CoPE-A on LiftWing.

After lowering the maximum input token length to 4096, we seem to be able to process all incoming requests. I will figure out optimizations we could make to allow bigger input lengths, but the current 4096 token limit should already be good enough for testing our policies.

Mar 17 2026, 9:06 AM · Patch-For-Review, Product Safety and Integrity, Machine-Learning-Team

Mar 16 2026

BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

If I understand you correctly (and if I don't, please don't hesitate to correct me), you're arguing that we might have uses that can't be satisfied, which would force a product team to build an HTTP API to serve them, one that would otherwise have also worked as the lambda (while providing an example of a hypothetical use-case). Or put another way, that (a, above) we might have past use cases with extant HTTP APIs, and (b) we might have (unavoidable) future ones too.

Mar 16 2026, 3:11 PM · User-Eevans, Data-Persistence
BWojtowicz-WMF added a comment to T418832: Deploy CoPE-A on LiftWing.

After deployment, CoPE-A-9B model server was successfully processing small requests of less than 500 input tokens.

Mar 16 2026, 11:49 AM · Patch-For-Review, Product Safety and Integrity, Machine-Learning-Team
BWojtowicz-WMF added a comment to T418832: Deploy CoPE-A on LiftWing.

The CoPE-A-9B model is now deployed on LiftWing.

Mar 16 2026, 10:54 AM · Patch-For-Review, Product Safety and Integrity, Machine-Learning-Team

Mar 13 2026

BWojtowicz-WMF added a comment to T392833: Q1 FY2025-26 Goal: Make article topic data available at scale and within SLOs for Year in Review.

Weekly Update

Mar 13 2026, 2:37 PM · Machine-Learning-Team, Patch-For-Review, OKR-Work, Goal
BWojtowicz-WMF added a comment to T419734: RfC: Use of gRPC as Lambda interface for linked artifact caching.

That said, there's a practical concern on the ML Platform side and other non-ML services willing to integrate. The vast majority of our internal services communicate over HTTP, and our infra (mesh, ingress, routing) is built around that. To integrate with gRPC-only Hoarde, each service team would need to either deploy an adapter/proxy alongside their service, or deploy a gRPC-capable replica. I've prototyped the adapter approach and discussed the replica route with Luca - both are technically feasible. But as Hoarde would onboard more use-cases, I can imagine this becoming a pattern where every integrating service carries extra infrastructure just to bridge the protocol gap. So we’re shifting complexity and maintenance from Hoarde to its clients.

To me the above is the main concern, since almost every service at the foundation runs HTTP..

Is this because you a) envision the service being used to do caching for already implemented systems, b) as-yet-implemented systems that will invariably need an HTTP service anyway (and if so, why), or c) because using something other than HTTP generally imposes a burden that exceeds the benefits of using grpc (and if so, how)?

Mar 13 2026, 9:03 AM · User-Eevans, Data-Persistence

Mar 12 2026

BWojtowicz-WMF added a comment to T418832: Deploy CoPE-A on LiftWing.

Small update on the progress.

Mar 12 2026, 8:50 AM · Patch-For-Review, Product Safety and Integrity, Machine-Learning-Team

Mar 11 2026

BWojtowicz-WMF moved T400602: Investigate reference-need persistently unavailable replicas alert from Ready To Go to 2025-2026 Q2 Done on the Machine-Learning-Team board.
Mar 11 2026, 12:51 PM · Essential-Work, Machine-Learning-Team
BWojtowicz-WMF closed T400602: Investigate reference-need persistently unavailable replicas alert as Resolved.
Mar 11 2026, 12:50 PM · Essential-Work, Machine-Learning-Team
BWojtowicz-WMF added a comment to T400602: Investigate reference-need persistently unavailable replicas alert.

Resolving this as this was a single time incident and the underlying concern about reference-need's high resource requests (22 CPUs, 6Gi memory) and its impact on cluster scheduling is now tracked as part of T414431, where we are optimizing resource utilization across all ISVCs.

Mar 11 2026, 12:50 PM · Essential-Work, Machine-Learning-Team
BWojtowicz-WMF moved T417860: Explore gpt-oss-safeguard-20b from Unsorted to 2025-2026 Q2 Done on the Machine-Learning-Team board.
Mar 11 2026, 10:34 AM · Machine-Learning-Team
BWojtowicz-WMF closed T417860: Explore gpt-oss-safeguard-20b, a subtask of T418267: Q2 FY2025-26 Goal: Host a content policy evaluation model on LiftWing, as Resolved.
Mar 11 2026, 10:34 AM · Machine-Learning-Team, Product Safety and Integrity, WE4.12 Content policy model evaluation, Goal
BWojtowicz-WMF closed T417860: Explore gpt-oss-safeguard-20b as Resolved.
Mar 11 2026, 10:34 AM · Machine-Learning-Team