Page MenuHomePhabricator

Additional capacity on the k8s Flink cluster for WCQS updater
Closed, ResolvedPublic2 Estimated Story Points

Description

As an operator of WCQS I want to ensure that there is enough capacity in our k8s cluster so that deploying WCQS Flink Updater isn't blocked.

For context, WCQS is a service similar to WDQS but consuming data from Commons. The WCQS Updater is similar to the WDQS Updater, but with a different source of data (Commons instead of Wikidata) and an expected lower rate of edits.

Note that with our current understanding of the resource requirement for WDQS Updater, we probably over provisioned the WDQS Updater. We might be able to reuse this excess capacity.

Current resources consumption for the WDQS updater running in Yarn: 13 cores / 15G RAM

AC:

  • estimation of the resources required
  • validation that there is enough capacity or a plan to increase the capacity

Event Timeline

Restricted Application added a subscriber: Aklapper. · View Herald Transcript

The most important parameter to streaming updater is related to storage - we have a huge surplus of computing resources in WDQS streaming updater, since edit rate is much lower in SDoC, there is no risk when it comes to throughput.

The storage needs on the other hand are are derived from the data size itself, best represented as triple count. Currently WDQS sits at ~13B. WCQS is ~2.8T, which means about 21% of storage used by WDQS. Since we planned 2x resources required for WDQS Streaming Updater, we are safe to proceed. It would make sense to add some resources to offset WCQS growth in the near future.

Hey! We are working on Wikimedia Commons Query Service (WCQS - a service similar to WDQS, but with the data from Commons). This will add some load to our Flink Updater. By our estimates, it will be fairly low compared to the WDQS related load and we over provisioned for WDQS, so we don't expect any issues. Could someone in ServiceOps confirm?

Hi. Given the 13 core/15GB RAM requirement, I can verify that we have that capacity free lying around[1], so we expect no problems there.

Is T280485#7275149 related to blazegraph and not flink ? I am not sure what 13B triplets vs 2.8T triples means storage wise and in which context.

[1] https://w.wiki/4Q7J

Is T280485#7275149 related to blazegraph and not flink ? I am not sure what 13B triplets vs 2.8T triples means storage wise and in which context.

The number of triples in Blazegraph is roughly linearly correlated with the local storage requirement on the Flink side.

Is T280485#7275149 related to blazegraph and not flink ? I am not sure what 13B triplets vs 2.8T triples means storage wise and in which context.

The number of triples in Blazegraph is roughly linearly correlated with the local storage requirement on the Flink side.

OK, so we are talking about an increase that's on the order of 200 times. However, I am still not clear though where exactly that change would be depicted. Swift? The fs (I guess /tmp mostly) of the container of the jobmanager or the taskmanager? And do we expect the increase to actually be on the order mentioned above? Or are the internal flink databases more or less efficient than that crude calculation?

Is T280485#7275149 related to blazegraph and not flink ? I am not sure what 13B triplets vs 2.8T triples means storage wise and in which context.

Oh, I now see the confusion! Wrong units (typo) in the initial message. The current Flink updater takes data from Wikidata, which has ~13B triples. The new Flink updater will add support for getting data from Commons, which has ~2.8B triples. So the new updater will add ~20% more resource consumption (assuming a linear cost).

This will mean:

  • additional storage on Swift (I assume this is trivial given the size of Swift and can be ignored)
  • additional CPU / RAM usage on k8s
  • additional local storage (/tmp) on the containers

It isn't super clear to me if our strategy is to increase the size of the current Flink cluster, or have a new cluster dedicated to the Commons updater (to be decided later today).

Duplicate the existing cluster would provide additional isolation between the 2 workflows. This is also the worst case scenario in terms of resource needed. The additional estimated resources are:

  • manager: 1 more pod at 1.6G, cpu: 500m
  • workers: 3 pods at 2.1G ram, cpu: 1000m

Is T280485#7275149 related to blazegraph and not flink ? I am not sure what 13B triplets vs 2.8T triples means storage wise and in which context.

Oh, I now see the confusion! Wrong units (typo) in the initial message. The current Flink updater takes data from Wikidata, which has ~13B triples.

😆😆. OK, thanks for clearing that up. That 200 times increase had me worried.

The new Flink updater will add support for getting data from Commons, which has ~2.8B triples. So the new updater will add ~20% more resource consumption (assuming a linear cost).

OK, that's nothing then.

This will mean:

  • additional storage on Swift (I assume this is trivial given the size of Swift and can be ignored)
  • additional CPU / RAM usage on k8s
  • additional local storage (/tmp) on the containers

We got enough on all of those 3, no worries there.

It isn't super clear to me if our strategy is to increase the size of the current Flink cluster, or have a new cluster dedicated to the Commons updater (to be decided later today).

Cool. Let us know what you decide. On our side, it probably isn't much more than 1 more deployment on the k8s cluster. That being said, and assuming my memory is up to date with how flink works in session cluster mode, I 'd expect it's able to handle this internally without needing another k8s deployment. It's also fine to increase the number of worker pods if that's something that would make things easier for you.

Duplicate the existing cluster would provide additional isolation between the 2 workflows. This is also the worst case scenario in terms of resource needed. The additional estimated resources are:

  • manager: 1 more pod at 1.6G, cpu: 500m
  • workers: 3 pods at 2.1G ram, cpu: 1000m

Even in this case, we got that capacity.

small precision:
If we reuse the same cluster (same k8s namespace):

  • it's 3 more pods at 2.1G ram, cpu: 1000m each

If we use a separate cluster (new k8s namespace):

  • add a pod at 1.6G, cpu: 500m to the 3 pods mentioned above

Sorry all for the confusion my typo caused, different name for that magnitude in my native language is confusing me my whole life :).

small precision:
If we reuse the same cluster (same k8s namespace):

  • it's 3 more pods at 2.1G ram, cpu: 1000m each

If we use a separate cluster (new k8s namespace):

  • add a pod at 1.6G, cpu: 500m to the 3 pods mentioned above

The difference is small enough to make resource allocation irrelevant for this. I 'd suggest you don't try to optimize for CPU/RAM when deciding.

I'd opt for "reuse the same [flink] cluster" from the perspective that we treat this snowflaky-ish in the k8s clusters. So less flink-clusters means less snowflakes (at some point it does become a snowball, right? 😃 ).

Sorry all for the confusion my typo caused, different name for that magnitude in my native language is confusing me my whole life :).

No worries, that's fine. You actually just triggered my encyclopedic inner beast which sent me down a small translation rabbithole and finally into https://pl.wikipedia.org/wiki/Przedrostek_SI and I must say Thanks. Today I learned :)

Btw, the equivalent in Greek is Διεθνές_σύστημα_μονάδων#Προθέματα_μονάδων and yes a similar pattern exists in Tera, Peta and Exa which all are -1 of what the sounds would implicitly suggest to a Greek person. So, I feel you.

Let's go with a single cluster at least for the start. If we hit any real life issue, we can revisit this solution.