We need a split_subgraphs Airflow task that runs
org.wikidata.query.rdf.spark.transform.structureddata.dumps.ScholarlyArticleSplit"
parametrized to read / write in the wikidata namespace.
References
We need a split_subgraphs Airflow task that runs
org.wikidata.query.rdf.spark.transform.structureddata.dumps.ScholarlyArticleSplit"
parametrized to read / write in the wikidata namespace.
References
| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Open | gmodena | T431253 [DE2.4.2] WDQS v2 Scaling | |||
| Stalled | gmodena | T422179 Set up regular bulk data ingestion and indexing pipeline | |||
| Resolved | gmodena | T428237 data-reload: spark processing jobs need to be orchestrated in the wikidata dag. | |||
| Resolved | gmodena | T429252 data-relaod: implement split_subgraphs task |