In order to test triple stores, we need main and scholarly graphs available on test nodes (wdqs1029, wdqs1030, wdqs1031, wdqs1032).
Could you initiate the transfer for us?
As a follow-up, I’d like us to start thinking about automated ways to move datasets generated by the batch pipelines.
It would be great if we could use storage that supports the S3 protocol, such as Ceph. The database backends we’re considering to replace Blazegraph have fast enough data ingestion to support reconciliation in a lambda-architecture style (links to Google Doc). This is a path I’d like to explore as we redesign the WDQS architecture.
We are thinking of the following setup:
wdqs1029 -> qlever main graph
wdqs1030 -> qlever scholarly graph
wdqs1031 -> virtuoso main graph
wdqs1032 -> virtuoso scholarly graph
wdqs1028 will be kept as full graph node, with qlever and virtuoso in multi-tenant that we can use for experimentation/break things at will.
We just need the datasets, and will take care of indexing ourselves.
AC
- main and scholarly triples datasets are available on test eqiad nodes.