User Details
- User Since
- Aug 23 2016, 11:49 PM (519 w, 3 d)
- Availability
- Available
- LDAP User
- Unknown
- MediaWiki User
- Pfps [ Global Accounts ]
Thu, Aug 6
Thu, Jul 30
Tue, Jul 21
The claim should be made more specific. My understanding is that the QLever SPARQL engine does not have built-in prefixes. But the UI at qlever.dev does automatically include expansions for common prefixes in the query as you type them. So QLever does have tooling to manage prefixes, but different tooling from Blazegraph.
Mon, Jul 20
The "why" seems obvious - the WDQS label service is a non-standard construct that violates the assumptions built into the SPARQL SERVICE construct.
Wed, Jul 15
See https://www.mediawiki.org/wiki/Wikibase/Indexing/RDF_Dump_Format#WDQS_data_differences for more information.
schema:name triples are stripped out of the RDF dump because they are redundant.
Mon, Jul 13
The issue here is, I think, whether there will be a UI feature that makes it easy to add labels, etc., to query results. At one time the QLever UI had a feature where you could just click a button and labels would be added to the query results. This wasn't as general as the WDQS label service, but it was very simple to use, which is a prime characteristic for the UI.
Jun 18 2026
You might want to look at
May 26 2026
GeoSPARQL is a standard, but not in the SPARQL 1.1 standard. Is there more geo here than in GeoSPARQL?
Is this problem being worked on? If the graph split is going to continue, it should be done right.
May 13 2026
Here is a change that I did to check out that there were no flags on this kind of edit:
May 6 2026
May 5 2026
Are there any plans to rationale the constraint mechanisms? I was on a call where we talked about class-based constraints, which are not possible in Wikidata but some of the constraints on P31 would better be based on classes.
After clicking around for a while
https://gerrit.wikimedia.org/g/mediawiki/extensions/WikibaseQualityConstraints/+/20697c11216ba1b98989095f1d6673335416df64/README.md gives good information about how to set up the system. But it is missing an overview of what is going on. Something that has "WBQC adds constraint checks that are invoked ... and can ...." and "The checks are invoked ... and do their work via ... and produce results by ..."
May 4 2026
Where is a good place to go to find actual documentation on the property constraint system?
Apr 27 2026
And remove Blazegraph query hints.
Apr 26 2026
And named subqueries.
Apr 24 2026
This use case is further described in https://phabricator.wikimedia.org/T423425
One question that comes up is whether some change to Wikidata will break software that accesses Wikidata using SPARQL queries. If more information was known about these usages, it would be easier to improve Wikidata.
Apr 15 2026
The accesses should be appropriately anonymized before being made available, of course.
Apr 8 2026
Mar 29 2026
If your name on the GSoC site is different from your ID in Phabricator you need to say what your GSoC ID is. Otherwise your interaction here will not be associated with your proposal and it will be rejected.
A reminder that you need to submit your proposals to Google for them to be considered, in addition to anything done in Phabricator.
Mar 28 2026
This project is medium difficulty and large (350 hours) in scale.
This project is medium difficulty and large size.
Mar 27 2026
Hi Meghana: If you email me at pfpschneider@gmail.com we can set up a meeting in the afternoon, US East Coast time.
Hi Arina: By this time you should have been interacting with us, the potential mentors, for quite some time and have done some of the microtasks. We will evaluate your proposal even so, but iteration at this late date is not likely.
Mar 26 2026
That's good. The next determination is whether to do a complete comparison or an incomplete one. Then there is the issue of whether to include a third or fourth engine so that compliance can be estimated.
Mar 25 2026
I don't think that the problem is solvable in general. You can't even rerun queries and always expect the same results because of updates.
Mar 23 2026
Hi Meghana: The deadline for proposals is next week so you are starting very late in the process. By now you should have tried out some of the suggestions in the initial comment. If you want to proceed we can set up a call as described in my earlier comment.
Mar 19 2026
@Olea That's interesting (to me). Feel free to contact me at pfpschneider@gmail.com
Mar 18 2026
I'm not looking for a daily or even weekly timeline in your proposal. What I am looking for is a breakdown of the overall task into a few pieces that can be tracked. I am also looking for one or more pieces that can be done by the midterm of the coding period. If your proposal is selected we will be using the familiarization period to further refine the work to be done.
As far as I am concerned, the minimal deliverable at the end of the project is a playable game that implements fixes to some kinds of constraint violations. It is in your interest to have a proposal that can support this minimal deliverable, and also has significant optional parts.
The guidance from Google in https://google.github.io/gsocguides/student/writing-a-proposal is to have deliverables and timelines in your proposals. This is a good idea, but it is possible to go too far in this area. What we are looking for is a sense that you can break down the overall project into several pieces, each probably including design, coding, and documentation. What is important is a plan to have something that can be evaluated at the midpoint. That doesn't have to be a full system, but there should be some coding involved.
This project has areas where you can decide how much or how little to do and still have a working result. That may make it different from other GSOC projects. It is a good idea to make pieces of your proposal optional, so that at the end you (and we) can claim victory even if everything in the proposal is not implemented.
In the end, a big part of GSOC is to get people interested in open-source projects. In my view it's a win for GSOC if you end up doing significant open-source work in the future, even if not all of your proposal ends up being implemented.
Mar 17 2026
Sorry all. Due to some other issues, I have not been adequately responsive. But there is light at the end of the tunnel. I'll get through all the backlog today.
Mar 11 2026
I can check triple counts on my benchmark machine when the current benchmark run finishes, which may take another day or so.
Also, I never used munged files with QLever. I don't know whether that would slow down or speed up QLever.
Mar 10 2026
The parsing part of the QLever ingestion process can be heavily parallelized if the input file or files are well-behaved. Because of this need for well-behaved input files a special flag must be set. The RDF dump files are well-behaved.
https://www.wikidata.org/wiki/Wikidata:Scaling_Wikidata/Benchmarking/Virtuoso#Run_the_Virtuoso_server_and_load_the_Wikidata_files_into_the_server already notes the slowdown in ingestion rate with Virtuoso as the process proceeds, and postulates a cause.
Mar 9 2026
Mar 6 2026
The reason that opened this ticket is that the team has been poor at responding to comments in the Wikidata wiki. So I looked in https://www.mediawiki.org/wiki/Wikidata_Platform#How_to_contact_us and the only contact methods there say to open Phabricator tickets. So I did.
Mar 5 2026
So who can I ask to see the interaction with the Privacy and Security team?
One problem I have here is that I haven't seen any of the interaction with the privacy and security people so I don't know what their requirements for anonymization are.
My belief is that this is a result of slow processing of transitive closure operations in Blazegraph and is likely to not be a problem in optimized SPARQL engines.
I just learned that the initial work to anonymize the queries was supported by the Wikimedia Foundation, https://meta.wikimedia.org/wiki/Research:Understanding_Wikidata_Queries
I am disappointed in this abrupt ending, particularly after nearly four months.
Mar 4 2026
This is the difference between join and left join. Regular joins (butting two triple patterns together) are associative and commutative, so the join order doesn't matter. Left joins (OPTIONAL) are associative (I think) but not commutative, so the join order does matter. So the order of the OPTIONALs matters. For official information you need to dig deeply into https://www.w3.org/TR/sparql11-query/. To extract information from that document sometimes requires a good grounding in the theory of querying.
Mar 3 2026
@Harikrshnaa I'll take a look this week.
Mar 2 2026
For a more realistic example, try https://qlever.dev/wikidata/xAiP0y then https://w.wiki/J5kE then https://w.wiki/J5kT then https://w.wiki/J5kH
The SCHOLIA replacement also works better if the label variable already has a value.
SPARQL defines what the result of evaluation is, just like in a programming language. Implementations are free to do anything so long as the specified result is reached. The above SPARQL must produce "only" de labels if there are any, according to the definition of SPARQL.
Moreover, the SCHOLIA work was already reported to the team in https://phabricator.wikimedia.org/T414453
It would be worthwhile for you to have more knowledge of the community activities in this area. Your team knows that the SCHOLIA queries have been rewritten into QLever so your team should have either reached out to them to find out what they did or looked at the changes that they made. There you would have seen a better transformation which instead of several OPTIONALs on different variables followed by a BIND uses a sequence of OPTIONALs on the same variable. Further, in some SCHOLIA queries there is an initial BIND that eliminates the problems that happen if the variable is unbound.
Mar 1 2026
I strongly suggest reaching out to the community to find out the current best practices for replacing the label service.
It appears from https://gitlab.wikimedia.org/repos/wikidata-platform/wdqs-query-proxy/-/blob/main/src/test/java/org/wikimedia/wdqs/WikibaseLabelParserTest.java?ref_type=heads that the replacement being evaluated is much less than ideal.
It would be useful to also rewrite named subqueries, as other SPARQL systems may not handle them either. And also remove Blazegraph optimization hints, which will interfere with results in other systems.
Is there a specification of exactly what the label service does?
Feb 27 2026
Feb 25 2026
One more thing that would be useful, if possible, is whether the query was syntactically legal according to Blazegraph. I could get this information by running the query through Blazegraph, but if this information is in the log I could use that instead.
Feb 24 2026
I want to mock up a server to investigate its load, so all I need for that is the anonymized query and a relative (or absolute) timestamp. User agent categories could be useful to better estimate future loads. Anonymizing string literals will be a problem for me, but I understand if this has to be done.
Looking through the code there appears to be use of SPARQL (for at least distinct values). Is this the main place where Wikidata itself depends on the WDQS? How does the split of the WDQS factor into the use of SPARQL? Does this mean that the distinct values constraint can't work on scholarly articles (or, actually, at all if it misses values from the scholarly graph)?
So the feature includes complex constraints?
Is there a good description of how the constraint system is implemented, preferably including the role of third-party tools?
Given that KrBot is a third-party closed-source (?) tool, I would be happy having it replaced.
Feb 23 2026
What is the difference, if any between the "quality constraints feature" and Wikidata property constraints? Having some notion of what is being investigated would be useful if the community is to provide useful input into the process.
I'm confused as to just what this ticket is about.
I just noticed the paper https://arxiv.org/abs/2602.14594
Feb 18 2026
It turns out that fallback functionality has become more important with the introduction of the mul "language". If I want English labels I need to have mull as a fallback or I may miss many labels.
Feb 17 2026
Thanks.
Feb 15 2026
It is very frustrating to have this task languish without any way of contacting the team that appears to be blocking progress.
The label service is indeed pervasive. How much are the mwapi services used?
Jan 22 2026
Jan 21 2026
Jan 20 2026
@Hridyesh_Gupta Thank you for your interest. It might be a bit early to start on the microtasks, as the period where potential contributors interact with mentors isn't for a while. As well, the topic will be getting some updates over the next little while.
Jan 14 2026
How can this be escalated?