Page MenuHomePhabricator

[Team Challenges #07E] Wikidata–Wikisource Lexicographical Data Visualization Tool
Open, Needs TriagePublic

Description

Mapping Wikidata lexicographical data within Wikisource text corpora, color-coding words by their lexical category (noun, verb, adjective, etc.) using data from Wikidata lexemes and their forms.

Words with no matching lexeme are left uncolored, making Wikidata lexicographical data gaps visible so the community can create the missing lexemes. In the long run this could also generate statistics on the lexicographical content of Wikisource, and potentially be implemented on-wiki (gadget/user script) so readers don't have to leave Wikisource.

Challenge: #7 — Stream data with Wikidata
Team: #07E Asie | Wikimania 2026 Team Challenges

Scope (Wikimania sprint)

Toolforge-hosted web tool. English Wikisource as test language; architecture should allow other languages later. On-wiki integration (hover pop-ups with lexeme details) and the statistics/analytics page are stretch goals for after Wikimania.

Workflow

  1. User enters an English Wikisource URL (to start and later for all Wikisource language domains)
  2. Tool retrieves the page text and tokenizes it into words
  3. Words are matched against Wikidata lexemes and forms
  4. Lexical categories are color-coded and displayed with a legend
  5. Unmatched words indicate data gaps in Wikidata

Subtasks

Overview:

  • Decide on tool name
  • Create GitLab repo
  • Set up Toolforge tool account
  • Backend: extract and tokenize text from Wikisource (API)
  • Backend: retrieve lexicographical data from Wikidata
  • Frontend: design UI layout (Codex)
  • Frontend: implement color-coded display
  • Test with sample Wikisource pages
  • Documentation
  • Showcase presentation

Links

Event Timeline

Tool name which is proposed to the team is Leximap or Wikisource leximap, since the tool intends to map lexicographical data from Wikisource corpora

Hello @Gnoeee. I hope you are doing well. I am currently working on the backend. Could you add me as a maintainer to the Toolforge project so I can have access. Thanks.

Hello @Gnoeee. I hope you are doing well. I am currently working on the backend. Could you add me as a maintainer to the Toolforge project so I can have access. Thanks.

@Kengkong1 , I have added you there as maintainer

Hello Team. I hope everyone is having a good day. As I work on the backend I realise that some words can have multiple classifications like fast- noun, adjective, adverb
so the front end will have to be able to display multiple options for words. To correct this we would need to do Natural Language Processing which is out of our scope.

Good afternoon team ,

Kengkong1 has identified an important limitation. Some English words belong
to more than one lexical category depending on context. Resolving that
accurately would require Natural Language Processing, which is beyond the
scope of this sprint.

Let's discuss a simple approach.
Me I'm suggesting , what if we approach as displaying all matching
categories and record context aware classification as a future enhancement.

Regards
Nantale Rajat
User :[ BabyJat ] I

Hello Team. I hope everyone is having a good day. As I work on the backend I realise that some words can have multiple classifications like fast- noun, adjective, adverb
so the front end will have to be able to display multiple options for words. To correct this we would need to do Natural Language Processing which is out of our scope.

@Kengkong1 , we can display multiple lexical categories as a mixed color. We donot have to work on NLP as of now. @Dagmawi-M has created a prototype tool, where all kinds of lexical categories can be displayed. Kindly coordinate with him so that both of your workflows and codes can sync.

Hi @Dagmawi-M. I hope your day is going well. I am currently working on the backend right now. Could you use the gitlab repo so that we may align and phabricator so we can track who is doing what task. For the front end I believe the goal is to use the Wikimedia Codex library. Also since our project is using license GPL-3.0, it means that all imported libraries also have to be open-source as well. Thanks.

But yeah the protoype looks pretty cool!

Hi @Kengkong1, hope yours is going well too

yes... I can share the sample code on repo as soon as I get time at the airport, still in a hurry

image.png (1,866×1,352 px, 312 KB)

Hi team the backend functionality is ready. I am unsure if i have time to implement caching but will see if i have time tomorrow. i will be at work tho. if anyone wants to use this functionlatiy you can find it on the dev branch in /composables. and see it implemented on views/home. thanks

Hi everyone! I hope you are all doing well. I really wish I could have joined everyone in Paris, hopefully I can meet everyone next Wikimania. I would humbly like to present a js version of the Leximap tool. I put a lot of hours into it and am quite proud of it.

Pros: Custom colouring of categories per view (if you click on the colour you can change it), ability to see multiple categories for one word, ability to toggle categories, high matching performance
Cons: The longer the article the longer the load (loading will take some seconds), front end is humble, the entire wiki article url has to be pasted in the search bar (though pdf pages are supported), its only built around 'En" right now, limited categories- there appear to be 100+ different lex categories so the filters are limited and hardcoded

I would like to suggest that testers and group members conduct testing on both https://leximap.toolforge.org/ and https://leximap-js.toolforge.org/ [[ https://toolsadmin.wikimedia.org/tools/id/leximap-js | https://toolsadmin.wikimedia.org/tools/id/leximap-js ]]using the same articles for correctness, and other initial criteria to find a solution to demo.

example test links (recent from WikiSourcce home-different sizes):
https://en.wikisource.org/wiki/The_Defendant/In_Defence_of_a_New_Edition
https://en.wikisource.org/wiki/The_Journey,_or_Cross_Roads_to_Conqueror%27s_Castle
https://en.wikisource.org/wiki/Page%3ARicardo_Essay_on_Profits_1815.djvu/23
https://en.wikisource.org/wiki/Narrative_of_Sojourner_Truth,_a_Northern_Slave/Preface
https://en.wikisource.org/wiki/The_Atlantic_Monthly/Volume_70/Number_417/In_a_Japanese_Garden

Thanks everyone!

hi @Kengkong1 @HakanIST

Can you send us your photo as quickly as possible? (Mkidane4@gmail.com is my email)

very urgent

File:Kengkong (Kisenge Anthony Mbaga) Photo.jpg|Kengkong1-Kisenge
Mbaga-Developer

here thanks

hi @Dagmawi-M,
here thanks File:Kengkong (Kisenge Anthony Mbaga) Photo.jpg|Kengkong1-Kisenge Mbaga-Developer

Hello!
I tested the tool and it seems very promising! I already love it

This are my two cents:

image.png (1,598×760 px, 154 KB)

1- It says no scanned source, even when it is backed by scan (La reina de Rapa Nui/Capítulo I at Spanish Wikisource)
2- All the text surrounded by braces is just the header template and some hidden metadata (display:none). It would be nice if hidden elements weren't present for analysis

image.png (204×249 px, 15 KB)

3- It would be nice to display the grammatical features for each lexeme. If I go to the verb page L235668, I find this grammatical features that give context to the lexeme form we're examining: masculino (male, Q499327, grammatical gender), singular (single, Q110786, grammatical number) and participio (participle, Q814722, grammatical mode). This features are essential to understand the possible meaning of the words we're examining! Maybe display them using some icons or abbreviated way, or simply display their labels

Good luck!!