Page MenuHomePhabricator

Proposal: Bulk OCR Improvements
Closed, DeclinedPublic

Description

Profile Information

Name: Gideon Fiadowu Norkplim
User Page: Gideon_Fiadowu_Norkplim
Location: Accra, Ghana.
Typical working hours 3 PM to 8 PM(+0)

Synopsis

  • I propose to implement the missing features for Bulk OCR Improvements in the Wikisource extension (T415145). The project will add a user-friendly interface that lets authorized users (e.g. admins, experienced proofreaders) select multiple pages or an entire book, choose from different OCR engines (Tesseract, Kraken, etc.), review results in a preview/feedback modal, and safely insert the OCRed text into the page text layers.

This will increase the digitization of public-domain books on Wikisource, reduce repetitive manual work for volunteers, improve text quality and accuracy, and encourage more contributors to participate in large-scale proofreading projects. By restricting bulk operations to trusted users and adding proper feedback + error handling, we also maintain wiki safety and data integrity. The result will be a more efficient, accessible, and volunteer-friendly Wikisource digital library.

  • Possible Mentor(s)

Parthiv Menon (@theprotonade), @Satdeep Gill (@SGill)

Deliverables

Wikimedia projects have been an essential part of my learning journey ever since I started exploring the internet. Wikipedia and Wikisource are the reliable and freely accessible sources of knowledge I know, and I have been a user and contributor. As a Wikipedia editor, I have added and improved content, and I have also participated in a Wikimedia hackathon, which gave me my first real taste of contributing code and collaborating with the community.
For this Bulk OCR Improvements project, it was important for me to first understand the current state of the OCR tools on Wikisource. I went through the existing BulkOcrWidget, the Wikimedia OCR service, and related Phabricator tasks. To get familiar with the codebase and show my commitment.

About Me

I am an undergraduate student pursuing a Bachelor’s degree in Information Technology at Central University, Ghana. I am currently in my 3rd year and expect to graduate in 2027.
I first heard about the Google Summer of Code program after participating in a Wikimedia hackathon 2025 and I have been excited about the possibility of contributing code to improve tools that help Wikimedians like me. I will have no major time commitments. I will treat this project as a full-time 40-hour-per-week effort. I currently have no exams, internships, or planned vacations that would interfere with the program.

Past Experience

  • Links to any feature or bug fix you have written for a Wikimedia project during the application phase:
  • I participated in the Wikimedia Hackathon 2025 (Wiki-Mentor-Africa track), where I worked on task T388192: “BE - Design function to get lexeme forms”. The goal was to implement a backend function that takes a lexeme ID (or list of lexemes) and returns detailed form information. I designed the function and submitted a Gerrit patch to the labs/tools/wdaudiolex-be repository. Although the patch was later marked as a duplicate, the experience taught me how to work with MediaWiki-related backend code, understand Wikibase/Lexeme structures, navigate the Gerrit code review process, and collaborate with the Wikimedia technical community in a fast-paced environment.
  • I am an active editor and contributor on Wikipedia, Wikimedia Commons, and Wikisource. On Wikipedia I have improved articles, while on Wikisource I have done proofreading and used the existing OCR tools. This hands-on experience helped me deeply understand the pain points volunteers face when digitizing books, especially the repetitive and time-consuming nature of manual OCR which is exactly what the Bulk OCR Improvements project aims to solve.

    During the hackathon, I made my first code contribution to a Wikimedia tool. I have also worked on smaller personal and academic projects involving JavaScript, HTML, and CSS. These experiences strengthened my skills in object-oriented programming and reading/extending existing codebases — skills I am now applying while working on the Wikisource extension.

Any Other Info

I am familiar with the Wikimedia OCR service and have used it while contributing to Wikisource.
I am happy to create simple UI wireframes for the OCR engine selector and the preview/feedback modal if the mentors would find them useful. I have carefully read the main task T415145 and all related subtasks.
Even though I have not submitted microtasks before this application, I am highly motivated and ready to start contributing code immediately. I plan to dedicate the full 350 hours to delivering a clean, well-tested, and well-documented Bulk OCR feature that will make a meaningful difference for the Wikisource community.

Event Timeline

Gopavasanth subscribed.

Hi, thank you for your submission and the effort you put into your proposal. This year we received over 380 strong applications, and unfortunately we were not able to offer you a slot. This was a very competitive process, and many high quality proposals could not be selected. We truly encourage you to stay engaged and continue contributing to Wikimedia projects. Over the years, many contributors who were not selected for Google Summer of Code have gone on to make impactful contributions and become long term members of the community. Please do not see this as a failure, but as a step forward in your journey. We would love to stay in touch and support your continued involvement.

If you would like guidance on how to contribute to our projects outside GSoC, feel free to reach out to any of the mentors or org admins, they will be happy to help you get started.

You can get started or continue contributing here:

We hope to see your contributions in our community soon.