← Latest papers
💬 NLP

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

This paper introduces MUDIDI, a two-stage framework leveraging Vision Language Models and Large Language Models to effectively digitize multilingual dictionaries into machine-readable formats, supported by a new annotated dataset and benchmarks demonstrating the superior performance of LLMs in handling complex scripts and layouts.

Original authors: David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive library of old, dusty dictionaries written in hundreds of different languages. Some are written in English, but many are in rare, endangered languages with unique scripts that look like ancient runes or intricate calligraphy. These books are treasures for linguists and the communities that speak these languages, but right now, they are stuck in "scan mode." They are just pictures of pages. You can't search them, you can't count the words, and you can't easily use them to teach a language or build a translation app.

The problem is that these dictionaries are messy. They have weird layouts, tiny abbreviations, and symbols that standard computer scanners (like the ones you use to scan a receipt) get completely confused by.

The authors of this paper, MUDIDI, built a new "digital translator" to solve this. Think of it as a two-step assembly line that turns a picture of a dictionary page into a clean, organized digital database.

The Two-Stage Assembly Line

Stage 1: The "Super-Scanner" (Reading the Page)
Imagine a very careful librarian who looks at a chaotic page of text and tries to type it out exactly as it appears, preserving the bold words, the italics, and the order of the columns.

  • The Challenge: Standard scanners often miss letters or mix up the order of columns.
  • The Solution: The authors tested various "AI librarians" (specifically Large Language Models and Vision Language Models). They found that the smartest AI models (like Gemini) are like super-human librarians. They can read complex scripts (like ancient cuneiform or Arabic-based scripts) and keep the formatting perfect, far better than traditional scanning software.
  • The Trick: Sometimes, giving the AI a "cheat sheet" (a list of valid letters for that specific language) helps it read better, but not always. Sometimes, just letting the AI look at the picture is enough.

Stage 2: The "Organizer" (Sorting the Entries)
Once the text is typed out, it's still just a wall of words. It needs to be sorted into neat little boxes. A dictionary entry isn't just a word; it has a definition, a part of speech (noun/verb), examples, and cross-references.

  • The Challenge: The AI needs to look at the wall of text and say, "Okay, this bold word is the Headword, this italicized part is an Example, and this symbol means Cross-Reference."
  • The Solution: The AI is given a specific rulebook (called the MDF schema, which is like a standard filing system for dictionaries). The AI learns to cut the text into individual entries and tag every piece of information correctly.
  • The Boost: The authors found that if they gave the AI the dictionary's introduction page (which explains how the book is organized) along with the rulebook, the AI became much better at sorting the data. It's like giving a new employee the employee handbook before asking them to organize the filing cabinet.

What They Discovered

  1. Smart AI beats old scanners: The new "general-purpose" AI models are much better at reading these tricky, old dictionaries than the specialized scanning software we've used for decades. They handle weird fonts and layouts with ease.
  2. Context is King: If you want the AI to sort the dictionary entries correctly, you have to give it the "context." Showing the AI the dictionary's introduction page helps it understand the rules of that specific book, leading to much fewer mistakes.
  3. It works for the hard stuff: This system works on a huge variety of languages, from common ones like Greek and Japanese to endangered ones like Chukchi and Syriac.

The Result: A Digital Treasure Chest

The authors didn't just build the tool; they also built a test kit. They gathered 30 different dictionaries, had human experts check the work, and created a dataset to prove their system works.

In short: They created a way to take a picture of a dusty, complex dictionary page and turn it into a clean, searchable, machine-readable file. This allows communities and researchers to finally unlock the knowledge trapped in these old books, helping to preserve languages that might otherwise disappear.

The Catch: While the AI is amazing, it's not perfect. For very rare scripts or extremely messy layouts, human experts still need to double-check the work. But this new method makes that human job much faster and easier than before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →