← Latest papers
💬 NLP

LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages

The paper proposes LEVOS, a novel approach that leverages Sanskrit-based character-level segmentation and a Transformer model to improve the translation of technical terms into low-resource Indian languages, achieving significant performance gains and enhancing educational accessibility.

Original authors: Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan, Preethi Jyothi

Published 2026-03-26
📖 5 min read🧠 Deep dive

Original authors: Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan, Preethi Jyothi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a child in a remote village how to understand complex science concepts like "photosynthesis" or "quantum mechanics." The problem is that the only textbooks available are in English, a language the child doesn't speak. You need to translate these words into their local language (like Marathi, Kannada, or Odia).

But here's the catch: these local languages are "low-resource," meaning there aren't many digital dictionaries or computers trained on them. If you just ask a standard AI to translate "photosynthesis," it might guess a word that sounds right but means something totally different, or it might just give up and leave the English word as is.

This paper, LEVOS, proposes a clever solution to this problem by using a "linguistic time machine": Sanskrit.

Here is the story of how they did it, broken down into simple steps:

1. The Problem: The "Lost in Translation" Game

Think of technical words (like "revaluation" or "injection") as complex Lego structures. In English, they are built with specific blocks. In Indian languages, they are also built with blocks, but the instructions for how to put them together are often missing because there isn't enough data to teach the computer.

When you try to translate these words directly from English to a local language, the computer gets confused. It's like trying to build a house using a blueprint written in a language you barely understand.

2. The Secret Ingredient: Sanskrit as the "Grandparent"

The authors realized that many Indian languages (like Hindi, Marathi, and Tamil) are like distant cousins. They all share a common ancestor: Sanskrit.

  • The Analogy: Imagine English is a modern smartphone. The local Indian languages are older, rugged phones. Sanskrit is the original blueprint that both the smartphone and the old phones were based on.
  • Even though the local languages have changed over time, they still share a huge amount of vocabulary and building blocks with Sanskrit.

The team's idea was simple: Don't translate directly from English to the local language. Instead, use Sanskrit as a bridge.

3. Step One: The "Word Splitter" (CharSS)

Sanskrit is famous for its "Sandhi" rules. This is like a magical glue. When two Sanskrit words are put together, they often melt into one long, unbroken word.

  • Example: "Nara" (man) + "Indra" (king) becomes "Narendra."

To use Sanskrit as a bridge, the computer first needs to know how to un-glue these words. The authors built a special AI tool called CharSS (Character-level Sanskrit Segmentation).

  • The Metaphor: Think of CharSS as a master chef who can take a giant, fused loaf of bread (a compound Sanskrit word) and perfectly slice it back into its original, individual ingredients (the sub-words).
  • They trained this AI using a powerful model called ByT5, which is great at looking at text letter-by-letter (or byte-by-byte) to find the hidden patterns.

4. Step Two: The "Translator's Assistant"

Once they could split Sanskrit words, they used a clever trick to help the main translator (an AI called NLLB).

  1. The Setup: They have English words and their Hindi translations (Hindi is "rich" in data, meaning the computer knows it well).
  2. The Magic: They took the Hindi translation, stripped away the Hindi-specific parts, and used their "Word Splitter" (CharSS) to turn the Hindi word into its Sanskrit "roots."
  3. The Boost: They fed the computer the English word plus these Sanskrit roots as extra hints.

The Analogy: Imagine you are trying to guess a word in a game of "Charades."

  • Without the trick: You only know the word is "Mass." You might guess "Heavy object" (physics) or "Crowd of people" (social). You are guessing in the dark.
  • With the trick: You are given a hint card that says, "This word comes from a Sanskrit root meaning 'people'." Suddenly, you know the answer is "Crowd," not "Heavy object."

5. The Results: A Smoother Bridge

The team tested this on three difficult fields: Administration, Biotechnology, and Chemistry.

  • The Outcome: By adding these Sanskrit "hints," the AI made significantly fewer mistakes. It translated technical terms much more accurately.
  • The Human Test: They even asked humans to check the results. The AI, using this Sanskrit bridge, produced translations that humans found more useful and accurate than the standard method.

Why Does This Matter?

This isn't just about computer science; it's about education and equality.

  • Right now, if you want to learn advanced science in a local Indian language, you often can't because the words don't exist in digital form.
  • By using this method, we can automatically generate high-quality technical dictionaries for languages that have very few resources.
  • The Big Picture: It's like building a bridge over a river of ignorance. The "Sanskrit bridge" allows knowledge to flow from English textbooks into the hands of students in villages who speak Marathi, Kannada, or Odia, making education accessible to everyone.

In a nutshell: The authors taught a computer to speak the "ancient language" of Sanskrit to help it understand how to translate modern technical words into local Indian languages, making learning easier for millions of people.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →