← Latest papers
💻 computer science

The power of context: Random Forest classification of near synonyms. A case study in Modern Hindi

This study demonstrates that a Random Forest classifier trained on Hindi word embeddings can successfully distinguish between Sanskrit and Perso-Arabic synonyms based solely on contextual usage patterns, providing quantitative evidence that etymological origin leaves a detectable imprint on language even when words share the same meaning.

Original authors: Jacek Bąkowski

Published 2026-04-03
📖 5 min read🧠 Deep dive

Original authors: Jacek Bąkowski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Can a Word's "Ancestry" Be Heard in Its Voice?

Imagine you have a room full of twins. They look exactly the same, they wear the same clothes, and they say the exact same things. In fact, they are so similar that if you asked them to describe a "cup of tea," they would both give you the same definition.

In linguistics, these are called synonyms. The paper asks a tricky question: Even if these words mean the same thing, do they "sound" different because of where they come from?

The author, Jacek Bąkowski, investigated this using Modern Hindi. Hindi is a unique language because it's like a giant cultural sandwich. It has two main layers of vocabulary:

  1. The Sanskrit Layer: The ancient, native roots of India (like the bread).
  2. The Perso-Arabic Layer: Words borrowed centuries ago from Persian and Arabic speakers who ruled the region (like the delicious filling).

For centuries, these two layers have lived side-by-side. For example, the word for "Army" can be Senā (Sanskrit) or Fauj (Persian). They mean the same thing, but the paper wanted to know: If you only look at the company these words keep (the other words around them), can a computer tell which one is "native" and which one is "foreign"?

The Experiment: The "Word Detective"

To solve this mystery, the author didn't ask human experts. Instead, he built a Random Forest, which is a type of computer brain (Machine Learning) that acts like a super-smart detective.

Here is how the detective worked:

  1. The Clues (Context): The computer didn't look at the dictionary definitions. It looked at Word Embeddings. Think of this as a giant map where every word is a dot. The position of the dot is determined by the other words it hangs out with. If a word is always seen with "kings," "palaces," and "poetry," it sits in a different part of the map than a word seen with "farmers," "fields," and "tools."
  2. The Training: The computer was shown 135 pairs of synonyms (like Nation vs. Qaum). It was told, "Here is the word; tell me if it's Sanskrit or Persian."
  3. The Test: The computer had to guess the origin of words it had never seen before, based only on the company they kept.

The Results: The Computer Got It Right!

The result was surprising and powerful. The computer got it right about 88% of the time.

What does this mean?
It means that even though Senā and Fauj mean the same thing, they live in different neighborhoods.

  • The Persian words tended to hang out in contexts related to administration, court, and specific cultural vibes.
  • The Sanskrit words tended to hang out in broader, more general, or religious contexts.

The computer didn't need to know the history books. It just needed to listen to the "chatter" around the words. The context (the other words) was loud enough to reveal the etymology (the origin).

The "Twin" Analogy: Why Some Twins Look Alike

The paper also looked at which words the computer got wrong. It found a funny pattern:

  • Specialized words (like specific terms for "Religion" or "Army") were easy to identify. They are like twins who wear very distinct uniforms.
  • Common, everyday words (like "Water," "Dream," or "Reality") were harder to identify. They are like twins who wear plain white t-shirts. Because they are used so often in so many different situations, they blend into the background.

The author suggests that the more common a word is, the more "polysemous" (having many meanings) it becomes. It's like a popular celebrity who is seen everywhere; it's hard to tell where they are from because they are everywhere. But the "specialized" words are like niche artists; you know exactly where they belong.

The "Magic" of the Map

One of the coolest parts of the study was testing how much of the map the computer needed to see.

  • The computer was given a map with 200 dimensions (a very complex, multi-layered map).
  • The author tried giving the computer only the first 50 dimensions (a blurry, partial map).
  • Surprise: The computer still worked almost as well!

This suggests that the "origin" signal is so strong and woven into the language that you don't need a perfect, high-definition map to see it. The signal is everywhere, like a scent that lingers in a room even if you only catch a whiff.

The Takeaway: Words Have a "Soul"

The main conclusion of this paper is that words are more than just definitions.

Even if two words are perfect synonyms, they carry a "cultural fingerprint."

  • Using a Persian word might subtly signal a certain level of formality, a connection to history, or a specific social class.
  • Using a Sanskrit word might signal something more traditional or religious.

The author calls this a "Semantic Frame." Imagine a picture frame. The picture inside is the definition (e.g., "Army"). But the frame itself is made of the word's history. The frame changes how you look at the picture.

In simple terms:
Language isn't just a list of instructions. It's a living history book. Even when we think we are just swapping one word for another, we are actually choosing a different perspective, a different cultural vibe, and a different piece of history. And thanks to this study, we now have mathematical proof that computers can hear that difference, too.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →