← Latest papers
💬 NLP

Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages

This paper presents the first cross-lingual visual reasoning audit for six Indian languages, revealing that multilingual Vision-Language Models suffer significant accuracy drops and exhibit degraded chain-of-thought performance in non-English contexts, particularly for Dravidian languages, demonstrating that current multilingual pretraining fails to transfer visual reasoning capabilities across linguistic and script boundaries.

Original authors: Swastik R

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Swastik R

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart robot tutor named "Visionary." This robot is amazing at looking at pictures of math problems, science diagrams, and geometry puzzles, and then solving them. In fact, if you ask it in English, it gets about 60–65% of the answers right. It's a star student.

But here's the catch: India has hundreds of millions of students who learn in languages like Hindi, Tamil, Telugu, Bengali, Kannada, and Marathi. The big question this paper asks is: "If we ask Visionary to solve these same problems in their native languages, does it stay a star, or does it suddenly become confused?"

The author, Swastik R, decided to find out by giving Visionary a massive "cross-lingual audit." Here is what happened, explained simply.

1. The Translation Test

The researcher took 980 tricky visual questions (like "What is the area of this shape?" or "Why is the sky blue?") and translated them from English into six major Indian languages. The pictures stayed exactly the same; only the words changed.

Then, they asked eight different versions of the robot (from open-source models to the most advanced ones like GPT-4o) to solve them.

2. The Big Drop in Performance

The results were a bit of a shock. When the robot switched from English to an Indian language, its performance took a nosedive.

  • The "English Advantage": Even the best models dropped in accuracy by 10 to 25 percentage points.
  • The "Dravidian Gap": The drop was even worse for languages from the Dravidian family (Tamil, Telugu, Kannada) compared to the Indo-Aryan family (Hindi, Bengali, Marathi).
    • Analogy: Imagine a runner who is great on a smooth track (English). When you switch them to a muddy field (Indo-Aryan languages), they slow down a bit. But when you put them in deep swamp water (Dravidian languages), they struggle to move at all.

3. The "Chain of Thought" Trap

You've probably heard that asking AI to "think step-by-step" (Chain-of-Thought) helps it solve hard problems. The researcher tried this in Indian languages.

  • The Result: It backfired! Instead of helping, asking the robot to "think step-by-step" in Bengali or Kannada actually made it worse.
  • Analogy: Imagine asking a person who is barely fluent in a foreign language to write a detailed essay explaining their logic. They get so stuck trying to form the sentences that they forget the actual math problem. The robot's "thinking process" is locked in English, so forcing it to think in another language just creates a jumbled mess.

4. The "Hidden English" Secret

The researcher peeked under the hood to see how the robots were thinking.

  • Some models (like Llama-4-Maverick) were secretly doing all their thinking in English, even when the question was in Tamil. They would just translate the final answer back to Tamil at the very end.
  • Analogy: It's like a student taking a test in French but doing all their mental math in their head in English. They might get the right answer, but if you asked them to explain why in French, they would likely fail. This is dangerous for education because a real tutor needs to explain concepts fluently in the student's language.

5. Size Doesn't Fix Everything

Usually, making a robot "bigger" (giving it more brain power) makes it smarter.

  • The researcher compared a smaller robot (7B parameters) to a huge one (32B parameters).
  • The Result: The bigger robot did slightly better, but the gap between English and Indian languages didn't close much.
  • Analogy: Giving a person a bigger dictionary helps them know more words, but it doesn't teach them how to speak the language fluently or solve logic puzzles in that language. You need specific training, not just more data.

6. The "Science Picture" Exception

There was one interesting twist. For some science questions that relied heavily on diagrams (like a picture of a cell or a machine), one specific model (Gemma 3-27B) actually did better in Indian languages than in English!

  • Analogy: Sometimes, a picture is worth a thousand words. If the image is clear enough, the robot doesn't need to read the words perfectly to guess the answer. But for math word problems, where the language is the key, the robot struggled immensely.

The Bottom Line

This paper is a wake-up call for EdTech companies and schools in India.

  • The Problem: If you deploy these AI tutors in regional schools today, non-English speaking students will get significantly worse help than English-speaking students.
  • The Cause: The robots aren't "multilingual" in their reasoning; they are just multilingual in their vocabulary. They haven't learned how to think in these languages.
  • The Solution: We can't just translate the questions. We need to train these robots specifically on how to reason in Indian languages, not just how to read them.

In short: The robot is fluent in English, but it's still a toddler in Tamil, Telugu, and Kannada. Until we teach it to "think" in those languages, it's not ready to be a teacher for India's millions of regional-medium students.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →