Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR
This paper introduces PuMVR, the first benchmark quantifying significant orthographic bias in multilingual Vision-Language Models by demonstrating that these systems exhibit substantial performance gaps and inconsistent reasoning across Punjabi's three active scripts (Gurmukhi, Shahmukhi, and Roman), thereby challenging the assumption that one language equates to a single writing system.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant, multilingual robot assistant. You tell it, "I speak Punjabi," and it proudly says, "Great! I can help you in Punjabi." But here's the catch: Punjabi isn't just one way of writing. It's like a language that wears three different masks:
- Gurmukhi: The script used in India (like a blocky, structured font).
- Shahmukhi: The script used in Pakistan (like a flowing, cursive style).
- Roman: The language written with standard English letters (like typing "Hello" instead of "नमस्ते").
The paper argues that current AI models are playing a trick. They assume "One Language = One Script." They think if a model knows Punjabi in Gurmukhi, it automatically knows it in Shahmukhi and Roman. This is false.
The authors built a test called PuMVR (think of it as a "script-switching obstacle course") to prove this. They took 375 puzzles—ranging from identifying cultural objects like drums to solving visual riddles—and presented the exact same puzzle to AI models in all three scripts.
Here is what they found, using simple analogies:
1. The "Script Gap" (The Blind Spot)
When the AI solved a puzzle in Gurmukhi, it might get it right 90% of the time. But when you handed it the exact same puzzle written in Shahmukhi, its performance could drop to 60%.
- The Analogy: Imagine a chef who can perfectly cook a dish when the recipe is written in English, but if you hand them the same recipe written in French, they suddenly forget how to boil water. The ingredients (the visual image) and the goal (the answer) are the same, but the way the instructions are written breaks the chef's brain.
2. The "Visual Boost" Doesn't Fix the Problem
You might think, "Well, the AI can see the picture! That should help."
The paper tested this by giving the AI the image plus the text.
- The Analogy: It's like giving a person a map (the image) while they are trying to read a sign in a language they don't know. The map helps them a little bit, but it doesn't teach them how to read the sign. The AI still struggled with the Shahmukhi script even when it could see the picture. The bias wasn't fixed; it just got a tiny boost.
3. The "Reasoning Split" (The Identity Crisis)
The researchers asked the AI to "think out loud" (Chain-of-Thought) to see how it solved the problems.
- The Analogy: When the AI read the puzzle in Gurmukhi, it thought, "This is a Sikh cultural symbol, so I'll look for religious clues." But when it read the exact same puzzle in Shahmukhi, it suddenly thought, "This is an Islamic cultural symbol, so I'll look for different clues."
The AI wasn't reasoning about the picture; it was reacting to the shape of the letters. The script itself acted like a switch that changed the AI's entire personality and strategy.
4. The "Consistency Score" (The Real Test)
The paper introduces a new score called SCR (Script Consistency Rate).
- The Analogy: Imagine a student taking a math test. If they get 90% on the test written in Times New Roman, but only 60% when the test is written in Comic Sans, are they actually good at math? No. They are just good at reading Times New Roman.
The paper found that for many top-tier AI models, their "consistency score" was shockingly low (sometimes as low as 25%). This means they are failing to be truly "multilingual" because they can't handle the same language in different writing styles.
The Bottom Line
The paper concludes that we are fooling ourselves by saying AI is "multilingual." It's actually "multi-script" only in a very limited way. If you speak a language that uses multiple scripts (like Punjabi, Serbian, or Kurdish), your AI assistant might be reliable for you in one script but completely unreliable in another.
The authors aren't saying the AI is useless; they are saying it's unreliable for billions of people who switch between scripts. To make AI truly fair, we need to stop testing it on just one version of a language and start testing it on all the ways people actually write it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.