VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
This paper introduces VMMU, a Vietnamese benchmark comprising 2,500 multimodal questions across seven tasks designed to evaluate vision-language models' ability to perform genuine visual-textual reasoning beyond English, revealing that current models struggle primarily with multimodal grounding rather than OCR despite strong language-specific performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, multilingual robot assistant. You want to see if it can help a Vietnamese student solve a tricky homework problem that comes in a picture. The picture isn't just a photo; it's a page from a textbook or a driving test that has Vietnamese text mixed with diagrams, charts, and traffic signs.
The paper introduces a new "test" called VMMU to see how good these robots are at this specific task. Here is the breakdown of what they found, using simple analogies:
1. The Test: A "Multitask" Obstacle Course
Think of VMMU as a gym for AI, but instead of lifting weights, the robots have to solve 2,548 different puzzles. These puzzles cover seven areas: Math, Physics, Chemistry, Biology, Geography, Driving Tests, and IQ tests.
The catch? The questions aren't just typed out on a screen. They are embedded in images. To answer, the robot must:
- Read the Vietnamese words inside the picture (like reading a sign).
- Look at the picture part (like a graph or a map).
- Connect the two to figure out the answer.
2. The Big Surprise: Reading Isn't the Problem
The researchers first asked: "Is the robot failing because it can't read Vietnamese?"
They tested the robots' ability to simply transcribe (read out loud) the text inside the images.
- The Result: The robots were actually excellent at reading the Vietnamese text. They got it right almost 95% of the time.
- The Analogy: Imagine a student who can read every word on a page perfectly but still gets the math problem wrong because they don't understand how the numbers relate to the picture. The problem isn't the eyes; it's the brain.
3. The Real Struggle: Connecting the Dots
When the robots tried to solve the full problems (reading + looking + reasoning), their scores dropped significantly.
- The Result: Even the smartest commercial robots only got about 66% of the answers right.
- The "Thinking" Factor: The researchers found a huge difference between "fast" robots and "thinking" robots.
- Fast Robots: They guessed quickly and got about 50–70% right.
- Thinking Robots: These models pause to "think" through the steps. They scored much higher (73–86%).
- The Analogy: It's like the difference between a student who blurts out the first answer that comes to mind versus one who stops, draws a diagram, and walks through the logic step-by-step. The "thinking" ones did much better.
4. The "Clutter" Problem
The researchers noticed that when the text and the picture were crammed together in one messy image, the robots got confused.
- The Experiment: They took the text out of the picture and gave it to the robot as a clean list, while handing the picture separately.
- The Result: The robots got better at solving the problem.
- The Analogy: It's like trying to solve a puzzle while someone is shouting instructions at you from across the room. If they hand you the instructions on a clean piece of paper and let you focus on the puzzle pieces, you do much better. The robots struggled to ignore the "noise" of the text when it was mixed with the image.
5. Translation Doesn't Help
The researchers wondered: "Maybe the robots just don't know Vietnamese well enough. What if we translate the question to English?"
- The Result: No. Translating the question to English actually made the robots worse.
- The Analogy: Imagine a driving test where the signs are in Vietnamese, but the instructions are in English. The robot gets confused because the sign says "Stop" in Vietnamese, but the English instruction doesn't match the visual context. The robot needs to understand the whole Vietnamese context to solve the puzzle.
6. The "Guessing" Habit
Finally, they tried removing the picture entirely and just giving the robot the text question.
- The Result: The robots' scores dropped, but they didn't drop to zero. They still got about 51% right (compared to a random guess of 25%).
- The Analogy: It turns out the robots are like students who have memorized the "tricks" of the test. Even without the picture, they can guess the right answer based on the wording of the question alone, using patterns they've seen before. They aren't actually "seeing" the answer; they are just good at guessing.
The Bottom Line
The paper concludes that current AI models are great at reading Vietnamese text in images, but they are still struggling to connect that text with the visual evidence and reason through the problem. They need to learn how to "think" more deeply and handle the messy reality of text and images mixed together, rather than just relying on shortcuts or guessing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.