← Latest papers
💬 NLP

Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education

This paper introduces the Visual Reasoning Benchmark (VRB), a dataset of 701 authentic primary school visual problems from Zambia and India, to evaluate Multimodal Large Language Models and reveal their significant limitations in dynamic spatial reasoning tasks despite proficiency in static skills, highlighting critical risks for classroom deployment.

Original authors: Mohamed Huti, Alasdair Mackintosh, Amy Waldock, Dominic Andrews, Maxime Lelièvre, Moritz Boos, Tobias Murray, Paul Atherton, Robin A. A. Ince, Oliver G. B. Garrod

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Mohamed Huti, Alasdair Mackintosh, Amy Waldock, Dominic Andrews, Maxime Lelièvre, Moritz Boos, Tobias Murray, Paul Atherton, Robin A. A. Ince, Oliver G. B. Garrod

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant new student named "AI." This student has read every book in the library and can solve complex riddles, write poetry, and do advanced math on paper. You might think, "Great! Let's put this student in a primary school classroom to help kids learn."

But before you do, you need to give them a specific test. Not a test of what they know from books, but a test of how they see the world.

This paper is about that test. The researchers built a special exam called the Visual Reasoning Benchmark (VRB). Here is the story of what they found, explained simply.

1. The Test: Real Classroom Problems

Most AI tests are like high-level chess puzzles or complex math equations written in text. But primary school kids don't just read; they look at pictures. They look at patterns, shapes, and diagrams.

The researchers went to real classrooms in Zambia and India. They grabbed actual exam papers used for 6th and 7th graders. These papers are full of pictures: "Which shape is the odd one out?" or "If you fold this paper, what does it look like?"

They made sure the pictures were raw and unedited. They didn't clean them up. They kept the smudges, the faint lines from photocopying, and the messy edges. Why? Because in many schools, that's what the kids actually see. If an AI can't handle a slightly blurry photocopy, it's not ready for the real world.

2. The Results: The "Jagged Frontier"

When they ran these tests on the smartest AI models available, the results were a mix of "Wow" and "Oh no."

Think of the AI's ability like a jagged mountain range.

  • The High Peaks: The AI is amazing at some things. It can count dots, spot big vs. small shapes, and follow simple lines. It's like a super-accurate scanner.
  • The Deep Valleys: But when the task requires the AI to "imagine" moving things in its head, it falls flat.

The researchers call this the "Spatial Ceiling."

  • Static Skills (The Peaks): If you ask, "How many circles are there?" or "Which one is bigger?", the AI gets it right most of the time.
  • Dynamic Skills (The Valleys): If you ask, "If I fold this paper in half, what shape do I get?" or "If I rotate this shape 90 degrees, what does it look like?", the AI gets confused. It's like asking a person to close their eyes and visualize a map, but the AI is trying to solve it by just looking at the pixels on the screen. It struggles to "mentally rotate" objects.

3. The "Ghost in the Machine" Problem

Here is the scary part for teachers: The AI sometimes gets the right answer for the wrong reason.

Imagine a student who guesses the answer to a math problem. They get it right, but they didn't understand the concept. If a teacher sees a correct answer, they might think, "Great, the student gets it!" But the student actually has a misunderstanding.

The AI does the same thing. It might guess the right answer to a folding puzzle because it recognizes a pattern, not because it understands how folding works. If a teacher uses this AI to grade papers or help a student, the AI might:

  • Mark a wrong answer as right (or vice versa).
  • Give a student a "hint" that sounds smart but is actually based on a misunderstanding.
  • Reinforce a wrong idea in the student's head.

4. The "Dirty Paper" Test

The researchers also tested the AI with "noisy" images—pictures that were faded, had coffee stains, or were poorly copied.

  • The Top Models: When the picture was perfect, they were great. But as soon as the picture got a little messy, their performance dropped sharply. It's like a race car that drives perfectly on a smooth track but crashes on a bumpy road.
  • The Bottom Models: They were bad at everything, but they didn't get much worse when the picture was messy. They were consistently confused.

This suggests that the smartest AIs are relying too much on "perfect" details and aren't truly understanding the structure of the image.

5. The Bottom Line: Not Ready for the Classroom (Yet)

The paper concludes that while AI is getting better, it is not yet ready to be a solo teacher for visual problems in primary school.

  • The Gap: There is a huge gap between what an AI can do with text (reading/writing) and what it can do with visuals (seeing/imagining).
  • The Risk: If we let AI grade these tests or help kids solve these puzzles without a human watching, we risk confusing the kids and reinforcing bad habits.
  • The Future: We need AI that can "think" in 3D space, not just read 2D pictures. Until then, these tools should be used with a "human in the loop"—a teacher who double-checks the AI's work.

In a nutshell: The AI is a brilliant librarian who can read any book, but it's still learning how to play with building blocks. Until it masters the blocks, we can't let it teach the kids how to build.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →