MM-THEBench: Do Reasoning MLLMs Think Reasonably?
This paper introduces MM-THEBench, a comprehensive benchmark designed to evaluate hallucinations in the intermediate reasoning steps of multimodal large language models, addressing the gap in existing assessments that overlook the internal thinking processes of post-trained reasoning models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a team of super-smart, multi-talented detectives (these are the Reasoning Multimodal Large Language Models, or MLLMs). These detectives are great at looking at a picture, reading a chart, or watching a video, and then solving a mystery.
In the past, these detectives would just shout out their final answer: "The suspect is the butler!" But recently, they've been trained to think out loud before speaking. They write down a long list of clues and logical steps (called Chain-of-Thoughts or CoTs) to explain how they reached that conclusion.
The problem? Sometimes, even if the detective writes a long, fancy list of steps, some of those steps are total lies (hallucinations). Yet, by sheer luck or by skipping over the lies, they still shout out the correct final answer.
This paper introduces a new tool called MM-THEBENCH to catch these detectives in the act of lying during their thinking process.
The New Detective Training Ground: MM-THEBENCH
Think of MM-THEBENCH as a high-tech "Lie Detector Test" specifically designed for the thinking process, not just the final answer.
1. The Three Types of Lies (The Taxonomy)
The authors realized that when a detective lies, it usually falls into one of three categories. They created a checklist to spot them:
- The "Fake Memory" Lie (Knowledge): The detective makes up a fact they learned in school. Example: "I know for a fact that the Eiffel Tower is in London." (It's not).
- The "Bad Eyesight" Lie (Perception): The detective looks at the photo but sees things that aren't there. Example: "I see a red car in the picture," when the car is actually blue.
- The "Bad Logic" Lie (Reasoning): The detective sees the right things but connects the dots wrong. Example: "The car is blue, and blue cars are fast, therefore this car is a Ferrari." (Bad logic).
2. The Grading System (Rubrics)
Instead of just checking if the final answer is right or wrong, MM-THEBENCH breaks the detective's thinking down into tiny atomic steps.
- Imagine a math problem. The "Gold Standard" solution has 5 specific steps.
- The system checks the detective's 20-step thinking process against those 5 gold steps.
- It asks: "Did you actually see the car? Did you know the rule about triangles? Did you do the math right?"
- It gives a score for every single step, marking exactly where the "lie" happened.
What They Found (The Results)
The authors tested 14 of the world's smartest detective models (including big names like GPT-5, Claude, and Gemini) using this new test. Here is what they discovered:
1. The "Correct Answer" Trap
Just because a detective gets the right final answer doesn't mean they thought correctly.
- Analogy: Imagine a student taking a math test. They write down a paragraph of nonsense, make up numbers, and get the wrong logic, but then they guess the final number and get it right.
- Finding: Many models got the right answer even though their "thinking steps" were full of hallucinations. The thinking process was not a faithful explanation of how they solved it.
2. The "Bad Eyesight" vs. "Bad Logic" Difference
This was a surprising discovery:
- Perception Hallucinations (Bad Eyesight): These happen all the time. The models often misidentify objects or colors. However, this rarely ruins the final answer. The model might think a car is red when it's blue, but it still figures out the car is moving.
- Reasoning Hallucinations (Bad Logic): These are rare, but deadly. If the model messes up the logic (e.g., "If A then B" when it's actually "If A then C"), the final answer is almost always wrong.
- Takeaway: It's better to have a detective with bad eyesight than one with a broken brain.
3. The "Overthinking" Problem
The paper found that when models are forced to think longer (more steps), they don't necessarily get smarter.
- Analogy: Imagine a detective who keeps writing notes. The first 5 notes are brilliant. The next 15 notes are just them rambling, repeating themselves, or making up new facts.
- Finding: As the thinking gets longer, the "precision" drops. The models generate a lot of extra, useless, or hallucinated steps just to fill the space. They are "overthinking" and drifting away from the truth.
The Bottom Line
The paper argues that we can no longer just trust the final answer a model gives us. We need to look at how they got there.
MM-THEBENCH is like a magnifying glass that forces us to look at the detective's notebook. It shows us that while these AI models are getting better at solving problems, their internal "thinking" is often messy, full of small lies, and sometimes completely disconnected from reality. To make them truly reliable, we need to fix their thinking process, not just their final answers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.