When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
This paper systematically demonstrates that for visual spatial reasoning in vision-language models, indiscriminately injecting additional spatial, commonsense, or reasoning information often degrades performance, whereas targeted, selective information injection yields superior results.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a tricky puzzle with a friend who is very smart but sometimes gets confused by too many details. This friend is a Vision-Language Model (VLM)—a type of AI that looks at pictures and answers questions about them.
The specific puzzle you are giving them is Visual Spatial Reasoning (VSR). This means asking questions like: "Is the dog to the left of the cat?" or "Is the cup behind the book?"
Even though these AI friends are brilliant, they often get these spatial questions wrong. So, researchers (the authors of this paper) thought, "Maybe if we give them a little extra help, like a map or a hint, they will get it right!"
They tried giving the AI three types of "help":
- Spatial Cues: Giving them exact coordinates (like a GPS address) or descriptions of where things are.
- Commonsense Knowledge: Telling them facts like, "Dogs usually look at moving things."
- Chain-of-Thought (CoT): Asking them to "think step-by-step" before answering.
The Big Surprise: The researchers found that giving the AI more help often made it perform worse. It's like trying to navigate a city while someone shouts directions, traffic laws, and history facts at you all at once. You get overwhelmed and crash.
Here is the breakdown of their findings using simple analogies:
1. The "One Good Map" vs. "The Encyclopedia" (Spatial Context)
- The Idea: The researchers thought, "If we give the AI a bounding box (a box around the dog) AND its depth (how far away it is) AND its size, it will be perfect!"
- The Reality: It's like trying to read a map while someone is also shouting the latitude, longitude, street names, and the history of the buildings.
- The Lesson: Less is more. The AI works best when you give it one single, clear hint that matches the question. If you give it a bunch of different spatial facts at once, it gets confused and ignores the good ones. It's better to give it a clear arrow pointing left than a whole textbook on geometry.
2. The "Useless Fact" Problem (Commonsense Knowledge)
- The Idea: "Let's tell the AI that 'dogs chase cats' so it can guess the relationship!"
- The Reality: Sometimes this helps, but often it's like a detective trying to solve a crime while reading a biography of the suspect's great-grandfather. If the fact isn't exactly relevant, it becomes noise.
- The Lesson: You have to be very picky. Only give the AI facts that are highly relevant to the specific picture. If you give it too many facts, even the smart ones, it starts to hallucinate or get distracted.
3. The "Step-by-Step" Trap (Chain-of-Thought)
- The Idea: "Let's ask the AI to explain its thinking process before answering. That usually helps humans!"
- The Reality: This is a double-edged sword.
- Scenario A (Good): The picture is clear (e.g., a cat sitting on a table). When the AI thinks step-by-step, it confirms the obvious and gets it right.
- Scenario B (Bad): The picture is tricky (e.g., is the teddy bear to the left or right of the cat? It depends on whose perspective you take). If the AI tries to think step-by-step here, it might convince itself of a wrong answer because it's trying too hard to make a logical story out of a confusing situation.
- The Lesson: Asking an AI to "think step-by-step" only works if the starting facts are 100% clear. If the picture is ambiguous, forcing the AI to write a long explanation just makes it dig a deeper hole of confusion.
The Main Takeaway
The paper's title, "When More Is Less," is the golden rule here.
Imagine you are cooking a complex dish. If you add a pinch of salt, it tastes great. If you add a whole shaker of salt, a cup of sugar, and a bottle of hot sauce, the dish is ruined.
The researchers found that for AI to be good at spatial reasoning:
- Don't dump a pile of data on it.
- Select the one piece of information that matters most.
- Make sure the "thinking instructions" only happen when the picture is clear.
In short: To make AI smarter at seeing the world, we don't need to feed it more information; we need to feed it the right information, carefully and selectively.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.