From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
This paper introduces SFI-Bench, a video-based benchmark comprising over 1,500 expert-annotated questions that evaluates multimodal large language models on structured spatial and functional reasoning, revealing their current inability to effectively integrate spatial memory with functional inference for grounded intelligence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to navigate your house. You might think the hardest part is teaching it to recognize a chair or a coffee cup. But according to this paper, that's actually the easy part. The real challenge is teaching the robot two much harder things: where everything is in relation to each other, and what everything is actually for.
The researchers created a new test called SFI-Bench (Spatial-Functional Intelligence Benchmark) to see if today's smartest AI models can do these things. Think of it as a "driver's license exam" for AI, but instead of driving a car, the AI has to navigate a video of a room and answer tricky questions about it.
Here is a breakdown of what they found, using simple analogies:
1. The Two Big Skills: The Map and the Manual
The paper argues that true intelligence requires two complementary skills, which they call "Where Things Are" and "What They Are For."
- The "Where" (Spatial Reasoning): This is like building a mental map in your head. It's not just seeing a sofa; it's knowing that the sofa is to the left of the TV, that there are three pillows on it, and that if you walk past the sofa, you'll hit the wall.
- The Test: The AI is shown a video of a room and asked things like, "How many blue pillows are on the sofa that is furthest from the door?" or "If I walk from the kitchen to the bedroom, what object will I pass first?"
- The "What For" (Functional Reasoning): This is like understanding a user manual. It's knowing that a remote control is for the TV, not the toaster, and knowing exactly which buttons to press to cancel a washing machine cycle.
- The Test: The AI sees a strange device and must figure out, "What does this do?" or "My glasses are coming out of the dishwasher with a rainbow stain; what's wrong and how do I fix it?"
2. The Results: Good at Seeing, Bad at Thinking
The researchers tested many of the world's most advanced AI models (like GPT-5, Gemini, and open-source models) on this benchmark. Here is the verdict:
- The "Eyes" are Open: The AI models are excellent at simple perception. If you ask, "Is there a cat in the video?" they get it right almost every time.
- The "Brain" is Stuck: When the questions get harder, the models struggle.
- The Memory Gap: Imagine trying to remember a room after walking through it. The AI often forgets where it started. If you ask it to connect a dot from the beginning of the video to the end, it often gets lost or "teleports" objects to the wrong places.
- The "Overthinking" Trap: The researchers found that when they told the AI to "think harder" (by giving it more time to generate a long chain of reasoning), it didn't get smarter. In fact, it often got dumber. It started making up facts or getting confused by its own long explanations. It's like a student who keeps writing more and more on a test until they accidentally erase the right answer.
- The "Manual" Problem: For the "What For" questions, the AI often failed unless it was allowed to use a search engine to look up real-world manuals. Without that external help, the AI was just guessing, often relying on common sense that didn't fit the specific situation (e.g., assuming any remote works for any TV).
3. The Surprising Findings
- Time Doesn't Matter Much: Humans build mental maps by moving through time. If you shuffle the video frames so they play out of order, a human would be totally lost. Surprisingly, the AI didn't care much; it just looked at all the pictures at once and tried to guess. It doesn't seem to understand the "flow" of time like we do.
- Pictures Beat Words: The researchers tried feeding the AI a text description of the room instead of the video. The AI's performance dropped significantly. It turns out, even if you describe a room perfectly in words, the AI still needs to see the video to build a correct mental map.
- Bigger isn't Always Better: Sometimes, smaller, more efficient models performed just as well as massive ones, provided they didn't get bogged down in long, unnecessary reasoning chains.
The Bottom Line
The paper concludes that while AI is getting very good at "seeing" the world, it is still struggling to "understand" the world. It can recognize a toaster, but it doesn't truly grasp how the toaster fits into the kitchen or how to fix it if it breaks.
To build truly intelligent agents (robots or software that can act on our behalf), we need to move beyond just teaching them to recognize objects. We need to teach them to build coherent mental maps, remember where things are over time, and know how to use tools based on real-world knowledge, not just guesswork. Currently, they are like a tourist who can take a great photo of a landmark but doesn't know how to get there or what to do once they arrive.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.