ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
ReVSI is a newly proposed benchmark and evaluation protocol that addresses the systematic invalidity of current VLM 3D reasoning assessments by re-annotating 381 scenes and regenerating QA pairs to ensure questions are answerable and accurate under realistic, sparsely sampled video inputs, thereby revealing previously obscured failure modes in spatial intelligence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to test a student's ability to navigate a house and count the furniture inside. You give them a video of the house and ask, "How many pillows are on the sofa?"
In the world of Artificial Intelligence, researchers have been using a standard test called VSI-Bench to grade how well AI models (specifically Vision-Language Models, or "VLMs") understand 3D spaces. But, according to this paper, that test is broken. It's like grading a student's math skills using a test paper where the numbers are blurry, the questions are missing key details, and the answer key is wrong.
The authors introduce a new, rebuilt test called ReVSI to fix these problems. Here is the breakdown of what they found and what they did, using simple analogies.
1. The Problem: The "Broken Map" and the "Blindfold"
The paper identifies two main reasons why the old test (VSI-Bench) was giving misleading grades:
The "Broken Map" (Annotation Drift):
The old test relied on "maps" (3D data) created years ago for different purposes. Think of it like using a sketchy, hand-drawn map from 1990 to navigate a city today.- Missing Objects: The old map might say a chair isn't there because the scanner missed it, but in the actual video, the chair is clearly visible. The AI gets penalized for seeing the chair, even though the "answer key" says it shouldn't be there.
- Wrong Labels: The map might label a "cup" as a "notebook" because the 3D shape was distorted. The AI gets confused because the video clearly shows a cup.
- Bad Geometry: The map might calculate a room's size as 20 square meters because the scanner got confused by a mirror, but the video shows it's actually 40 square meters.
The "Blindfold" (Frame Sampling):
Modern AI models often don't watch the whole video; they only look at a few snapshots (frames) to save time.- Imagine asking someone, "How many people are in this room?" but you only show them a photo taken when the room was empty.
- The old test assumed the AI could see the entire video. But in reality, if the AI only sees 16 snapshots out of 1,000, it might miss the object entirely. The old test didn't account for this, so it asked questions that were impossible to answer with the limited view the AI actually had.
2. The Solution: ReVSI (The "Freshly Renovated House")
The authors didn't just tweak the old test; they tore it down and rebuilt it from scratch. They call this ReVSI.
- Re-annotating the House: They hired human experts to look at the raw videos and manually draw new, perfect "maps." They fixed the missing chairs, corrected the wrong labels, and measured the rooms accurately based on what is actually in the video, not what a noisy scanner thought was there.
- The "Frame-Aware" Rule: They changed the rules of the game. Now, if an AI is only allowed to look at 16 frames, the test questions are generated specifically for those 16 frames. If the object isn't in those 16 frames, the question changes or is removed. This ensures the AI is only tested on what it can actually see.
- The "Dummy Video" Stress Test: This is their most creative tool. They created "dummy videos" where they digitally erased the object the AI was supposed to find.
- The Test: They show the AI a video of a living room but remove all frames containing the "pillows."
- The Goal: A smart AI should say, "I can't see any pillows, so the answer is zero."
- The Result: Many AI models, instead of saying "zero," guessed "2" or "3" because they had memorized that living rooms usually have pillows. This revealed that these models were hallucinating (guessing based on memory) rather than seeing (looking at the evidence).
3. The Shocking Results
When they ran the new test, the rankings of the AI models changed dramatically:
- The "Proprietary" Models (Big Tech): Models like GPT-5.2 and Gemini 3 actually did quite well. They seemed to rely more on what they actually saw in the video rather than guessing.
- The "Open-Source" Models: Many models that looked great on the old test suddenly looked much worse. Their scores dropped significantly because they were relying on the "broken map" and "biases" of the old test.
- The "Specialized" Models: Some models that were specifically trained to be good at 3D tasks actually got worse on the new test than their untrained versions. This suggests they had over-fitted to the bad data in the old test, learning the "wrong answers" instead of the right reasoning.
The Bottom Line
The paper argues that we can't trust the current scores of AI spatial intelligence because the tests were flawed. ReVSI is a new, stricter, and more honest exam. It forces AI to prove it is actually looking at the video frames provided, rather than just guessing based on what it thinks a room should look like.
In short: The old test was like a teacher grading a student based on a blurry photo and a wrong answer key. ReVSI is a teacher who hands the student a clear, high-definition video and asks questions that match exactly what is visible in that video.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.