EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
The paper introduces EPIC-Bench, a fine-grained benchmark comprising 6,600 annotated tuples across 23 embodied tasks, which reveals that current vision-language models struggle with complex visual-text alignment for physical interactions despite their advanced reasoning capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot assistant. You tell it, "Go pick up the red mug on the table." To do this, the robot needs to be able to see the mug, understand that "red mug" refers to that specific object, find a path to it without tripping, and know exactly where to grab it.
For a long time, we tested these robots' "eyes and brains" using simple quizzes. We'd show them a picture and ask, "Is there a red mug?" or give them multiple-choice answers like "A) Yes, B) No."
The problem? The robots were cheating. They weren't really "seeing" the mug; they were just guessing based on how the question was worded or using common sense. It's like a student who memorizes the answer key instead of learning the subject.
This paper introduces EPIC-Bench, a new, much harder test designed to stop the cheating and see if robots can actually see and interact with the real world.
Here is a breakdown of what they did, using some everyday analogies:
1. The New Test: "The Masking Game"
Instead of asking "Yes or No," EPIC-Bench asks the robot to draw a mask (like a digital sticker) over the exact object it needs to find.
- The Old Way: "Is there a cat?" (Robot guesses "Yes" because it sees a living room).
- The EPIC-Bench Way: "Here is a picture of a messy living room. Please highlight only the cat that is sleeping under the chair."
- Why it matters: If the robot highlights the whole chair or the floor, it fails. It proves the robot actually looked at the pixels, not just the words.
2. The Three Levels of the Test
The benchmark is like a video game with three levels of difficulty, covering everything a robot needs to do in a house:
Level 1: Target Localization (The "Spot the Difference" Game)
The robot has to find specific things based on tricky descriptions.- Basic: "Find the red cup."
- Harder: "Find the taller cup."
- Tricky: "Find the left half of the screen" or "Find the part of the human that is below the ears."
- The Catch: Sometimes there are zero cups, sometimes five. The robot has to count them perfectly, not just find one.
Level 2: Navigation (The "GPS" Game)
The robot has to figure out how to move.- Ground Detection: "Show me the floor I can walk on." (It must ignore rugs that are too slippery or holes in the ground).
- Feasible Path: "Draw a line from here to the door." The line must stay on the floor and not go through walls.
- Visual Matching: "Here is a picture of a toy in one room. Find that same toy in a different picture taken from a different angle."
Level 3: Manipulation (The "Handyman" Game)
The robot has to figure out how to use things.- Affordance: "Where do I grab this kettle?" (It needs to find the handle, not the hot bottom).
- Contact: "Which objects are touching the bread?"
- Placement: "Can I put this heavy box on that flimsy chair?" (The robot has to say "No" if it would break).
3. The Results: The Robots Are Still Learning
The researchers tested 89 different AI models (both famous commercial ones and open-source ones) on this new test.
- The Good News: The smartest models, especially those with "thinking" modes (like a human taking a moment to reason), did better. They are getting closer to understanding the world.
- The Bad News: Even the best models struggled.
- Counting is hard: If you ask them to find "three apples," they often find one or hallucinate a fourth.
- Parts are confusing: They are great at finding a whole "chair," but terrible at finding just the "leg of the chair."
- Direction is tricky: They get confused about left vs. right, or what "facing the window" means.
- The "Cheating" is gone: Because the test requires drawing precise masks, the models can't just guess. They have to actually look.
4. Why This Matters
Think of EPIC-Bench as a driving test for AI. Before, we just asked the car, "Do you know what a stop sign is?" Now, we are putting the car on a real road with obstacles, asking it to drive to a specific spot, and checking if it actually stayed in the lane.
The paper concludes that while our AI "eyes" are getting sharper, they still lack the deep understanding needed to safely and reliably interact with physical objects in our messy, real-world homes. This new benchmark gives researchers a clear map of exactly where the robots are failing so they can fix those specific problems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.