EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
The paper introduces EgoGapBench, a diagnostic benchmark revealing that current multimodal large language models struggle with Egocentric Action Selection in multi-agent scenes by incorrectly adopting others' actions, a gap that persists even after fine-tuning on existing first-person data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are standing in a crowded room, watching a baseball game. You are not the batter, the pitcher, or the umpire. You are just an observer standing behind home plate.
Now, imagine a robot is standing right next to you, looking at the exact same scene through your eyes. The robot's job is to decide: "What should I do right now?"
According to this paper, current AI robots are terrible at this specific job. They keep getting confused and thinking, "Oh, I see a batter holding a bat! I must be the batter!" even though they are clearly standing behind the plate, not in the batter's box.
Here is a breakdown of the paper's findings using simple analogies:
1. The Problem: The "Body Cue" Crutch
For a long time, scientists tested AI on "first-person" videos (like a GoPro strapped to a person's head). In these videos, you can often see the person's hands, feet, or the tools they are holding.
- The Analogy: It's like a student taking a test while cheating by looking at their own hands. If the AI sees hands holding a bat, it knows "I am the batter."
- The Flaw: The paper argues that AI isn't actually understanding perspective; it's just recognizing body parts. If you flip the video upside down or remove the hands, the AI gets lost. It doesn't know who it is; it only knows what it sees.
2. The New Test: EgoGapBench
The researchers built a new test called EgoGapBench to see if AI can truly understand perspective without cheating.
- The Setup: They showed the AI images of busy scenes (like a baseball game or a wedding) where no one's hands or body are visible from the camera's point of view.
- The Trap: In the picture, there is a batter swinging a bat. The AI is supposed to be the umpire standing behind the plate.
- The Question: "What should you do next?"
- The Wrong Answer: The AI often picks "Swing the bat" because it sees the batter doing it. This is called a "Motor Intrusion" error. It's like the AI thinking, "I see someone running, so I must be the runner," even though it's standing still.
3. The Results: Humans vs. Robots
- Humans: When people took this test, they got it right almost every time (94.5%). We naturally understand, "I am the umpire, so I should lean forward to judge the pitch, not swing the bat."
- AI Models: Even the smartest AI models (like GPT-5 or Gemini) scored much lower (around 66% for the best, and near random guessing for others). They consistently tried to copy the actions of the people they saw in the picture.
4. The "Training" Mistake
The researchers tried to fix the AI by teaching it with more "first-person" videos (videos where you see hands and bodies).
- The Analogy: It's like trying to teach someone to drive a car by showing them videos of people driving, but never letting them sit in the driver's seat.
- The Result: This actually made the AI worse at the new test. The AI got even more addicted to looking for body cues. When those cues were missing in the new test, the AI crashed.
5. The Conclusion
The paper concludes that seeing a scene is different from acting from a specific perspective.
- Current AI is great at describing what it sees ("There is a batter").
- Current AI is bad at figuring out what it should do based on where it is standing ("I am the umpire, so I should watch").
The researchers say we need to stop just showing AI videos of people doing things and start teaching it how to understand its own "seat" in the room, separate from everyone else. They have released this test (EgoGapBench) so other scientists can try to build robots that don't get confused by the people standing next to them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.