A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
This paper introduces CrossHOI-Bench, a unified multiple-choice benchmark designed to overcome the limitations of existing HOI evaluation methods, revealing that while large vision-language models achieve competitive zero-shot performance, they lack the multi-action reasoning and precise person-action assignment capabilities of specialized HOI methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand what people are doing in a photo. Specifically, you want it to spot Human-Object Interactions (HOI)—like a person riding a bike, cutting a cake, or hugging a dog.
For years, we've had two types of "robots" (AI models) trying to do this:
- The Specialists: Models built specifically for this one job. They are like a master chef who only cooks pasta. They are fast and good at their specific task, but they might get confused if you ask them about sushi.
- The Generalists (VLMs): Massive, "all-knowing" Vision-Language Models (like the brains behind advanced chatbots that can see). They are like a Swiss Army Knife. They can write poetry, solve math, and describe a sunset. The big question was: Can a Swiss Army Knife actually cook pasta better than the master chef?
The problem was that we didn't have a fair way to test them.
The Problem: The "Unfair Test"
The old way of testing these robots was like a multiple-choice quiz where the teacher only wrote down one correct answer, even if there were two or three valid ways to describe the picture.
- The Scenario: A photo shows a person standing near an airplane.
- The Reality: The person could be boarding the plane, exiting the plane, or just waiting for it. Both "boarding" and "exiting" make sense visually.
- The Old Test: The teacher's answer key only said "Boarding."
- The Result: If the Generalist robot said "Exiting," the test marked it wrong. If the Specialist robot guessed "Boarding," it got it right.
This was unfair to the Generalists because they are designed to be flexible and creative. They naturally see multiple possibilities. The old test punished them for being smart and flexible, making them look worse than they actually were.
The Solution: CrossHOI-Bench (The New, Fair Exam)
The authors of this paper built a new testing ground called CrossHOI-Bench. Think of it as redesigning the exam to be fair for both the Specialist and the Generalist.
Instead of asking, "What is the one thing happening?", they ask:
"Here is a picture. Here are four options: A, B, C, and D. Which of these are happening?"
- The Options:
- (A) Boarding the plane
- (B) Exiting the plane
- (C) Eating a sandwich
- (D) Flying the plane
- The New Rule: If the person is doing both boarding and exiting (or if the image is ambiguous), the model gets credit for picking both A and B.
They also made the test harder. They removed the "easy" pictures (like a single person holding a cup in a white room) and focused on the messy, real-world stuff:
- Crowded scenes with 5 people doing different things.
- Tricky moments where it's hard to tell if someone is holding a phone or texting on it.
- Confusing situations where Person A is holding a ball, but Person B is throwing it, and the robot has to know who is doing what.
What They Found (The Plot Twist)
When they ran this new, fair test, the results were surprising:
The Generalists (VLMs) are surprisingly strong:
The "Swiss Army Knives" (like Qwen2.5 and InternVL) performed better than the Specialists in many areas, even without any special training! They are great at understanding the context and the story of the image. They can reason, "Oh, that person is looking at the cake, so they are probably cutting it."But they have a specific weakness:
The Generalists sometimes get confused in a crowd. If two people are standing next to each other, the Generalist might mix them up. It might say, "The guy on the left is riding the horse," when actually, the guy on the right is. They struggle to keep track of who is doing what when everyone is close together.The Specialists still have a job:
The "Master Chefs" (Specialist models) are still very good at the technical stuff. They are better at pinpointing exactly which person is doing the action and handling multiple actions happening at once (like "riding" AND "hugging" the horse simultaneously).
The Big Takeaway
This paper tells us that we don't need to choose between the "Specialist" and the "Generalist."
- Generalists are amazing at understanding the big picture and reasoning through complex scenes.
- Specialists are still the kings of precise localization and handling crowded, multi-action scenes.
The authors created a new "fair exam" (CrossHOI-Bench) that finally lets us see the true strengths and weaknesses of both types of AI. It turns out, the future of seeing and understanding the world isn't about one robot doing everything; it's about combining the creativity of the Generalist with the precision of the Specialist.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.