LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
The paper introduces LIBERO-PRO, a robust benchmark that exposes the severe reliance of current Vision-Language-Action models on rote memorization by demonstrating their complete performance collapse when evaluated under generalized perturbations of objects, states, instructions, and environments, thereby urging the community to adopt fairer evaluation standards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are training a robot chef to make a salad. You show it a video of someone picking up a bottle of salad dressing and putting it in a basket. The robot watches, learns, and then, when you ask it to do the same thing, it does it perfectly. You give it a high score: 95%! You think, "Wow, this robot is a genius at following instructions."
But here's the twist: The robot isn't actually thinking. It's just acting like a parrot that has memorized a specific script.
This is the core story of the paper LIBERO-PRO. The authors argue that the current way we test these "Vision-Language-Action" (VLA) robots is broken. It's like giving a student a math test where the questions are identical to the ones they practiced, just with the numbers shifted by one millimeter. The student gets 100%, but if you change the numbers slightly, they fail completely.
Here is a breakdown of what the paper found, using simple analogies:
1. The "Rote Memorization" Trap
The current standard test (called LIBERO) is too easy. The robots are trained on a set of tasks and then tested on the exact same tasks, with only tiny, almost invisible changes (like moving a bowl an inch to the left).
- The Result: Robots score over 90%.
- The Reality: They aren't understanding the task; they are just replaying a memorized movie of what they saw during training. They are "cheating" by memorizing the layout of the room and the sequence of moves, rather than actually understanding what a "basket" or "salad dressing" is.
2. The "LIBERO-PRO" Stress Test
To fix this, the authors created LIBERO-PRO. Think of this as a "surprise pop quiz" designed to catch cheaters. They systematically messed with four things to see if the robot could actually adapt:
- The Object Swap (The "Imposter"):
- The Test: You tell the robot, "Pick up the salad dressing." But in the scene, there is no salad dressing. Instead, there is a can of alphabet soup.
- The Robot's Reaction: It ignores the soup, looks at the empty space where the dressing usually sits, and tries to grab thin air. It's so stuck on its memory that it doesn't even notice the object is missing or different.
- The "Nonsense" Instruction:
- The Test: You tell the robot, "Pick up the salad dressing," but then you replace the words with gibberish like "fdsgfdsgsd."
- The Robot's Reaction: It doesn't pause or say, "I don't understand." It just grabs the salad dressing anyway. It proves the robot isn't actually listening to your words; it's just guessing based on what it sees.
- The "Moving Target" (Position Shift):
- The Test: You move the object just a little bit away from where it was in the training videos.
- The Robot's Reaction: It reaches for the old spot and misses the object completely. It has no concept of "where" the object is right now; it only knows "where it was supposed to be."
- The "Combo" Task:
- The Test: You ask the robot to do two simple things it knows how to do separately (pick up a bowl, then turn on the stove) but combine them into one new instruction.
- The Robot's Reaction: It fails. It can't stitch two learned actions together to make a new plan. It's like a musician who can play a scale and a chord, but if you ask them to play a song, they freeze.
3. The Big Reveal
When the authors ran these "stress tests" on the top robots (like OpenVLA and Pi0), the results were shocking.
- Standard Test Score: 90%+ (Looks amazing).
- LIBERO-PRO Score: 0% (Total failure).
The robots collapsed instantly. They couldn't handle a different object, a moved object, a weird instruction, or a new combination of tasks.
The Takeaway
The paper concludes that the current "gold standard" for testing robots is misleading. It's like judging a driver's skill only by how well they can drive a car on a perfectly straight, empty track they've driven a thousand times.
The authors are calling on the scientific community to stop using the old, easy tests. They want everyone to use LIBERO-PRO instead, which forces robots to prove they can actually see, understand, and adapt to the messy, unpredictable real world, rather than just reciting a memorized script.
In short: The robots aren't smart yet; they are just really good at memorizing. LIBERO-PRO is the test that finally exposes the difference.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.