← Latest papers
💻 computer science

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

This paper introduces Act2Answer, a novel evaluation protocol that adapts VLM benchmarks into action-based tabletop tasks to systematically measure and analyze knowledge retention in Vision-Language-Action models, revealing that while these models maintain basic concepts, they exhibit significant gaps in richer semantic knowledge compared to their source VLMs.

Original authors: Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin, Nikita Kurlaev, Daria Pugacheva, Albina Burlova, Mikhail Kolosov, Denis Shepelev, Andrey Kuznetsov, Elena Tutubalina, Aleksandr I. Panov, Alexey K.
Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin, Nikita Kurlaev, Daria Pugacheva, Albina Burlova, Mikhail Kolosov, Denis Shepelev, Andrey Kuznetsov, Elena Tutubalina, Aleksandr I. Panov, Alexey K. Kovalev, Vlad Shakhuro

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant student who has read every book in the library and knows the world inside out. You call them "The Knowledgeable One." Then, you hire them to work as a robot butler. You teach them how to pick up cups, open doors, and move furniture.

The big question this paper asks is: After all that training on how to move their hands, does the robot butler still remember the facts they learned from the books? Or did the "how to move" training accidentally wipe their brain clean?

The authors of this paper, led by Nikita Kachaev and colleagues, decided to find out. They built a new test called Act2Answer (Action to Answer).

The Problem: The "Success" Trap

Usually, when we test robots, we ask: "Did the robot successfully pick up the cup?"

  • If the robot picks up the cup, we say, "Great job!"
  • If it drops the cup, we say, "Bad job."

But this is like grading a student only on whether they can write their name, without checking if they actually know the answer to the math problem. A robot might drop a cup because its hand is clumsy (bad motor skills), not because it doesn't know that the cup is full of hot coffee and shouldn't be touched. The old tests mix up "clumsy hands" with "empty brains."

The Solution: The "Action Quiz"

To fix this, the researchers created a game called Act2Answer.

Imagine a table with two pictures on it:

  1. A picture of a sad person.
  2. A picture of a happy person.

The robot gets a voice command: "Put the cube on the sad person."

Instead of the robot talking back and saying "The sad person is on the left," the robot has to physically move a cube and place it on the picture of the sad person.

  • If it puts the cube on the sad picture, it gets a point.
  • If it puts it on the happy picture, it gets zero.

This is a "low-stakes" test. The robot doesn't need to walk across the room or lift a heavy sofa. It just needs to point with a cube. This isolates the robot's knowledge from its clumsiness.

What They Found: The "Simple vs. Complex" Gap

The researchers tested 7 different robot brains (VLA models) and compared them to their "parent" versions (the smart text-and-image models before they were trained to move).

Here is the story their data tells:

1. The Robot is Great at "What" (Simple Stuff)
If you ask the robot, "Is this a red ball or a blue ball?" or "Is this a circle or a square?", the robots are almost perfect. They can still see colors and shapes perfectly. It's like the robot remembers how to recognize a stop sign.

2. The Robot is Lost at "Why" and "Who" (Complex Stuff)
But when the questions get deeper, the robots start to fail.

  • Emotions: They struggle to tell the difference between a "sad" face and a "happy" face.
  • Social Rules: They don't seem to know that you shouldn't put a cube on a picture of a baby if the instruction implies safety.
  • Facts: They forget who famous people are or how to count objects correctly.

It's as if the robot learned to move its arm so well that it forgot how to be a human. The "clumsy hands" training seems to have overwritten the "smart brain" training for anything that isn't a simple shape or color.

3. The "Parent" vs. The "Child"
The researchers found a huge gap between the "Parent" (the smart text model) and the "Child" (the robot).

  • The Parent could answer 90% of the complex questions correctly.
  • The Child (the robot) often dropped to near 50% (which is just guessing).

This suggests that when we turn a smart AI into a robot, we often lose a lot of its common sense.

The "Hidden Knowledge" Mystery

The researchers also peeked inside the robot's brain (its internal layers) to see if the knowledge was truly gone or just hidden.

They found that the knowledge was still there in the middle layers of the robot's brain. The robot "knew" the answer internally. But when it tried to translate that knowledge into a physical action (moving the cube), the signal got lost.

The Analogy: Imagine a person who knows the answer to a riddle but has a speech impediment that makes them stutter so badly they can't say it out loud. The knowledge is there, but the "action" part of the brain can't access it.

The Takeaway

The paper concludes that simply teaching a smart AI how to move a robot arm isn't enough. We are currently training robots to be great at "doing" but terrible at "thinking" about the world.

To build a truly helpful robot assistant, we need to figure out how to teach them to move without making them forget the common sense they started with. The paper suggests that keeping the robot connected to language and world knowledge while it learns to move might be the key, but right now, that balance is missing.

In short: The robots can pick up the cup, but they might not remember that the cup is hot, or that the person holding it is sad. They are excellent movers, but they are forgetting how to be smart.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →