Probing Perceptual Constancy in Large Vision-Language Models
This study evaluates 155 Vision-Language Models across 236 experiments on color, size, and shape constancy, revealing significant performance variability and a distinct dissociation between shape constancy capabilities and those of color and size.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a red apple. If you walk away from it, it looks smaller. If you shine a blue flashlight on it, it looks purple. If you look at it from the side, it looks like an oval instead of a circle.
A human brain is amazing because it knows the apple is still a big, red, round apple, regardless of those changes. It doesn't get fooled by the distance, the weird light, or the angle. This superpower is called Perceptual Constancy.
This paper is like a giant report card for 155 different "AI brains" (called Vision-Language Models) to see if they have this same superpower. The researchers built a special test called ConstancyBench to check three specific skills:
1. The Three Skills Tested
Think of these skills as three different levels of a video game:
- Color Constancy (The "Lighting" Level): Can the AI tell that a white wall is white, even if the room is lit by a yellow lamp or a blue neon sign?
- The Test: The AI looks at a Rubik's cube in weird lighting and has to guess the true color of the middle square.
- Size Constancy (The "Distance" Level): Can the AI understand that a car far away is the same size as a car right next to it, even though the distant one looks tiny on the screen?
- The Test: The AI looks at two missiles, one far back and one up close, and has to decide if they are actually different sizes or just look different because of perspective.
- Shape Constancy (The "Angle" Level): Can the AI know that a round plate is still a circle, even if it's tilted so it looks like an oval?
- The Test: The AI looks at a door that is slightly open and tilted, making it look like a trapezoid, and has to realize it's actually a rectangle.
2. The Results: Who Passed?
The researchers tested 155 different AI models, ranging from small, lightweight ones to massive, super-smart ones (like GPT-4o). Here is what they found:
- The "Shape" Skill is Easy: The AI models were surprisingly good at Shape Constancy. It's like they are great at recognizing the "outline" of things. Even smaller models could figure out that a tilted rectangle is still a rectangle.
- The "Color" and "Size" Skills are Hard: The models struggled much more with Color and Size. They often got tricked by the lighting or the distance. It's as if they are looking at a flat picture and forgetting that the world is 3D and has changing light.
- Bigger Brains Do Better: There was a clear rule: The bigger the AI model, the better it was at these tests.
- Small models often failed, getting confused by simple tricks.
- The massive models (the "giants" of the AI world) did much better, especially at understanding size and distance. It seems that as you give the AI more "brain power" (parameters), it gets better at understanding the 3D world.
3. The Big Takeaway
The paper concludes that AI doesn't have a single, perfect "human-like" vision yet. Instead, it has a stratified (layered) vision:
- It's good at simple geometry (Shape).
- It's getting better at understanding the world's scale and lighting (Size and Color), but only if the model is huge.
The authors suggest that while these AI models are getting smarter, they might not be "seeing" the world the way humans do. They might just be very good at guessing based on patterns they've seen in their training data, rather than truly understanding physics.
In short: If you ask an AI to identify a shape, it's usually smart. If you ask it to figure out how big or what color something really is when the lighting or distance is tricky, it's still learning. And the bigger the AI, the less likely it is to get fooled.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.