← Latest papers
🤖 AI

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

The paper introduces NumerosityVLM, a cognitively inspired benchmark with 10,800 controlled synthetic images that reveals Vision-Language Models' numerosity perception is primarily determined by their architecture and language components rather than visual conditions, with linearly separable numerosity signals emerging early in the vision encoder.

Original authors: Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan

Published 2026-08-18
📖 4 min read☕ Coffee break read

Original authors: Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Human beings possess a remarkable, innate ability to recognize how many items are in a group without counting them one by one. A baby can tell the difference between a pile of three crackers and a pile of four before they ever learn to speak, relying on a sense of number that emerges naturally in the developing brain. This skill, known as numerosity perception, is a fundamental part of how we understand the world. In recent years, scientists have turned their attention to artificial intelligence, specifically a type of computer system called a vision-language model. These systems are designed to look at images and answer questions about them, often performing with human-like fluency on complex tasks. However, a lingering question remains: do these machines truly understand the concept of "how many," or are they simply guessing based on visual patterns they have memorized?

To answer this, a team of researchers created a new testing ground called NumerosityVLM. They built a set of 10,800 computer-generated images designed to strip away the tricks that might fool a machine. In the real world, counting is often confused by other visual clues. For instance, a group of large objects might take up more space than a group of small objects, leading a system to think there are more of the large ones simply because they cover more area. To prevent this, the researchers controlled every detail of their images. They created scenarios where the number of items changed while the total space they occupied stayed the same, and vice versa. They also removed texture, shape, and color in stages, leaving behind simple dots, to see if the models relied on the look of the objects or the actual count. They tested seven different open-source models, ranging from smaller, older systems to larger, more advanced ones, asking them to count the items in each image.

The results revealed a clear divide between the models. The more advanced systems, which had larger internal structures, performed significantly better, achieving high accuracy even when the visual clues were misleading. The smaller, older models struggled, often failing to distinguish between different quantities and producing the same limited set of answers regardless of what was in the picture. The researchers found that the biggest factor determining success was not the visual conditions of the image, but the architecture of the model itself. The way the machine was built mattered far more than the specific visual tricks used in the test. This suggests that the ability to count is not just a matter of seeing clearly, but of having the right internal structure to process that information.

Digging deeper into how these models worked, the team looked at the different stages of the machine's processing pipeline. They discovered that the ability to distinguish numbers actually appears very early, right when the machine first looks at the image. Even the less successful models could detect the number of items in their initial visual layers. The problem, it turned out, happened later. The difference between a model that could count well and one that could not lay in how it translated those visual signals into words. The more successful models were better at taking the numerical information they had already seen and converting it into a correct text answer. The less successful models seemed to lose that information as they moved from seeing to speaking.

The study also examined how these machines handle the scale of numbers. The best-performing models were nearly perfect when counting small groups of up to four items, a range humans can recognize instantly. As the numbers grew larger, their accuracy dropped, and they began to consistently guess lower than the actual count. This pattern of underestimation grew stronger as the numbers increased, mirroring a known limitation in human perception where estimating large quantities becomes less precise. However, the most advanced models showed this bias much later than the weaker ones, suggesting they had a more robust grasp of larger numbers.

Ultimately, the research indicates that current vision-language models do possess a form of numerical understanding, but it is fragile and heavily dependent on their design. The ability to see "how many" is present in the early visual processing of these systems, but the final step of turning that perception into a reliable answer is where the machines often fail. The study concludes that the gap in performance is not due to a lack of visual data, but rather to how different models are built to interpret and express that data. By isolating these factors, the researchers have provided a clearer map of where artificial intelligence stands in its journey toward genuine numerical cognition, showing that while the eyes of the machine can see the number, the mind of the machine must still learn to count.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →