CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
This paper introduces CompareBench, a comprehensive benchmark suite featuring TallyBench, OmniCaps, and a 1,200-question visual comparison dataset, to evaluate and reveal systematic weaknesses in current vision-language models regarding object counting, geometric, spatial, and temporal reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human vision is not merely about seeing; it is about comparing. When we look at a room, we do not just register that chairs and tables exist; we instantly judge which chair is taller, which table is closer, how many cups are on the counter, or whether a photograph shows a scene from the past or the present. This ability to weigh one thing against another is a fundamental part of how we understand the world. In recent years, computers have learned to see with remarkable skill, capable of describing images and answering questions about what they contain. These systems, known as vision-language models, are the engines behind many modern tools that can read a chart, describe a photo, or solve a visual puzzle. Yet, while these machines have become adept at recognition and description, it has remained unclear whether they possess the same intuitive grasp of comparison that humans use every second of the day.
A researcher has now built a specialized test to answer this question, creating a new evaluation suite designed to isolate the specific act of comparing visual information. They found that while the most advanced computer systems perform well on general tasks, they struggle significantly when asked to make direct comparisons about quantity, size, distance, or time. The study reveals that even the smartest models can fail at counting objects that are clearly visible, misjudge which of two items is longer, or get confused about which of two people is standing closer to the camera. The researcher discovered that these systems often stumble on tasks that are trivial for a human, suggesting that the ability to compare visual details remains a systematic weakness in current artificial intelligence.
To investigate this gap, the researcher constructed a comprehensive set of tests called CompareBench, which is supported by two other resources they developed: TallyBench and OmniCaps. TallyBench is a collection of two thousand images, each paired with a simple question asking the computer to count a specific type of object, such as dogs, books, or spoons. This resource serves as a foundation for testing how well a model can count individual items and also provides the source images for the quantity comparison task. OmniCaps is a smaller set of seven hundred and sixteen images featuring historical events, famous landmarks, and notable people, each tagged with a specific date. This collection allows the researcher to test whether a model can understand the passage of time by ordering images chronologically. The main test, CompareBench, brings these elements together into twelve hundred questions that ask the computer to make direct comparisons. These questions are divided into four categories: comparing quantities, comparing geometric properties like length and thickness, judging spatial relationships like depth and height, and ordering events in time. The quantity comparison subset specifically draws from TallyBench to form six hundred questions.
The researcher tested nine different versions of leading computer vision systems from three major technology companies. They asked these models to solve the twelve hundred comparison questions and the two thousand counting tasks. The results showed a clear divide between human performance and machine performance. Humans achieved near-perfect scores on almost every task, correctly counting objects and judging sizes with ease. The computer models, however, showed persistent errors. On the counting tasks, the best models got roughly eighty-seven percent of the answers correct, meaning they missed or miscounted objects in more than one out of every ten images. On the comparison tasks, the performance varied by category. The models were reasonably good at comparing quantities, but they struggled significantly with spatial reasoning, often failing to determine which object was closer to the camera or which point was higher off the ground. They also had trouble with geometric comparisons, such as deciding which of two books was thicker.
One of the most surprising findings emerged in the category of temporal, or time-based, comparison. In this section, the models were asked to look at images of historical scenes, landmarks, or famous people and decide which one appeared earlier in history. Here, the computer models actually outperformed the human test subjects. The humans, who were limited to looking only at the visual evidence in the pictures, often could not determine the correct order because the images looked very similar. The computer models, however, had access to a vast amount of pre-existing knowledge about history, architecture, and famous figures. They could draw on this internal database to correctly order the images, even when the visual clues alone were insufficient. This suggests that while the models lack a true visual understanding of space and size, they can sometimes compensate for this by using their massive memory of facts to solve problems that require external knowledge.
Despite these occasional successes, the study highlights that the core ability to compare visual elements is still not fully mastered by artificial intelligence. The researcher observed that the models frequently confused the length of an object with its height, or failed to distinguish between a foreground object and a background one. They also found that the models often missed partially hidden objects when counting, a task that humans handle instinctively. The researcher concluded that current systems are not yet reliable for tasks that require fine-grained visual comparison, and that these failures are not just random mistakes but systematic limitations in how the models process visual information. By providing a focused set of tests that isolates these specific skills, the new benchmark offers a clear way to measure progress in this area. The researcher hopes that by making their data and methods public, other scientists will be able to use these tools to build better models that can eventually match the human ability to see, compare, and understand the world around us.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.