Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures
This paper introduces the "render ceiling," a model-free benchmarking method that uses inverted camera rendering of known crystal structures to definitively separate visual perception errors from reasoning failures in vision-language models, revealing that many models struggle with geometric extraction rather than logical inference.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of materials science, researchers spend their days trying to understand how atoms arrange themselves to form the solids that make up our world. These arrangements, known as crystal structures, are often visualized as three-dimensional drawings where spheres represent atoms and sticks represent the bonds holding them together. For decades, scientists have relied on these images to communicate complex spatial relationships, but a new generation of artificial intelligence is now being asked to read them. These vision-language models are designed to look at an image and answer questions about it, acting as a bridge between what a computer sees and what a human understands. The hope is that these systems could eventually help design new materials by interpreting scientific figures just as a human researcher would. However, a persistent problem has made it difficult to know if these machines are truly learning to see or if they are simply guessing based on patterns in the text. When a model gets an answer wrong, it is nearly impossible to tell if it failed because it could not interpret the picture correctly or because it failed to use logic to solve the puzzle. Without a way to separate these two skills, the scientific community cannot know where to focus its efforts to improve these tools.
A team of researchers has now introduced a new way to measure these models that bypasses the need for guesswork entirely. Instead of relying on another computer program to act as a judge, they created a system based on the exact mathematical rules used to draw the images in the first place. They started with a known set of crystal structures, whose atomic positions were perfectly defined, and used a standard camera setup to generate thousands of images. Because they knew exactly how the camera was positioned and what the original object looked like, they could mathematically reverse the process. By inverting the camera angles and re-calculating the positions of the atoms from the different views, they could recover the original structure with perfect precision. This process created a "ceiling" of perfect accuracy, a benchmark that represents the absolute maximum amount of information available in the images. If the images themselves contain the answer, this method proves it. If a model fails to find the answer, the failure must belong to the model, not the picture.
The researchers tested this method on over two thousand crystal structures, checking to ensure that no two atoms ever appeared to overlap in a way that would confuse the calculation. They found that for every single one of these structures, the mathematical reversal worked perfectly, recovering the correct answer every time. This meant the ceiling was solid; the images held the full truth. They then asked fourteen different vision-language models to identify the crystal system of these same images. The results were revealing. While the models performed better when they were given the exact atomic coordinates as text, they still fell far short of the perfect score, even with that extra help. For most of the models, the gap between their performance and the perfect ceiling was not caused by an inability to see the image. Instead, the majority of their errors happened after they had supposedly understood the picture, during the reasoning phase. The study showed that for thirteen out of the fourteen models, the difficulty lay in their logic, not their vision.
To prove that the images themselves were readable, the researchers trained a simple computer vision system that had no language capabilities at all. This system looked only at the pixels of the images, without any text or chemical knowledge. Remarkably, this simple system scored higher than every single vision-language model tested, correctly identifying the crystal structures in nearly ninety percent of cases. This finding suggests that the images are not too complex for a machine to read; rather, the complex models are struggling to process the information they see. The study also uncovered a hidden flaw in how these systems work. When asked to extract the coordinates of atoms from the images, a powerful model produced lists of numbers that looked perfectly formatted but were completely wrong, matching almost none of the actual atoms in the picture. In a standard test, this kind of error would be blamed on poor reasoning, but this new method showed it was actually a failure to see the image at all. The model was fabricating data, creating a false reality that looked correct on the surface but was empty underneath.
The implications of this work extend beyond just grading test scores. The researchers demonstrated that the way cameras are placed around an object matters more than the number of cameras used. They found that a specific arrangement of five views was enough to eliminate any ambiguity, while fewer views left room for error. This provides a clear rule for scientists who create benchmarks: the quality of the view matters more than the quantity. More importantly, the study offers a new tool for auditing artificial intelligence in scientific workflows. In the future, as these models are used to design new materials, this method can act as a check to ensure that the intermediate steps of their reasoning are based on real data and not on hallucinated numbers. By separating what the machine sees from how it thinks, the researchers have provided a way to build more reliable tools for science, ensuring that when a model claims to understand a crystal, it is actually seeing it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.