← Latest papers
💬 NLP

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

This paper introduces OmniPhys, a large-scale multimodal benchmark derived from Chinese educational corpora that evaluates and advances physics understanding and structured diagram generation in Multimodal Large Language Models through 15,246 questions and 19,850 images.

Original authors: Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang

Published 2026-08-27
📖 5 min read🧠 Deep dive

Original authors: Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Physics is the study of how the universe works, from the way a ball rolls down a hill to the invisible forces that guide electricity through a wire. For centuries, humans have learned these rules by reading words, looking at diagrams, and solving problems that require connecting text with images. Today, computers have become incredibly good at reading text and looking at pictures, often called multimodal models because they can handle both types of information at once. These machines can write stories, answer questions, and even describe what is happening in a photograph. However, when it comes to the strict, logical world of physics, simply recognizing an image or reading a sentence is not enough. True understanding requires the computer to see how a diagram and a description fit together to form a single, unbreakable chain of logic. If a machine cannot do this, it is merely guessing, not reasoning.

A team of researchers from East China Normal University has built a new tool to test exactly how well these computers can think like physicists. They created a massive collection of problems called OmniPhys, drawn from real Chinese textbooks and exams ranging from middle school to university. This is not just a list of questions; it is a rigorous test that forces computers to do more than just pick the right answer from a list. The collection includes over 15,000 questions and nearly 20,000 images, covering mechanics, electricity, light, heat, and sound. The researchers designed it to be difficult, ensuring that the problems require the computer to interpret complex drawings and text simultaneously. Crucially, the test also asks the computers to do something they rarely do: draw their own diagrams. Instead of just selecting a picture, the computer must generate a new image that correctly shows how light reflects or how forces act, adhering to the strict laws of physics.

When the researchers ran their tests, they found that even the most advanced computers, including the most powerful ones available to the public, struggled significantly. The best models could solve about 70 percent of the problems correctly, but when the researchers looked closer at how the models arrived at those answers, the picture changed. Many models got the right answer by accident or by memorizing patterns, rather than by truly understanding the logic. They often missed key steps in their reasoning or failed to connect the visual clues in the diagram with the text. The situation was even worse for the task of drawing. When asked to generate a physics diagram, the computers frequently produced images that looked plausible at first glance but contained fundamental errors. They might draw a light ray bending the wrong way or a force pointing in the opposite direction, violating the very laws of nature they were supposed to demonstrate.

The study also revealed that these computers get progressively worse as the problems get harder. They performed reasonably well on middle school questions but their accuracy dropped sharply as the problems moved into high school and university levels. This decline suggests that while these machines are good at handling simple tasks, they have not yet mastered the deep, layered reasoning required for complex scientific thinking. The researchers also discovered that giving the computer a picture of the problem helped it perform better than if it only had the text, proving that visual information is essential for solving these puzzles. However, they found that some models actually performed worse when given an image, suggesting that the computer sometimes treats the picture as confusing noise rather than helpful information.

Perhaps the most surprising finding was that the computers used to grade the work were often too kind. When the researchers asked an artificial intelligence to grade the diagrams drawn by other computers, the grader gave high scores to images that human experts rated as incorrect. The automated grader missed subtle physical errors that a human would catch immediately. This means that relying on machines to evaluate the quality of scientific reasoning is currently unreliable. The human experts, who reviewed the work carefully, found that the computers were failing to grasp the core concepts of physics, even when the final answer happened to be right.

This work does not claim that computers are useless for science, but it does show that they have a long way to go before they can truly understand the physical world. The researchers have made their collection of problems and images available to the public so that other scientists can use it to build better models. By identifying exactly where these machines fail—whether in drawing a correct diagram, following a logical chain, or understanding a complex visual scene—the study provides a clear map for future improvements. The goal is not just to make computers that can pass a test, but to create systems that can reason through the laws of nature with the same rigor and accuracy as a human physicist. Until then, the gap between what these models can say and what they truly understand remains wide.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →