← Latest papers
🔢 mathematics

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

The paper introduces PhysElite, a large-scale bilingual multimodal benchmark containing 11,586 Olympiad-level physics problems with step-by-step solutions, which reveals that even the most advanced multimodal large language models currently achieve only 33.7% accuracy on these complex reasoning tasks.

Original authors: Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang

Published 2026-08-27
📖 4 min read🧠 Deep dive

Original authors: Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Physics has always been a rigorous test of human intelligence, a discipline where understanding the world requires more than just memorizing facts. It demands the ability to look at a complex situation, often depicted in a drawing or a diagram, and mentally untangle the invisible forces at play. To solve a problem, one must identify the correct laws of nature, apply them step-by-step, and weave a logical chain of reasoning that leads to a single, precise answer. For decades, this kind of deep, multi-step thinking was considered a uniquely human strength, a frontier that machines could not easily cross. Today, artificial intelligence has made incredible strides, with computers capable of reading text and recognizing images with startling accuracy. Yet, a critical question remains: can these machines truly reason through the kind of difficult, real-world physics problems that challenge the world's best high school students?

A team of researchers has now built a new, massive testing ground to find the answer. They created a collection of over eleven thousand physics problems, drawn from the daily training materials of top-performing students in China. These are not simple questions with multiple-choice answers; they are open-ended challenges that require the solver to derive a solution from scratch. Crucially, the researchers did not just gather the questions; they also collected the detailed, step-by-step solutions that experts use to solve them, along with the diagrams that illustrate the physical setups. This new dataset, which includes problems in both English and Chinese, covers everything from the motion of planets and the flow of electricity to the behavior of light and the mysteries of the atomic world. By pairing every problem with its visual diagram and its full logical derivation, the team created a benchmark that mirrors the true complexity of expert-level physics reasoning.

When the researchers put eighteen of the most advanced artificial intelligence models to the test, the results were sobering. Even the strongest models, which had been trained on vast amounts of data and could perform complex reasoning tasks, struggled immensely. The best-performing model managed to get the final answer correct in only about one-third of the cases. This means that for every three difficult physics problems presented, the most powerful artificial intelligence currently available gets two of them wrong. The gap between these machines and human experts is substantial; the human students who served as a baseline for the study solved nearly half of the problems correctly, a significant lead over the artificial systems. The researchers found that this difficulty was not limited to a specific type of problem; the models stumbled across mechanics, thermodynamics, and optics alike, though they seemed to struggle most with problems involving light and geometric reasoning.

The study went deeper than just counting correct answers. The researchers examined the actual thought process of the machines, breaking down their solutions step-by-step to see exactly where they went wrong. They discovered that the errors were rarely simple calculation mistakes. Instead, the models often failed at the very beginning of the reasoning chain. They would misinterpret the diagram, misunderstand the physical setup, or apply the wrong law of physics to the situation. Once the initial setup was flawed, the rest of the solution, no matter how mathematically sophisticated, led to the wrong conclusion. In many cases, the models could produce a plausible-looking derivation that followed the right general path but collapsed under the weight of a single fundamental misunderstanding. This suggests that while these systems are excellent at pattern matching and generating text, they still lack the deep, intuitive grasp of physical reality that allows a human to visualize a problem and know which principles apply.

The researchers also tested whether giving the models extra help would improve their performance. They provided some models with additional diagrams that showed the intermediate steps of the solution, hoping this visual guidance would bridge the gap. While this did help slightly, the improvement was modest, and in some cases, adding extra text hints actually made the models perform worse. This indicates that the core issue is not a lack of information, but a fundamental difficulty in processing that information correctly. The models are not simply missing data; they are struggling to synthesize the visual and textual clues into a coherent physical model. The study concludes that despite the rapid progress in artificial intelligence, these systems are still far from mastering the kind of rigorous, multi-step reasoning required to solve authentic, high-level physics problems. The path forward will require new methods to teach machines not just to answer questions, but to truly understand the physical world they are asked to describe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →