← Latest papers
🤖 AI

PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving

The paper introduces PHYSICS, a comprehensive benchmark of 1,297 expert-annotated university-level physics problems, which reveals that even the most advanced foundation models like o3-mini struggle with high-level scientific reasoning, achieving only 59.9% accuracy despite various optimization strategies.

Original authors: Kaiyue Feng, Yilun Zhao, Yixin Liu, Tianyu Yang, Chen Zhao, John Sous, Arman Cohan

Published 2026-05-22
📖 5 min read🧠 Deep dive

Original authors: Kaiyue Feng, Yilun Zhao, Yixin Liu, Tianyu Yang, Chen Zhao, John Sous, Arman Cohan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've built a fleet of incredibly smart, super-fast robots. These robots have read almost every book in the library and can solve complex math puzzles faster than a human calculator. You want to know: Can these robots actually think like a real physicist?

That's exactly what the researchers behind the PHYSICS benchmark wanted to find out. They didn't just ask the robots to solve simple riddles; they gave them a "PhD-level final exam" in physics.

Here is the story of their experiment, broken down simply:

1. The Exam: "The Ultimate Physics Test"

The researchers created a new test called PHYSICS. Think of it as a giant, digital question bank containing 1,297 difficult problems.

  • Where did the questions come from? They weren't made up by the researchers. They were taken from real, grueling qualifying exams that actual physics PhD students have to pass to become professional scientists.
  • What's on the test? It covers the big six areas of physics: how things move (mechanics), how atoms behave (quantum), heat and energy (thermodynamics), electricity and magnetism, light (optics), and atomic physics.
  • The Twist: Unlike school tests where you just pick "A, B, C, or D," these questions are open-ended. The robot has to write out the whole solution, showing its work, just like a human student.

2. The Grading Machine: "The Robot Teacher"

How do you grade a robot's math homework without a human reading every single page? The team built a super-strict automated grading system.

  • It uses a digital tool called SymPy (think of it as a math-checker robot) to verify if the robot's answer is mathematically correct, even if the robot wrote it in a slightly different way.
  • If the math is too weird for the robot to check, it calls in a "backup teacher" (a more advanced AI) to read the logic and decide if it makes sense.

3. The Results: "The Robots Hit a Wall"

The researchers tested 33 of the smartest AI models available today, including the "champions" like o3-mini, GPT-4o, and DeepSeek-R1.

The Scoreboard:

  • The Best Robot: Even the smartest model, o3-mini, only got about 60% of the questions right.
  • The Average Robot: Most other top models scored around 30% to 40%.
  • The Human Benchmark: The researchers noted that a human expert, given enough time, would typically score between 70% and 80%.

The Big Takeaway: Even the most advanced AI is still struggling to reach the level of a human physics expert. They are like a student who memorized the textbook but gets stuck when asked to apply the rules to a tricky, real-world scenario.

4. Why Did They Fail? (The "Hall of Errors")

The researchers looked closely at why the robots got the answers wrong. They found five main "bad habits":

  1. Making Up Rules: Sometimes, the robots invented facts that weren't in the question. It's like a student guessing the teacher's mind and adding extra rules that don't exist.
  2. Ignoring the Picture: Many physics problems come with diagrams. The robots often misread the picture, thinking a wire was connected one way when it was actually connected another.
  3. The "Over-Thinker" Trap: Some robots tried to think so hard and write so much that they ran out of space or got confused by their own long explanations. They got lost in their own reasoning.
  4. Math Slips: Even when the logic was perfect, they would make simple calculation errors in complex equations, like a human making a typo in a long math problem.
  5. Missing the Point: Sometimes, they just misunderstood what the question was actually asking, answering a different question entirely.

5. Did They Try to Help? (The "Cheat Sheet" Experiment)

The researchers tried a few tricks to help the robots do better:

  • Self-Reflection: They told the robots, "Hey, check your own work before you submit." This helped a little bit, making the answers slightly more consistent.
  • The "Google" Trick (RAG): They let the robots look up information online while they worked. This also helped improve scores, proving that the robots sometimes just needed a quick reminder of a specific formula or fact.

However, even with these helpers, the robots still couldn't fully close the gap with human experts.

The Bottom Line

The PHYSICS benchmark shows us that while AI is amazing at math and reading, solving deep, complex scientific problems is still a huge challenge. The robots are like very fast calculators that sometimes forget why they are doing the math. To get them to the level of a real physicist, we need to teach them to reason more deeply, read diagrams better, and stop making up facts.

This isn't just about physics; it's a wake-up call that for AI to truly help scientists, it needs to get much better at "thinking" like a human expert, not just "calculating" like a machine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →