Testing the capability and performance of Claude Sonnet 4 on the GRE physics test
This paper demonstrates that Claude Sonnet 4 achieved a perfect scaled score of 990 on a 70-question GRE Physics practice exam, significantly exceeding the authors' hypothesis of 90% accuracy and highlighting the model's strong capabilities in complex scientific problem-solving.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Benchmarking Claude Sonnet 4 on the GRE Physics Examination
Problem Statement
As artificial intelligence (AI) becomes increasingly integrated into K–12 and post-secondary education, there is a need to benchmark the capabilities of Large Language Models (LLMs) in solving complex scientific problems and demonstrating critical thinking. While previous research indicates that general-purpose AIs often struggle with upper-division university physics problems (achieving only 35–39% accuracy in fields like optics and thermodynamics), newer specialized models enhanced with advanced processing capabilities suggest improved potential. This study sought to evaluate the performance of Claude Sonnet 4, an LLM developed by Anthropic, against a standardized, graduate-level physics assessment to determine if it could exceed a 90% accuracy rate on complex problem-solving tasks.
Methodology
The researchers utilized a publicly available 70-question GRE Physics practice examination (Form GR1775) provided by the Educational Testing Service (ETS). The study employed the following protocol:
- Model: Claude Sonnet 4, identified at the time as Anthropic's most advanced free model.
- Input: Each of the 70 questions was submitted individually to the model as a screenshot, with no additional text or context provided.
- Scoring Protocol: Responses were compared against the official ETS answer key. Due to message limits, multiple chat sessions were initiated to process all questions. Partial credit was awarded if the model answered incorrectly on the first attempt but correctly on a second attempt; otherwise, questions were marked fully incorrect.
- Analysis: The total number of correct responses was tallied and converted to a scaled score using the official GRE conversion sheet. The study also analyzed performance across specific content areas (Classical Mechanics, Electromagnetism, Quantum Mechanics, Atomic Physics, etc.) and compared model performance against the mean accuracy of human test-takers for specific questions.
Key Results
The study yielded the following quantitative and qualitative findings:
- Overall Performance: Claude Sonnet 4 correctly answered 65 out of 70 questions, achieving an accuracy rate of approximately 93%. This corresponds to a scaled score of 990, which is the maximum possible score on the GRE Physics test.
- Content Area Breakdown:
- Quantum Mechanics & Atomic Physics: The model achieved 100% accuracy (16/16 questions).
- Electromagnetism: The model scored 12.5/13 correct.
- Classical Mechanics: This was the model's weakest area, with a score of 12.5/14 correct (89% accuracy).
- Correlation with Human Performance: No correlation was found between the model's performance and the mean accuracy of human test-takers. For the 17 questions where human accuracy was between 20–30%, the model lost only 0.5 points on a single problem.
- Error Analysis: When errors occurred, they were attributed to incorrect synthesis, such as miscalculating ratios or misanalyzing specific details within the problem statement, rather than a lack of conceptual understanding. The model consistently provided in-depth explanations for its chosen answers, which facilitated the identification of specific reasoning errors.
Significance and Claims
The paper claims that the results support the hypothesis that consumer-grade LLMs can solve and explain questions in advanced scientific domains with high proficiency. The authors conclude that:
- High-Level Capability: LLMs are capable of achieving maximum scores on graduate-level physics examinations, suggesting a transformative potential for scientific study and education.
- Educational Utility: These models serve as powerful tools for working through convoluted questions without human error, provided their limitations in reasoning regarding philosophical concepts or unresolved cosmological questions are acknowledged.
- Future Trajectory: Through continued machine learning and user training, the authors suggest models may eventually reach 100% accuracy. They also note that performance may vary based on model versions (e.g., Opus vs. Sonnet) and that paid subscriptions may offer further advantages in usage allotments and model quality.
The study emphasizes that while AI demonstrates significant promise, it remains subject to limitations in detail-oriented synthesis and does not yet possess the full generative capabilities (such as image or video generation) found in other competing models like ChatGPT.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.