Extended Reasoning Mode Improves Large Language Model Accuracy on the Orthopaedic In-Training Examination: A Paired Within-Model Comparison
This study demonstrates that utilizing extended-reasoning inference significantly improves GPT-5's accuracy on Orthopaedic In-Training Examination questions compared to standard mode, though the persistence of confident errors suggests these models should serve as supervised adjuncts rather than primary study resources for residents.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every year, thousands of orthopaedic surgery residents in the United States take a rigorous exam to measure their knowledge of bones, joints, and muscles. This test, known as the Orthopaedic In-Training Examination, is a critical checkpoint in their training, helping program directors decide if a trainee is ready to become a board-certified surgeon. In recent years, these students have increasingly turned to artificial intelligence tools to help them study. These tools, called large language models, are computer programs trained on vast amounts of text that can answer questions, explain complex ideas, and even mimic the style of a medical textbook. They have become so capable that they can often pass medical exams themselves. However, a crucial question remained unanswered: does the way a student asks the computer to think change the quality of the answer? Specifically, does forcing the computer to pause and reason through a problem step-by-step before speaking produce better results than asking it to answer immediately?
A team of researchers set out to find the answer by putting a single, advanced artificial intelligence model through a series of tests using the actual questions from the 2024 orthopaedic exam. They did not compare different brands of computer programs against each other. Instead, they used the same program twice for every single question. First, they asked the model to answer in its standard mode, where it generates a response quickly. Then, they asked the exact same model the exact same question, but this time they switched it to an "extended reasoning" mode. In this second setting, the computer is instructed to spend more time internally breaking down the problem, weighing evidence, and thinking through the logic before it ever types out a final answer. The researchers wanted to see if this extra mental effort, which takes about a minute longer per question, would actually make the computer smarter.
The results were striking. When the model answered in the standard, quick mode, it got 79 percent of the questions right. This score was already impressive, placing the computer's performance roughly on par with a senior resident who has been in training for five years. But when the researchers switched the model to the extended reasoning mode, the accuracy jumped to 93 percent. This was not a case of the computer getting lucky on some questions and unlucky on others. The data showed a clear pattern: every question the computer got right in the fast mode, it also got right in the slow mode. The improvement came entirely from the questions it had previously missed; in the extended mode, it corrected nearly all of its mistakes. The researchers found that this boost in performance happened across almost every area of orthopaedic surgery, from hand injuries to spinal conditions, and it did not matter if the question included a picture of an X-ray or was just text.
Despite this dramatic improvement, the study uncovered a subtle but important warning for anyone using these tools. Even when the computer was wrong, it sounded just as confident as when it was right. In the few instances where the model still gave an incorrect answer, even after thinking deeply, it did not hesitate or express doubt. It would provide a detailed explanation and even cite medical literature to support its wrong choice, making the error sound authoritative. The researchers noted that if a student does not already know the correct answer, they might be easily misled by this false confidence. The computer is also capable of being tricked; if a user suggests a wrong idea, the model will often accept it and build a confident, detailed explanation around that mistake.
The study concludes that while these artificial intelligence tools are powerful, they are not perfect replacements for human teachers or verified study materials. The researchers suggest that for students using these tools, the best approach is to always use the extended reasoning setting, as the extra minute of thinking time significantly increases the chance of a correct answer. However, the output should always be treated as a draft that needs verification. The computer is a useful assistant for organizing study plans or summarizing concepts, but it cannot yet be trusted to judge its own accuracy. The findings show that the way we interact with these machines matters just as much as the machines themselves; a simple switch in settings can turn a good answer into a great one, but it cannot eliminate the need for human oversight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.