Do AI-generated surgical cases improve clinical reasoning? A quasi-experimental comparison in surgical clerkship
This quasi-experimental study demonstrates that surgical clerkship students trained with DeepSeek-generated cases achieved significantly higher clinical reasoning scores and more favorable cognitive load profiles compared to those trained with traditional expert-authored cases, suggesting that AI-generated content can effectively enhance medical education when supported by expert review.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Surgery is as much a mental discipline as a physical one. Before a surgeon ever picks up a scalpel, they must learn to think like a detective, gathering clues from a patient's history and symptoms, weighing different possibilities, and committing to a course of action even when the picture is incomplete. This mental process, known as clinical reasoning, is the engine of safe medical practice. When it works, patients get the right care; when it fails, mistakes happen. Traditionally, medical students learn this skill by working through written stories of real patients, called cases, under the guidance of teachers. However, creating these stories is slow, expensive work for busy doctors, and the resulting collection of cases often feels repetitive, showing students only the most typical versions of diseases rather than the messy, confusing reality they will face in the hospital.
Artificial intelligence has recently offered a potential solution to this bottleneck. Large language models, which are computer programs trained on vast amounts of text, can now write these patient stories instantly. But a critical question remained unanswered: does reading a story written by a machine actually help a student think better than reading one written by a human? A team of researchers in Shanghai set out to find the answer by testing whether cases generated by an artificial intelligence system could improve the reasoning skills of surgical students more effectively than traditional, human-written cases.
The study took place over two separate six-month periods at a major teaching hospital in China. The researchers worked with two groups of fifth-year medical students, each group consisting of thirty students who were rotating through their surgical clerkship. The first group, acting as the control, spent their time working through twenty-four patient cases that had been written by experienced faculty members over the previous three years. These were the standard, expert-authored stories the department had always used. The second group, the experimental group, worked through a different set of twenty-four cases. These were not written by humans but were generated by a specific artificial intelligence model called DeepSeek. The researchers carefully matched the two sets of cases so that they covered the exact same types of diseases, such as appendicitis, gallbladder inflammation, and intestinal blockages, ensuring that the only major difference was who or what wrote the stories.
To measure how well the students were learning, the researchers did not simply ask them to recall facts or choose the right answer from a list. Instead, they used a sophisticated digital platform that mapped out the students' thought processes. As students worked through new, unseen patient scenarios, the system tracked every decision they made: what questions they asked, what tests they ordered, and how they interpreted the results. This allowed the researchers to see not just if the student got the right diagnosis, but how they arrived at it. They looked at whether the students missed important steps, jumped to conclusions too quickly, or followed a logical path that an expert would recognize. They also asked the students to report on how mentally taxing the work felt and how confident they felt in their own abilities.
The results showed a clear advantage for the students who trained with the artificial intelligence cases. On the final assessment, the group that used the AI-generated stories scored significantly higher on their overall reasoning performance than the group that used the human-written stories. The difference was substantial enough to be considered a large improvement. When the researchers broke down the scores, they found that the AI group was better at two specific things: they were more accurate in their final diagnoses, and they were much better at following a complete, logical path of inquiry without skipping steps or settling on an answer too early. Perhaps most importantly, the students in the AI group made fewer errors in judgment, such as ignoring contradictory evidence or letting their initial hunches cloud their thinking.
The study also shed light on why this might have happened. The students who used the AI cases reported feeling less mental strain from the way the material was presented, and they felt more capable of building a deep understanding of the subject matter. This aligns with a theory of learning which suggests that when students are presented with a variety of different scenarios, their brains are forced to work harder to find the underlying patterns, rather than just memorizing a single, predictable story. The artificial intelligence was able to generate these varied scenarios effortlessly. For every disease, the AI created three different versions: one that looked typical, one with unusual symptoms, and one complicated by other health issues. This variety forced the students to think more flexibly. In contrast, the human-written cases, while accurate, tended to follow a more standard pattern, which may have allowed students to rely on shortcuts in their thinking.
Efficiency was another major finding. The researchers found that generating these cases with the help of the artificial intelligence was dramatically faster than writing them by hand. While a human expert might spend four to eight hours crafting a single detailed case, the AI could produce a draft in moments. The human faculty still played a crucial role, spending about twelve minutes reviewing each AI-generated story to ensure it was medically sound and free of dangerous errors. One case was even rejected because the AI suggested a medication that was unsafe for the specific patient described. This review process meant that the department saved roughly one hundred and sixty hours of faculty time over the course of a year, time that could be redirected toward teaching and mentoring.
The researchers were careful to note that this success depended on having human experts in the loop. The artificial intelligence is a powerful tool for generating volume and variety, but it is not perfect. It can make mistakes, such as suggesting the wrong medication, which is why a human must always check the work before it reaches a student. The study also acknowledged that the students who used the AI cases might have been motivated simply because the material was new and novel, a factor that could have boosted their performance. Furthermore, the study was conducted at a single hospital with a specific group of students, so the results might look different in other settings or with different groups of learners.
Despite these limitations, the study provides strong evidence that artificial intelligence can be used to enhance medical education in a meaningful way. It suggests that by using AI to create a wider, more diverse library of patient stories, educators can help students develop the kind of flexible, robust thinking required for surgery. The technology does not replace the teacher; rather, it frees the teacher from the drudgery of writing endless case files so they can focus on guiding the students' minds. The ultimate goal is not to have students memorize the stories the computer writes, but to train them to reason through the uncertainty of real life, where every patient is a unique puzzle that cannot be solved by a simple formula.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.