← Latest papers
📄 medicine

GPT-5 versus Senior Anesthesiology-Intensive Care Residents on Sequential Clinical Cases: A Pilot Study of Documented Clinical Reasoning

In a pilot study comparing documented clinical reasoning on sequential cases, GPT-5 achieved significantly higher R-IDEA scores than senior anesthesiology-intensive care residents, suggesting its potential utility as a supervised educational scaffold for reasoning documentation rather than as a replacement for clinical competence.

Original authors: Elmahdi Ezzikouri¹, Sabah Benhamza¹, Mohammed Bennani Othmani¹, Mohamed Lazraq¹

Published 2026-08-18
📖 1 min read☕ Coffee break read

Original authors: Elmahdi Ezzikouri¹, Sabah Benhamza¹, Mohammed Bennani Othmani¹, Mohamed Lazraq¹

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: GPT-5 versus Senior Anesthesiology-Intensive Care Residents on Sequential Clinical Cases

Problem Statement
While Large Language Models (LLMs) have demonstrated high accuracy on medical licensing examinations, performance metrics on multiple-choice tests do not elucidate how explicitly a response documents the underlying clinical reasoning process. Clinical reasoning is a context-dependent cognitive activity involving problem representation, hypothesis generation, and evidence interpretation. Existing literature often lacks rigorous empirical designs comparing LLMs to trainees on these specific reasoning documentation skills. Furthermore, there is a scarcity of evidence regarding LLM performance in French-language postgraduate training environments, particularly within North African medical contexts. This study addresses the gap between "correct answers" and "documented reasoning" by comparing the quality of written clinical reasoning outputs between a state-of-the-art LLM (GPT-5) and senior medical trainees.

Methodology
The study was a single-center, cross-sectional comparative pilot conducted at the Centre Hospitalier Universitaire Ibn Rochd in Casablanca, Morocco.

  • Participants: The study utilized an exhaustive census of 14 fourth-year anesthesiology-intensive care residents (100% participation).
  • Stimuli: Twenty real, fully anonymized clinical cases (2020–2024) were selected and standardized. Cases were divided into two sequential parts: Part 1 (history, background, physical exam) and Part 2 (laboratory and imaging results).
  • Task Allocation: Each resident completed four distinct cases (56 total resident responses). GPT-5 completed all 20 cases once. The design was an incomplete-block structure where every resident was matched with GPT-5 on the same four cases for primary analysis.
  • Prompt Engineering: GPT-5 was accessed via the OpenAI API (model alias gpt-5, temperature 0.3, reasoning effort "medium"). The prompt assigned the model the role of an anesthesiology-intensivist and explicitly requested a structured output mirroring the Revised-IDEA (R-IDEA) components: a conceptual checklist, a concise problem representation, prioritized investigations, and a differential diagnosis with explicit evidence for and against each hypothesis. The model was not provided with the reference diagnosis or scoring rubric.
  • Evaluation: Two faculty members developed case-specific reference rubrics based on the R-IDEA instrument. One masked evaluator scored all responses (residents and GPT-5) against these rubrics. The R-IDEA score (0–10) assesses four domains: Interpretive Summary, Differential Diagnosis, Explanation of Lead Diagnosis, and Explanation of Alternative Diagnoses.
  • Statistical Analysis: The primary analysis used a paired Student's t-test comparing the mean R-IDEA scores of each resident (across their four cases) against the mean score of GPT-5 on those same four cases. Diagnostic accuracy was reported descriptively.

Key Results

  • Primary Outcome: GPT-5 achieved a significantly higher mean total R-IDEA score (9.5, SD 0.9) compared to the residents (7.0, SD 2.4). The mean paired difference was 2.5 points in favor of GPT-5 (95% CI 1.4–3.6; p < 0.001). The effect size was large (paired d_z = 1.39).
  • Domain Performance: GPT-5 outperformed residents in all four R-IDEA domains. The largest absolute difference occurred in the "Interpretive Summary" domain (1.0 point difference). GPT-5 scores were concentrated near the maximum of the scale, suggesting a potential ceiling effect, whereas resident scores showed greater variability (Coefficient of Variation: 34% for residents vs. 9% for GPT-5).
  • Diagnostic Accuracy: GPT-5 identified the faculty reference diagnosis in 75% of cases (15/20), while the mean diagnostic accuracy for residents was 68%. The study design precluded inferential statistical comparison of these percentages due to non-independent observations and unequal case replication.

Key Contributions

  • Conceptual Replication in a New Context: This study extends previous findings (e.g., Cabral et al.) regarding generative AI and clinical reasoning to a French-language setting, a Moroccan academic center, and a senior anesthesiology cohort.
  • Differentiation of Documentation vs. Competence: The study explicitly distinguishes between the quality of documented reasoning (which GPT-5 excelled at under the specific prompt) and actual clinical competence or bedside performance.
  • Methodological Framework: It demonstrates a protocol for comparing LLM outputs against trainee written work using structured rubrics (R-IDEA) and source-masking, highlighting the importance of prompt alignment with assessment criteria.

Significance and Claims
The authors conclude that under a prompt aligned with R-IDEA components, GPT-5 produces written clinical reasoning documentation that scores higher than that of senior residents. However, the paper maintains a modest stance on the implications of this finding:

  • No Claim of Superior Competence: The authors explicitly state that higher documentation scores do not establish superior clinical competence, diagnostic equivalence, or educational effectiveness.
  • Educational Hypothesis: The primary significance lies in the potential for GPT-5 to serve as a "reasoning-documentation scaffold." The authors propose that faculty-supervised, vetted GPT-5 outputs could be used as worked examples to help residents compare their own problem representations against a structured model, thereby making reasoning more visible.
  • Limitations and Future Directions: The study acknowledges limitations including the single-center design, the use of a single rater, the lack of a second model run to assess stochastic variability, and the specific alignment of the prompt to the rubric which may have favored the model. The authors suggest that future research should involve prospective, multi-rater educational trials to test whether using model-generated worked examples actually improves learner outcomes without inducing automation bias.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →