Performance of Large Language Model on Real-World Patient Histories for Imaging Appropriateness.
This study demonstrates that ChatGPT achieves moderate agreement with radiologists in classifying the appropriateness of lumbar spine MRIs based on ACR criteria, with its decision reliability significantly increasing when its reasoning quality is rated highly.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Performance of Large Language Model on Real-World Patient Histories for Imaging Appropriateness
Problem Statement
Low back pain (LBP) is a prevalent global musculoskeletal complaint, leading to frequent requests for lumbosacral spine magnetic resonance imaging (MRI). Systematic reviews indicate that approximately 29–34% of these imaging requests are inappropriate, often lacking "red flags" such as infection, malignancy, or trauma. Unnecessary MRIs increase healthcare costs, patient anxiety, and the risk of overdiagnosis. While clinical decision support tools and guidelines exist, their real-world implementation is hindered by workflow integration challenges and alert fatigue. Although Large Language Models (LLMs) have emerged as a scalable alternative for reducing unnecessary utilization, prior studies have largely relied on hypothetical clinical scenarios rather than real patient histories, utilized single-radiologist reference standards, and focused on earlier model versions (e.g., GPT-4) without systematically evaluating the link between reasoning quality and decision reliability.
Methodology
This cross-sectional study evaluated the diagnostic accuracy of GPT-5.2 in classifying the appropriateness of lumbosacral spine MRIs using real-world patient histories and American College of Radiology (ACR) criteria.
- Data Collection: Data were collected retrospectively from 160 adult patients (mean age 48 ± 15 years; 58% male) who underwent lumbosacral MRI between September and November 2025. Patients with insufficient history, prior spinal surgery, suspected infection, or acute trauma were excluded. A structured questionnaire, validated by five specialists, was used to gather pre-imaging clinical histories.
- Model Prompting: Patient histories were fed into GPT-5.2 via structured prompts assigning the model a radiologist role and requiring it to classify MRI indications as "Usually Appropriate" (UA), "May Be Appropriate" (MBA), "Usually Not Appropriate" (UNA), or "Insufficient Clinical Information" (ICI) based on ACR guidelines, including explicit reasoning.
- Reference Standard: Two independent radiologists (with 5 and 10 years of experience) served as the reference standard.
- Phase 1: Radiologists assessed appropriateness based on the same prompts provided to the LLM, blinded to the model's output.
- Phase 2: Radiologists evaluated the accuracy of the LLM's recommendations and the quality of its clinical reasoning using a 5-point Likert scale (1–5), blinded to each other's ratings.
- Statistical Analysis: Agreement was measured using weighted Cohen's kappa () and Bowker's test of asymmetry. For binary analysis, UA and MBA were grouped as "Appropriate," and UNA as "Inappropriate." Performance metrics included Sensitivity, Specificity, PPV, NPV, and Area Under the Curve (AUC). Subgroup analyses were conducted based on radiologist-rated reasoning quality (High: 4 vs. Low: <4).
Key Results
- Agreement: GPT-5.2 demonstrated moderate agreement with both radiologists ( vs. Radiologist 1; vs. Radiologist 2). This level of agreement was comparable to the inter-radiologist agreement ().
- Binary Classification: The model showed high sensitivity (~85% against both radiologists) but variable specificity (92.9% vs. Radiologist 1; 61.7% vs. Radiologist 2). The Negative Predictive Value (NPV) was moderate (57.8% vs. R1; 64.4% vs. R2), indicating a tendency toward conservative (restrictive) MRI use.
- Discriminatory Power: ROC analysis showed good to excellent discrimination against Radiologist 1 (AUC 0.891) and fair to good discrimination against Radiologist 2 (AUC 0.751).
- Reasoning Quality Correlation: A critical finding was the strong association between reasoning quality and decision reliability. When radiologists rated the model's reasoning as high (4), exact agreement on appropriateness was high (83.3% vs. R1; 90.4% vs. R2), and opposite discordance (the most dangerous error type) was low (2.3–4.7%). Conversely, when reasoning was rated low (<4), exact agreement dropped to 0–7.1%, and opposite discordance surged to 71.4–92.8%.
- Expert Evaluation: Radiologists rated the model's output quality highly, with median Likert scores of 5 (maximum) for both appropriateness and reasoning.
Significance and Claims
The authors claim that GPT-5.2 performs within the range of inter-radiologist variability when classifying lumbosacral MRI appropriateness using real patient histories. The study highlights that the quality of the model's reasoning is a robust predictor of its decision reliability. Specifically, high-quality reasoning outputs correlate with high concordance with expert judgment, while low-quality reasoning correlates with significant discordance.
The paper posits that LLMs like GPT-5.2 have the potential to serve as pre-screening support tools to reduce unnecessary imaging requests. However, the authors emphasize that this potential is contingent on the model's reasoning quality, suggesting a hybrid workflow where low-reasoning outputs are escalated for expert review. The study concludes that while the model shows promise, safe clinical integration requires continued human oversight, prospective multi-center validation, and the establishment of consensus reference standards to address inter-observer variability in ambiguous cases.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.