← Latest papers
📄 medicine

High Overall Performance but Limited Multidisciplinary Depth: A Cross- Sectional Evaluation of ChatGPT in Pathological Fractures of Plasma Cell Tumors

This cross-sectional study found that while ChatGPT generates clear, generally accurate, and patient-friendly responses regarding pathological fractures from plasma cell tumors, it demonstrates limited depth and consistency in complex multidisciplinary clinical scenarios.

Original authors: MÜJDAT ADAŞ, AHMET MURAT ÇÖREKCİ, İSMAİL DEMİRKALE, TANJU BERBER, EMRE UYSAL, EMRACAN AKDEMİR, FURKAN BARIŞ, Berna Akkus Yildirim

Published 2026-07-15
📖 1 min read☕ Coffee break read

Original authors: MÜJDAT ADAŞ, AHMET MURAT ÇÖREKCİ, İSMAİL DEMİRKALE, TANJU BERBER, EMRE UYSAL, EMRACAN AKDEMİR, FURKAN BARIŞ, Berna Akkus Yildirim

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Cross-Sectional Evaluation of ChatGPT in Pathological Fractures of Plasma Cell Tumors

Problem Statement
Large language models (LLMs) are increasingly utilized by patients to access medical information; however, their performance in complex, multidisciplinary oncologic scenarios remains under-evaluated. While prior studies have assessed AI in single-specialty frameworks, there is a gap in understanding how these models handle the coordinated decision-making required for pathological fractures associated with plasma cell tumors (e.g., solitary plasmacytoma and multiple myeloma). These conditions necessitate intricate collaboration between orthopedic surgeons and radiation oncologists regarding surgical stabilization, radiotherapy timing, and treatment sequencing. The study addresses the need to determine whether AI-generated responses can accurately and consistently address patient-oriented questions across these divergent clinical domains.

Methodology
This cross-sectional study evaluated the performance of ChatGPT-5.1 (OpenAI) using a structured framework:

  • Question Generation: A multidisciplinary panel of four senior specialists (two orthopedic surgeons and two radiation oncologists, each with >20 years of experience) developed 25 patient-oriented questions. These were derived from real-world outpatient encounters and tumor board discussions, categorized into five domains: (1) Diagnosis and Understanding, (2) Surgical Stabilization, (3) Radiation Therapy, (4) Treatment Sequencing, and (5) Recovery and Follow-up.
  • Response Generation: Questions were submitted individually to the AI via a standardized neutral prompt ("Please answer the following question in a way that a patient can easily understand") in separate temporary chat sessions to minimize memory effects. Only the first generated response was recorded.
  • Evaluation: Four expert physicians (the same panel) independently evaluated the responses using a 5-point Likert scale across five criteria: Scientific Accuracy, Information Sufficiency, Patient Comprehensibility, Clinical Relevance, and Currency of Information. Evaluators were blinded to each other's assessments.
  • Statistical Analysis: Inter-rater reliability was assessed using Intraclass Correlation Coefficients (ICC) based on a two-way random-effects model. Differences between evaluators and criteria were analyzed using the Friedman test. Readability was assessed using Flesch–Kincaid metrics.

Key Results

  • Overall Performance: The AI achieved a high mean overall score of 21.90 ± 0.83 out of 25, indicating generally high performance in accuracy and clarity.
  • Domain Variability: Performance was highest in "Diagnosis and Understanding" (22.55 ± 0.62) and "Radiation Therapy" (22.15 ± 0.74). Slightly lower scores were observed in multidisciplinary domains requiring complex reasoning, specifically "Surgical Stabilization" (21.60 ± 0.72) and "Treatment Sequencing" (21.65 ± 0.42).
  • Evaluation Criteria: "Patient Comprehensibility" consistently received the highest scores (up to 4.80 ± 0.21), whereas "Information Sufficiency" was consistently the lowest-rated criterion (mean ≈ 3.95), particularly in surgical and sequencing contexts. This suggests responses are clear but may lack technical depth.
  • Inter-Rater Reliability: Agreement among expert evaluators was low overall, with ICC values ranging from negative to moderate. Notably, negative ICC values were observed in "Surgical Stabilization" and "Treatment Sequencing," indicating that evaluators' judgments diverged significantly, often exceeding chance expectations.
  • Readability: The Flesch–Kincaid Grade Level was calculated at 8.05, corresponding to a 10th–12th-grade reading level.

Key Contributions

  1. Multidisciplinary Framework: Unlike previous single-specialty evaluations, this study explicitly tested AI performance across the interface of orthopedic surgery and radiation oncology, highlighting how disciplinary expectations influence the assessment of AI output.
  2. Identification of Variability as a Feature: The study reframes the observed low inter-rater reliability not merely as a statistical flaw, but as a reflection of the inherent complexity and context-dependence of multidisciplinary clinical judgment. The divergence in scoring underscores that "clinical adequacy" is interpreted differently by different specialties.
  3. Differentiation of Capabilities: The research delineates a clear distinction between the AI's strength in synthesizing general, patient-friendly information and its relative weakness in providing the nuanced, actionable depth required for complex treatment sequencing and surgical decision-making.

Significance and Claims
The paper concludes that while ChatGPT generates clear, generally accurate, and patient-friendly responses suitable for general education and conceptual explanation, it currently functions as a "generalist explainer" rather than a "specialist-level decision tool." The authors assert that in the context of pathological fractures in plasma cell tumors, AI should be viewed as an adjunctive informational tool to support patient literacy, but it is insufficient for replacing expert clinical judgment in complex multidisciplinary scenarios. The study emphasizes that the variability in expert evaluation reflects the real-world complexity of multidisciplinary care, suggesting that future AI integration must account for these specialty-specific nuances rather than relying on a single composite score of performance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →