← Latest papers
📄 medicine

Comparison of the Diagnostic Reliability of ChatGPT in Letournel–Judet Acetabulum Fracture Classification Based on Judet Radiographs According to Orthopaedic Residents

This study demonstrates that ChatGPT-4o exhibits fundamentally inadequate diagnostic reliability and negligible agreement with expert standards for classifying Letournel–Judet acetabular fractures, performing significantly worse than orthopaedic residents and thus failing as a viable decision-support tool for pelvic trauma.

Original authors: Muhammed Kilic, Ahmet Ozgur Yildirim, Ibrahim Alper Yavuz, Fatih Inci, Erman Ceyhan, Yakup Kahve, Tahsin Aydin, Ozdamar Fuad Oken

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Muhammed Kilic, Ahmet Ozgur Yildirim, Ibrahim Alper Yavuz, Fatih Inci, Erman Ceyhan, Yakup Kahve, Tahsin Aydin, Ozdamar Fuad Oken

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very complex, 3D puzzle made of broken bone inside a person's hip. To fix it, doctors need to figure out exactly how the pieces broke apart. They use a special map called the "Letournel–Judet classification" to describe the break. This map is tricky; it requires looking at three different flat X-ray pictures (like looking at a building from the front, the side, and the back) and mentally combining them to understand the whole structure.

This study asked a simple question: Can a super-smart computer chatbot (ChatGPT-4o) do this job as well as a human doctor in training?

Here is the breakdown of what happened, using simple analogies:

The Setup: The Test

The researchers took X-rays from 184 patients with broken hip sockets. They set up a "test" with three groups:

  1. The Expert: A seasoned trauma surgeon who saw the final 3D CT scans and the actual surgery results. This person is the "Answer Key."
  2. The Students: Two orthopedic residents (doctors in their final year of training). They only saw the flat X-rays, just like the computer.
  3. The AI: ChatGPT-4o. The researchers gave the AI the same X-rays and a specific list of questions (a checklist) to help it figure out the break type.

The Results: A Big Gap

The results were a massive shock to the researchers.

  • The Students (Residents): They were excellent. They got the diagnosis right 88% to 91% of the time. They agreed with the Expert almost perfectly.
  • The AI (ChatGPT): It struggled terribly. It only got the final diagnosis right 17.9% of the time. That means it was wrong in almost 8 out of 10 cases.

The Analogy:
Imagine a game where you have to identify a specific type of car crash by looking at three blurry photos.

  • The Residents are like experienced mechanics who looked at the photos and said, "That's a rear-end collision with a twisted frame." They were right almost every time.
  • The AI was like a student who looked at the photos and said, "That's a car that fell off a cliff," or "That's a car that hit a tree," even when it clearly wasn't. It was guessing wildly.

Where the AI Got Stuck

The study found that the AI wasn't completely blind.

  • The Good: The AI was actually pretty good at spotting simple, isolated things. It could correctly identify if a specific part of the hip bone (the "iliac wing") was broken about 82% of the time. It was like a student who can correctly point out "that's a broken wheel" but fails to understand the whole car.
  • The Bad: The AI failed miserably at the hard part: putting the pieces together. It couldn't figure out how the front and back parts of the hip were connected. It often confused a simple break with a "both-column" fracture (a very complex break), essentially seeing a simple scratch and calling it a total disaster.

The researchers noted that the AI seemed to suffer from "hallucinations." It would see a shadow on an X-ray and confidently say, "That is a specific sign of a fracture," even when it wasn't there.

The Conclusion: Don't Use It Yet

The paper's main takeaway is blunt: Do not use this AI to diagnose broken hips.

The authors state that while AI is great at simple tasks (like spotting a broken bone in a finger), it is currently fundamentally inadequate for the complex, multi-step thinking required for hip socket fractures. It cannot reliably combine the different X-ray views to create a correct diagnosis.

In short: If you handed these X-rays to a human resident, they would likely get the right answer. If you handed them to this specific AI, it would likely give you the wrong answer, which could lead to the wrong surgery. The researchers say we should stick to human doctors for this specific job until the AI gets much smarter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →