← Latest papers
📄 medicine

Blinded Expert Evaluation of Large Language Models for Osteoporosis Case Management in Older Adults: Impact of Direct Guideline Integration

This study found that while large language models can support geriatric osteoporosis management, the base ChatGPT 4.5 model outperformed guideline-integrated versions in accuracy and adherence, yet all models exhibited limitations in safety awareness and context-sensitive reasoning, underscoring the continued necessity for expert clinical supervision.

Original authors: Gokalp Kurthan AVLAGI, Seyda BILGIN, Duygu OZATA, Kubra CINGAR ALPAY, Tugba KANDEMIR, Tugce EMIROGLU GEDIK, Denız SUNA ERDINCLER

Published 2026-07-07
📖 4 min read☕ Coffee break read

Original authors: Gokalp Kurthan AVLAGI, Seyda BILGIN, Duygu OZATA, Kubra CINGAR ALPAY, Tugba KANDEMIR, Tugce EMIROGLU GEDIK, Denız SUNA ERDINCLER

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have four different "digital doctors" trying to solve a complex puzzle: how to treat osteoporosis (weak bones) in older adults. This isn't a simple puzzle; it involves checking for fall risks, managing multiple medications, watching kidney function, and following strict medical rules.

The researchers in this study wanted to see which digital doctor was the best at solving these puzzles and if giving them a "rulebook" (medical guidelines) helped them do better.

Here is the story of what they found, explained simply:

The Contestants

The study set up a blind taste test with four AI "chefs":

  1. The Base Chef (ChatGPT 4.5): A smart AI that knows a lot but didn't have the specific rulebook in front of it.
  2. The Rival Chef (Gemini 3): Another smart AI, also without the rulebook.
  3. The Guided Chef (ChatGPT 4.5 + Rulebook): The first chef, but this time they were handed the full 2025 Turkish Osteoporosis Guidelines to read before cooking.
  4. The Notebook Chef (NotebookLM + Rulebook): A different type of AI tool that was also handed the full rulebook.

The Twist: Two real-life human experts (geriatric specialists) tasted the dishes (read the answers) without knowing which chef made them. They scored the dishes on a scale of 1 to 5 for things like accuracy, clarity, safety, and whether they followed the rules.

The Results: Who Won?

Surprisingly, the Base Chef (ChatGPT 4.5 without the rulebook) won the contest with the highest score.

  • The "Rulebook" Didn't Help: You might think that giving a chef a specific recipe book would make the dish perfect. But in this case, handing the rulebook to the AI actually made the answers slightly worse or no better than the AI just using its own brain.
  • Why? The researchers suggest that dumping a huge, unorganized text document into the chat was like handing a chef a 500-page cookbook and saying, "Figure it out." The AI got confused about which parts of the book mattered for the specific puzzle, rather than using the book as a precise tool.
  • The Rival: The second chef (Gemini 3) did a good job, but the Base Chef was still the clear winner.

The Weak Spots

Even the winning chef had some blind spots. The study found that while the AI was great at reciting facts (like "what drug treats this?"), it struggled with the "human" side of the puzzle:

  • Safety Checks: All the AIs were a bit shaky when it came to spotting dangerous interactions or contraindications (things that could hurt the patient).
  • The "Holistic" View: They weren't great at looking at the whole picture of an older person's life (like their risk of falling or how many other pills they take).

It's like a very smart librarian who can instantly find the right book on a shelf but might forget to ask if the person reading it has a broken leg that makes walking to the shelf dangerous.

The "Hallucination" Check

The researchers also checked if the AIs made things up (called "hallucinations").

  • The good news: They almost never made things up.
  • The bad news: Even one made-up fact can be dangerous in medicine. One of the rival chefs made a small error in one case, but the others were clean.

The Bottom Line

The study concludes that these AI tools are like very knowledgeable assistants, but they are not doctors.

  • They can help organize information and suggest ideas.
  • However, they cannot replace a human doctor's supervision, especially when it comes to safety and looking at the whole patient.
  • Simply pasting a medical guideline into the chat doesn't make the AI smarter; it might just make it confused.

In short: The AI did a great job on the "textbook" questions, but it still needs a human expert to double-check the safety and the big picture before making any real decisions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →