← Latest papers
📊 statistics

On Using Large Language Models to Enhance Clinically-Driven Missing Data Recovery Algorithms in Electronic Health Records

This study demonstrates that combining large language models with clinical expertise to refine ICD-10-based "roadmap" algorithms enables the accurate and scalable recovery of missing Electronic Health Record data, achieving performance comparable to traditional, resource-intensive expert chart reviews.

Original authors: Sarah C. Lotspeich, Abbey Collins, Brian J. Wells, Ashish K. Khanna, Joseph Rigdon, Lucy D'Agostino McGowan

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Sarah C. Lotspeich, Abbey Collins, Brian J. Wells, Ashish K. Khanna, Joseph Rigdon, Lucy D'Agostino McGowan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Fixing the "Missing Puzzle Pieces"

Imagine a hospital's computer system (Electronic Health Records, or EHR) as a giant, massive jigsaw puzzle. Each patient is a picture, and the pieces are their health data: blood pressure, cholesterol levels, weight, and so on.

The problem is that many pieces are missing. Sometimes a doctor forgets to order a test, or the data didn't get saved correctly. When pieces are missing, the picture of the patient's health is incomplete, making it hard to see the full story or predict future risks.

In the past, to find these missing pieces, a team of human experts (like medical detectives) had to manually dig through thousands of paper and digital charts. They used a "roadmap" (a guide) that said, "If you see a diagnosis of 'Diabetes,' it's a safe bet that the missing blood sugar test was high."

While accurate, this manual detective work is incredibly slow, expensive, and can only be done for a tiny number of patients.

The New Idea: Teaching a Robot to Read the Map

This paper asks: Can we teach a computer (specifically a Large Language Model, or LLM) to do this detective work as well as the humans, but much faster?

The researchers wanted to see if they could use an AI to automatically fill in the missing health data for 1,000 patients, using the same logic the human experts used.

How They Did It: The "Roadmap" Upgrade

The team tried four different versions of their "roadmap" to see which worked best:

  1. The Human Map (Original): The original guide created by doctors. It had about 20 specific search terms (like "Diabetes" or "Infection").
  2. The AI Map (No Context): They asked the AI, "Here is a list of health markers. Please guess what other medical diagnoses might be related to them." The AI made a huge list of 656 terms.
    • The Result: The AI went wild and suggested too many things. It was like a detective guessing that "eating a sandwich" might mean "high blood pressure." It found very few actual matches.
  3. The AI Map (With Context): They tried again, but this time they showed the AI the original human map first and said, "Here is what the experts use. Now, expand on that list."
    • The Result: The AI was much smarter. It generated 950 terms, but they were much more relevant. It found many more missing pieces than the first attempt.
  4. The Hybrid Map (AI + Human Review): The AI generated the big list, and then two human doctors quickly reviewed it to say, "Yes, this makes sense," or "No, that's a false alarm."
    • The Result: This was the winner. It combined the speed and creativity of the AI with the safety check of the doctors.

What They Found

When they tested these methods on 100 patients (where they knew the "real" answer because humans had already checked the charts):

  • The Human Experts found about 45 missing pieces.
  • The AI (with human review) found almost the exact same number (45 pieces).
  • The AI (without human review) found fewer pieces or sometimes guessed wrong.

The Big Win: The best AI method didn't just match the humans; it could be applied to all 1,000 patients instantly. The human team could only check 100 patients because it took them months of hard work. The AI did the same job for the whole group in a fraction of the time.

The "Allostatic Load" (The Health Score)

The specific health puzzle they were solving was called the Allostatic Load Index (ALI). Think of this as a "wear and tear" score for the body. It adds up stress on the heart, metabolism, and inflammation systems.

  • Before: Because so many pieces were missing, the "wear and tear" score was often incomplete or looked lower than it actually was.
  • After: By using the AI to fill in the missing pieces, the researchers could calculate a more complete and accurate score for every single patient in the study.

The Bottom Line

This paper proves that we don't have to choose between accuracy and speed.

  • Old Way: Accurate but slow (Human detectives).
  • New Way: Accurate and fast (AI detectives guided by human experts).

By letting an AI expand the list of clues and having a human quickly double-check the results, hospitals can now fix missing data for thousands of patients at once. This makes the health data much more reliable for research and helps doctors get a clearer picture of patient health without spending years digging through charts.

What the paper does NOT claim:
The paper does not say this AI should be used to diagnose patients in real-time or replace doctors. It specifically focuses on cleaning up data for research and improving the quality of the "big picture" health data available in the system. It also notes that while the AI is great at finding missing pieces, it still needs human oversight to ensure it doesn't make up false connections.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →