Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset
This study demonstrates that large language model-assisted label cleaning effectively identifies and resolves clinically meaningful discrepancies between radiology reports and existing labels in the CT-RATE chest CT dataset, achieving high agreement with radiologist adjudication and supporting scalable quality improvement for public imaging data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a massive library of medical textbooks (radiology reports) describing thousands of chest CT scans. Alongside each book, someone has already created a "cheat sheet" (a dataset label) summarizing what's wrong with the patient, like "pneumonia" or "enlarged lymph nodes."
The problem is that sometimes the cheat sheet doesn't quite match the story in the book. Maybe the book says, "We might see a small spot," but the cheat sheet boldly says, "Pneumonia present." In the world of Artificial Intelligence (AI), if you train a robot to learn from these cheat sheets, it will learn the wrong lessons.
This paper is about a team of researchers who used a super-smart AI tool (a Large Language Model, or LLM) to act as a proofreader for these cheat sheets. Here is how they did it, explained simply:
The Detective Work
The researchers took a huge public dataset called CT-RATE, which contains about 24,000 chest CT reports and their corresponding labels. They asked a very advanced AI (GPT-5.4) to read the text of the reports from scratch and write its own cheat sheet.
Think of it like this:
- The Original Cheat Sheet: A hurried student's notes.
- The AI Proofreader: A meticulous librarian who reads the original book and writes a new set of notes based only on what is explicitly written.
- The Comparison: The researchers then compared the student's notes with the librarian's notes to see where they disagreed.
What They Found
- Mostly Good, But Not Perfect: The AI proofreader agreed with the original notes 96.4% of the time. That's a high score, but in a library of this size, even a 3.6% error rate means thousands of mistakes.
- The Tricky Spot: One specific category, lymphadenopathy (swollen lymph nodes), was a mess. The agreement was very low. The researchers found that the original notes were often too vague or too strict. For example, the report might mention "tiny lymph nodes" (which are normal), but the original label marked it as "disease present." The AI proofreader correctly ignored these tiny, normal nodes.
- The Human Verdict: To settle the arguments, real doctors (radiologists) looked at the cases where the AI and the original notes disagreed.
- In general disagreements, the doctors sided with the AI proofreader 74% of the time.
- For the tricky lymph node cases, the doctors sided with the AI 92% of the time. This suggests the original labels were often wrong, and the AI caught them.
The "Team Vote" Strategy
The researchers didn't just rely on one AI. They tried a second and third AI model and asked them to vote.
- Imagine three judges in a talent show. If two out of three judges agree on a score, that score is likely more reliable than just one judge's opinion.
- This "majority vote" method created the most accurate set of labels, even better than the original dataset or any single AI.
The Big Takeaway
The main conclusion is that AI can be a powerful quality control tool for medical data.
Just as a spell-checker helps you find typos in an essay, this LLM-assisted method helps find "label typos" in massive medical datasets. By cleaning up these errors, the researchers created a "refined" version of the dataset that is more honest about what the reports actually say. They plan to share this cleaner version with the public so that other scientists can build better AI models without learning from bad data.
Important Note: The researchers are careful to say that their AI is reading the text of the reports, not looking at the actual X-ray images. So, they are checking if the labels match the story, not necessarily if the story matches the reality of the patient's body. However, for the purpose of fixing the dataset's internal consistency, this method worked very well.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.