Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices
This study reveals that FDA clearance of AI-enabled medical devices does not guarantee trustworthy AI, as a comprehensive analysis of 519 reports shows significant and persistent gaps in documenting key principles like explainability and traceability, indicating that regulatory approval alone is insufficient to validate system trustworthiness.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where your doctor has a super-smart digital assistant that helps diagnose diseases, spot tumors in X-rays, or predict heart attacks before they happen. This isn't science fiction; it's Artificial Intelligence (AI) in healthcare. But just like a new car needs a safety inspection before it hits the road, these digital doctors need a "license" to work. In the United States, the Food and Drug Administration (FDA) acts as the traffic cop, checking the AI's math and making sure it works correctly before giving it the green light.
However, getting a license is only half the story. When you buy a car, you get a manual that tells you how it handles in the rain, how it brakes on ice, and what happens if a tire blows out. For AI, we need a similar "user manual" that explains if the system is fair to everyone, if it can be trusted when things go wrong, and if we can understand why it made a specific decision. This idea is called "Trustworthy AI." It's not just about being accurate; it's about being safe, fair, explainable, and ready for the real world. The big question is: When the FDA says an AI is "cleared" to use, does the public paperwork actually tell us enough to trust it with our health?
The Great AI Paperwork Mystery
A team of curious researchers decided to play detective and crack open the filing cabinets of the FDA. They wanted to see if the official "summary reports" for AI medical devices actually contained the good stuff we need to trust them. They didn't just look at whether the AI got a passing grade; they looked at the report card itself to see if it mentioned the six golden rules of Trustworthy AI: Fairness (is it biased?), Universality (does it work for everyone?), Traceability (can we track its history?), Usability (is it easy for doctors to use?), Robustness (does it break under pressure?), and Explainability (can it tell us why it made a choice?).
They scanned through 1,105 FDA reports published between 2021 and 2025. After a rigorous filter to make sure they were looking at the right documents, they ended up with 519 reports to study. Think of this like opening 519 mystery boxes to see what's inside.
What They Found: The "Missing Manual" Problem
The results were a bit like opening a box of fancy new gadgets and finding that most of them came with no instructions at all.
- The Silent Majority: Nearly one out of every four reports (24.7%) didn't mention any of the six trustworthiness rules. It was as if the manufacturer said, "Here's a robot doctor," but didn't tell you if it was fair, safe, or understandable.
- The One-Note Symphony: Most of the reports that did have information only talked about one or two rules. Almost none of the reports covered all six. In fact, zero reports managed to document evidence for every single principle. It's like a restaurant that only tells you the food tastes good but refuses to say where the ingredients came from or if the kitchen is clean.
- The Star Performer: The only rule that got a lot of attention was Robustness. About 57.6% of the reports mentioned this. This is the "does it crash if the Wi-Fi is slow?" test. Manufacturers love to show off that their AI is tough.
- The Ghosts in the Machine: The two most important rules for understanding the AI were almost completely ignored. Traceability (knowing the AI's history) was mentioned in only 8.3% of reports, and Explainability (the AI explaining its reasoning) was the biggest gap of all, showing up in just 3.5% of reports. It's like having a judge who gives you a verdict but refuses to tell you the law they used to decide it.
Does Time or Location Matter?
You might think, "Maybe the reports are getting better as time goes on?" or "Maybe the heart doctors are better at this than the eye doctors?" The researchers checked these ideas, but the answer was a flat no.
- Time Travel Didn't Help: Whether the report was from 2021 or 2025, the quality of the trustworthiness info didn't really change. The year the device was cleared didn't predict better transparency.
- Specialty Didn't Save the Day: It didn't matter if the AI was for Radiology (X-rays), Cardiology (hearts), or Pathology (tissue samples). The gaps were everywhere. Even in fields where you'd expect high-tech transparency, the "why" and "how" were still missing.
The Big Takeaway
The paper concludes that getting an FDA "stamp of approval" is not enough to prove an AI is trustworthy. Just because a device is cleared to use doesn't mean the public paperwork gives us the full picture. The current system is like a car that passes the speed test but doesn't tell you if the brakes work or if the airbags are real.
The researchers suggest that we need a new kind of "report card" for AI. Instead of just listing the technical stats, manufacturers should be required to fill out a standardized form that covers all six rules of trustworthiness. They need to explain how they checked for bias, how they make sure the AI works for different people, and how doctors can understand the AI's decisions. Until then, the gap between what we expect from a trustworthy AI and what the paperwork actually shows remains wide open.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.