Trust gap in clinical artificial intelligence: a meta-systematic review
This meta-systematic review of 130 clinical AI studies reveals a significant "trust gap" where technical performance is prioritized over meaningful ethical integration, as evidenced by a 21.4% alignment between claimed and actual engagement with trust-related considerations, prompting the proposal of a new multidimensional framework for more responsible AI implementation.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A House Built on a Shaky Foundation
Imagine the field of medical Artificial Intelligence (AI) as a massive construction project. For the last few years, engineers have been building incredibly tall, fancy towers (the AI systems) to help doctors diagnose diseases and plan surgeries. They have spent billions of hours making sure these towers are structurally sound, measuring their height, and checking if they can hold weight.
However, this paper argues that while the engineers are obsessed with the height and strength of the towers, they have completely forgotten to build a foundation or a doorbell so people can actually trust them.
The author, Asefeh Asemi, looked at 130 different "reports on reports" (called meta-systematic reviews) published between 2022 and 2025. These reports were supposed to be the final inspections of the AI towers. The big discovery? The inspectors are only checking the bricks, not the building's safety for the people living inside.
The Main Problem: "Ethics-Washing"
The paper introduces a concept called "Ethics-Washing."
Think of this like a car salesman who puts a big, shiny sticker on a car that says "100% Safe Family Vehicle." But when you look under the hood, the brakes are missing, and the seatbelts don't work. The salesman talks about safety, but they don't actually build safety into the car.
The study found that in the world of Clinical AI:
- The Talk: Many reports mention words like "ethics," "trust," or "fairness" in their summaries (like the shiny sticker).
- The Reality: When the author looked closely at the actual content of these reports, almost none of them actually explained how to measure trust or how to fix unfairness.
- The Result: They found a "Trust Gap." Out of 130 reviews, zero of them actually defined what "trust" means in a real, multi-dimensional way. They just assumed that if the AI is accurate, it is trustworthy.
The Four Missing Pillars of Trust
The paper argues that "Trust" isn't just one thing (like accuracy). It's like a table that needs four legs to stand up. Currently, the AI field is trying to balance the table on just one leg.
- The Technical Leg (The Only One Standing): This is about accuracy, speed, and math. The paper says 89% of the reviews focus only on this. They ask, "Does the AI get the diagnosis right?"
- The Interpersonal Leg (Missing): This is about the relationship between the doctor, the patient, and the machine. Does the doctor feel comfortable using it? Does the patient feel safe? The paper found almost no one is talking about this (only 2.3% of reviews).
- The Institutional Leg (Wobbly): This is about rules, laws, and who is responsible if the AI makes a mistake. If the AI kills a patient, who goes to jail? The doctor? The programmer? The hospital? The paper found this is barely discussed (18.5%).
- The Epistemic Leg (Completely Gone): This is about knowledge. How does the AI "know" what it knows? If the AI gives an answer, do we understand why? The paper found that 99% of the time, this is ignored.
The Danger Zone: High-Risk Areas
The most alarming part of the study is where the lack of trust is happening.
Imagine a hospital. The most dangerous places are the Emergency Room, the Operating Room, and the Cardiology Unit. These are places where a mistake can kill someone instantly.
The study found that the AI systems used in these high-risk areas are the ones getting the most attention for their technical performance. However, they are also the areas where the least amount of ethical checking is happening.
- Analogy: It's like giving a race car to a driver in a crowded city. The car is incredibly fast (high technical performance), but nobody checked if the driver has a license, if the car has seatbelts, or if the roads are safe for pedestrians. The faster the car goes, the more dangerous the crash will be.
The "Silo" Effect: People Not Talking to Each Other
The paper also noticed that the people building these systems aren't talking to the people who should be checking them.
- The Technicians are writing reports about math and code.
- The Doctors are writing reports about patient care.
- The Ethicists are writing reports about right and wrong.
They are all in separate rooms (silos). Very few reports try to bring all three groups together. The paper found that only 6.9% of the reviews were truly "interdisciplinary" (mixing all three groups). This means we are building tools that might be mathematically perfect but useless or dangerous in a real hospital.
The Solution: The CAIEE Framework
To fix this, the author proposes a new blueprint called the CAIEE Framework (Clinical AI Ethical Engagement).
Think of this as a new Inspection Checklist for building AI hospitals. Instead of just checking the height of the tower, the checklist forces builders to check:
- Technical Trust: Is it accurate?
- Interpersonal Trust: Do doctors and patients feel safe using it?
- Institutional Trust: Are the rules and responsibilities clear?
- Epistemic Trust: Do we understand how it thinks?
The Bottom Line
The paper concludes that the medical AI field is suffering from an "Authenticity Crisis." We are saying we care about ethics and trust, but our actions (and our research reports) show we only care about speed and accuracy.
Until we start building the other three legs of the trust table, the AI systems in our hospitals might be very smart, but they won't be safe, fair, or truly trusted by the people who need them most. The paper urges researchers to stop just "talking" about ethics and start actually measuring and building it into the systems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.