From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
This paper proposes an evidence-grounded LLM framework that shifts the focus from generating safety text to providing source-linked safety knowledge support, thereby addressing the limitations of current models in regulated medical device development by integrating artifact tracing, uncertainty checks, and expert review to assist with ISO 14971 and IEC 62304 compliance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where the things we rely on most—like the pacemaker keeping a heart beating or the insulin pump managing diabetes—aren't just mechanical boxes anymore. They are becoming like super-smart, connected smartphones that can think, learn, and talk to the cloud. This is the exciting reality of modern medical devices. But with great intelligence comes great responsibility, and in the world of medicine, that responsibility is called "safety." Before a device can touch a patient, engineers have to prove it won't cause harm. They do this by writing mountains of documents, creating "safety cases" that act like a giant puzzle. Every piece of the puzzle—a design choice, a software update, a complaint from a user—must be linked perfectly to show that the device is safe. If one piece is missing or doesn't fit, the whole picture is broken, and the device can't be used.
The challenge is that this puzzle is incredibly hard to solve. It requires rare experts who know both medicine and complex engineering, and they have to keep the puzzle pieces aligned as the device changes over time. It's slow, expensive, and exhausting. Enter Large Language Models (LLMs), the same kind of AI that can write poems, answer trivia, and chat about almost anything. Scientists have been wondering: "Can these AI chatbots help us build the safety puzzle faster?" The idea is that since safety work is mostly about reading and writing documents, maybe an AI could draft the safety notes for us. However, there's a catch. In a hospital or a factory, you can't just trust a chatbot to say, "I think this is safe." If the AI makes a mistake or guesses, people could get hurt. So, the big question isn't just "Can the AI write?" but "Can the AI write and prove exactly where it got its information, while a human expert double-checks everything?"
This is exactly what the paper "From Safety Documentation to Safety Knowledge" tackles. The authors, a team of researchers from Fraunhofer IESE, argue that we need to stop thinking of AI as a magic writer that spits out final safety reports. Instead, they propose a new way to use AI as a super-powered research assistant that never forgets its sources. They suggest a framework where the AI doesn't decide if a device is safe; it only prepares "candidate" safety notes that are tightly linked to the original evidence, like a student citing their textbook sources.
The paper finds that while current AI tools are good at generating text, they often fail in the strict world of medical devices because they lack "source linking." An AI might invent a plausible-sounding risk that doesn't actually exist in the device's design, or it might miss a rare danger because it's relying on general knowledge instead of the specific, private documents of that one device. The authors explicitly rule out the idea that AI can replace human safety engineers. They state clearly that the AI should never be the one to give the final "thumbs up" for a device's safety.
To fix this, the team proposes a step-by-step process. First, the AI gathers all the device's documents (requirements, design files, past complaints). Then, it uses these documents to draft safety items, but every single draft must come with a "receipt" showing exactly which document it came from. Next, the AI acts as a critic, checking its own work for duplicates or weak links. Finally, a human expert reviews the AI's work, accepting, rejecting, or editing the drafts, and recording their reasons. This creates a "source-linked safety item"—a piece of safety knowledge that is traceable, updatable, and always ready for a human to verify.
The paper suggests that this approach could make safety work faster and more consistent, especially when devices get updated or when new complaints come in. However, the authors are careful to note that this is a proposal and a framework, not a finished product that has been fully tested in a real hospital yet. They outline a plan to test this idea using new, private medical device cases to see if it actually helps experts find more risks and spend less time searching for documents. They don't claim to have solved the problem of medical safety; instead, they offer a new, safer way to use AI as a tool in the toolbox, ensuring that the final decision always remains in human hands.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.