← Latest papers
💻 computer science

Lessons Learned and Iterative Enhancements: A 2-Year Experience with an Established Responsible AI Framework

Over a two-year period, UNC Health enhanced its Responsible AI framework by implementing a three-tier risk-based system that significantly improved evaluation efficiency and reduced lead times while maintaining rigorous oversight, resulting in the successful review of 64 solutions and the establishment of robust post-deployment monitoring mechanisms with minimal adverse events.

Original authors: Torre Caparatta, Ada H. Tsoi, Darius K. Byramji, Shahryar Farooq, Kathryn Ruiz, Cassiopeia Frank, Gary Gartner, Noah Schwarz, Alexander Fenn, John P. Nazarian, Murotiwamambo Mudziviri, Steven W. Cotte
Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: Torre Caparatta, Ada H. Tsoi, Darius K. Byramji, Shahryar Farooq, Kathryn Ruiz, Cassiopeia Frank, Gary Gartner, Noah Schwarz, Alexander Fenn, John P. Nazarian, Murotiwamambo Mudziviri, Steven W. Cotten, Cheri Warren, Lisa M. Hall, Radhika Talwani Bombard, Kenny B. Taylor, Steven David McSwain, Ram Rimal

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine UNC Health as a massive, busy hospital system that suddenly found itself flooded with requests to use new "smart" computer programs (Artificial Intelligence) to help doctors and staff. Two years ago, they built a "gatekeeper" system to check if these programs were safe and fair. But as the flood of requests grew, the gatekeepers were drowning.

This paper is the story of how they upgraded their gatekeeping system to handle the rush without letting dangerous tools slip through. Here is how they did it, explained simply:

1. The "Traffic Light" System (Risk Tiers)

Before, every single request, no matter how small or big, had to go through the same long, complicated inspection line. It was like making a person buying a candy bar wait in the same security line as someone importing a nuclear reactor.

The Fix: They created a three-tier traffic light system:

  • Green Light (Tier 0): Low-risk tools. These are like digital calculators or tools that don't look at patient secrets. They get a quick "pass" with almost no waiting.
  • Yellow Light (Tier 1): Medium-risk tools. These might look at patient data but only give advice (like a suggestion), and a human must always double-check the work. These get a standard inspection.
  • Red Light (Tier 2): High-risk tools. These are the heavy hitters—tools that might summarize patient notes, make diagnoses, or work without a human watching every second. These get the "full body scan" inspection, involving many experts and strict tests.

The Result: This sorting system was like adding a fast lane to a highway. The time it took to approve low-risk tools dropped from 20 weeks to just 9 weeks.

2. The "Vendor Report Card" (The Survey)

To check the tools, the hospital sends a 25-question survey to the companies that make the AI. Think of this as a report card asking: "Are you fair? Are you transparent? Did you test this properly?"

The Problem: The companies (vendors) were often bad at filling out the report card.

  • The "No Risk" Lie: Many vendors just wrote "No risk" or "We have a human watching it" without explaining how. It's like a driver saying, "I'm safe because I have a seatbelt," but refusing to show you the car has brakes.
  • The "Copy-Paste" Issue: Some answers looked like they were copied from marketing brochures rather than written by the engineers who actually built the code.
  • The "Black Box" Mystery: Many tools were built on top of massive, pre-made AI brains (called Large Language Models). The vendors didn't build the brain; they just wrapped it in a new package. They often couldn't explain how the brain was trained, making it hard to check for fairness.

The Result: The hospital updated the survey questions three times to make them clearer. Even with better questions, the "Fairness" score was consistently the lowest, meaning vendors still struggle to prove their tools treat everyone equally.

3. The "Safety Net" (Post-Deployment)

Even after a tool is approved and put to work, the hospital knows mistakes can happen. They needed a way to catch errors after the tool is in the wild.

The Fix: They connected the AI tools to two existing safety nets:

  • The "Oops" Button: If a doctor sees the AI make a weird mistake (like writing a fake medical note), they can flag it immediately in a safety reporting system.
  • The "Complaint" Line: Patients can call or email if they are worried about how AI is being used in their care.

The Result: In the first six months of this new system, they caught 13 safety events. Interestingly, almost all of them were "hallucinations" (the AI making up facts) in medical notes. None of these events caused serious harm, and zero patients complained through the new phone line. This suggests the safety nets are working, but it also shows that most AI errors so far have been minor glitches rather than disasters.

4. The "Human-in-the-Loop" Trap

A major lesson learned is that you can't just say, "We put a human in the loop, so it's safe."

  • The Analogy: Imagine a self-driving car that says, "Don't worry, the driver is watching." But if the driver is tired, distracted, or trusts the car too much, they might not hit the brakes when the car fails.
  • The Reality: The hospital found that many vendors used "Human-in-the-Loop" as a magic excuse to avoid proving their AI was actually safe. They learned that humans often don't catch AI mistakes because they trust the computer too much (a problem called "automation bias").

5. The Future: "Agentic" AI

The paper ends by looking at the next generation of AI: Agentic Systems.

  • The Analogy: Current AI is like a calculator that gives you an answer. "Agentic" AI is like a robot butler that doesn't just give an answer, but goes out and does things (like booking an appointment, ordering supplies, or changing a record) based on a goal.
  • The Risk: If a calculator is wrong, you get one wrong number. If a robot butler is wrong, it might make a chain of bad decisions. The hospital admits their current rules aren't quite ready to handle these "doers" yet and will need to evolve.

The Bottom Line

UNC Health successfully sped up their AI approval process by sorting tools into risk levels and tightening their questionnaires. However, the biggest bottleneck remains the vendors themselves. If the companies making the AI won't be honest about risks or show how they tested their tools, the hospital's best rules can only do so much. The goal is to build a system where the "right" action (safe, fair care) is the easiest path for everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →