Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
This study demonstrates that a human-in-the-loop AI agent can effectively automate the gathering and drafting of translational impact summaries for clinical scholars, reducing the reporting workload from an estimated 15 hours to 14 minutes per scholar while maintaining high accuracy and capturing non-scholarly evidence often missed by routine processes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of medical research as a giant, bustling library where scientists are constantly writing new books about how to cure diseases, help communities, and save money. But there's a catch: the librarians (the people who manage these research programs) have a massive, impossible job. They need to prove that every single scientist in their library is actually making a difference in the real world, not just writing fancy papers. Usually, to do this, a librarian has to spend about 15 hours digging through the internet, checking grant records, and reading news articles for each scientist to write a one-sentence summary of their good deeds. It's like trying to find a specific needle in a haystack, then writing a poem about that needle, all by hand. If a scientist moves to a new city or changes jobs, the librarian has to start the search all over again. It's slow, exhausting, and doesn't work when you have hundreds of scientists to track.
This is where a new kind of helper comes in: an AI agent. Think of this AI not as a magic robot that knows everything, but as a super-fast, tireless research intern. This intern can scan thousands of websites, databases, and news feeds in seconds, gathering clues about what a scientist has done. But here's the twist: the AI isn't allowed to just hand over the final report. Instead, it acts like a draftsperson. It gathers all the evidence and writes a rough draft of the "impact summary." Then, a human librarian steps in to check the work. The human doesn't have to start from scratch; they just have to read the AI's draft, fix a few mistakes, and say, "Yes, this is true," or "No, this is wrong." The big question this paper asks is: Can this human-and-AI team do the job faster and better than the human working alone?
The Story of the Super-Intern
In this study, researchers at the University of Illinois Chicago built this "super-intern" AI agent and put it to work in a real-world setting. They tested it on 10 scientists (called "scholars") who were part of a special career-development program. The goal was to see if the AI could gather the evidence and write the first draft of their impact summaries, saving the human staff from those dreaded 15-hour marathons.
The AI agent worked like a detective with a very specific set of tools. It started with a "seed" (just the scientist's name and where they work) and then went on a wild goose chase across the internet. It checked official medical databases, looked at grant records, scanned news sites, and even dug up policy documents. It didn't just stop at the easy stuff; it tried to find "non-scholarly" evidence, like how a scientist helped a local community or influenced a new law. Once it gathered all these clues, it organized them into a neat file (a "dossier") and wrote a one-sentence summary for each discovery, explaining exactly how it helped the world.
But the AI wasn't left to its own devices. Two human staff members, who are experts in this field, acted as the editors. They looked at every single thing the AI found. For each finding, they had three choices:
- Accept: "Perfect! This is exactly right."
- Edit: "Close, but fix this detail."
- Reject: "Nope, this is wrong or not important."
They also had the option to add things the AI missed, though the AI was designed to be so thorough that it rarely missed anything.
What They Found
The results were pretty exciting. The human editors found that 81.7% of the AI's findings were usable right away or needed only a tiny tweak. That means for every 100 things the AI found, the humans kept 82 of them. This is a huge deal because it means the AI did the heavy lifting of gathering the information.
Here is the most important part: Time.
- Before: A human staff member estimated it took 15 hours to build a scientist's impact record from scratch.
- After: With the AI doing the first draft, the human staff only needed a median of 14 minutes per scientist to review and fix the work.
That is a massive shift! The humans went from being "hunters and writers" to being "editors and reviewers." They didn't have to spend hours searching; they just spent minutes checking the AI's work.
The AI was also surprisingly good at finding the right person. In a side test, the AI's ability to find a scientist's main profile page was almost as good as a human searching manually. It found the right person about 82% of the time, just like a human would.
The Hiccups and the Rules
Of course, the AI wasn't perfect. About 18% of the time, the humans had to reject a finding. Why?
- Weak Sources: Sometimes the AI found a piece of information on a website that wasn't very trustworthy (like a random directory page instead of an official news site).
- Not an Impact: Sometimes the AI found something the scientist did, but it wasn't actually a "big impact" (like a scientist giving a talk, which is good, but maybe not a life-saving discovery).
- Wrong Person: Occasionally, the AI mixed up two people with similar names.
The humans had to use their judgment to decide if a finding was "good enough" to count. The paper notes that the AI was designed to be "generous"—it found too much so that the humans could throw away the bad stuff, rather than the AI missing the good stuff. This is a smart strategy because it's easier to edit a long list than to go hunting for missing pieces.
The Verdict
This paper suggests that a human-and-AI team can work together to solve a problem that was previously too slow to handle. The AI acts as the first-pass author, doing the boring, time-consuming search and drafting work. The humans then step in to verify the facts and make the final call.
The study didn't prove that AI is perfect or that it can replace humans entirely. In fact, the paper argues that human judgment is still absolutely necessary because the AI can make mistakes with sources or misunderstand what counts as a "real" impact. But the study does show that with this new division of labor, tracking the impact of a whole group of scientists becomes possible. Instead of spending 15 hours per person, staff can now do it in minutes, making it feasible to keep track of everyone's contributions to science and society. It's like upgrading from a bicycle to a bicycle with a motor: you still have to steer and pedal, but you get there much, much faster.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.