Drivers of Oncologist Preference of AI-Generated Literature Review in a Randomized Mixed-Methods Study
This randomized mixed-methods study reveals that oncologists' trust and preference for AI-generated literature reviews depend more on report design features—such as conciseness, verifiable citations, and explicit uncertainty—than on accuracy alone, as evidenced by significantly lower utility ratings for evidence-graded reports compared to standard formats despite similar reference quality.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the high-stakes world of cancer care, doctors must constantly navigate a shifting landscape of new treatments, evolving guidelines, and complex molecular data. To make the right choice for a patient, they often need to find and understand the latest medical research quickly. Recently, artificial intelligence tools have emerged to help with this task, acting as digital assistants that can search through vast libraries of medical literature and summarize findings in plain language. However, a critical question remains: just because an AI can find the right facts, does it present them in a way that a tired doctor can trust and use? The challenge is not merely about whether the computer is correct, but about how the information is dressed up and delivered. If a report is hard to read, confusing, or seems to hide its sources, a physician might ignore a perfectly accurate answer. This gap between raw data and human trust is where the real work of clinical artificial intelligence lies.
A team of researchers at Stanford University and other institutions set out to explore this gap by asking thirty-four oncologists to review reports generated by different artificial intelligence systems. They did not simply ask the doctors which AI was the smartest; instead, they tested how the design of the report influenced the doctors' confidence. The study involved five different patient scenarios, ranging from radiation oncology to pediatric hematology, covering a mix of training levels from residents to experienced faculty. For each scenario, the doctors reviewed four different AI-generated summaries of the same medical question. These summaries came from three distinct sources: a specialized medical AI tool called OpenEvidence, a general-purpose chatbot known as ChatGPT, and a version of OpenEvidence that the researchers had manually tweaked. The researchers had a specific hypothesis for this tweaked version: they believed that if they reorganized the report to highlight the strength of the evidence—putting the most reliable, large-scale clinical trials at the top and burying weaker studies at the bottom—doctors would find it more useful and trustworthy.
The results of this experiment surprised the researchers. Despite using the exact same underlying facts and citations, the modified report that the team had carefully crafted to emphasize evidence strength was rated significantly lower in overall usefulness than the standard, unmodified report from OpenEvidence. In fact, the customized version was the least preferred option among the four. The doctors did not find the reorganized layout clearer or more logical; instead, they found it less organized and harder to scan. This finding suggests that simply rearranging the order of information to make the "hierarchy of evidence" more visible did not work as intended. The doctors did not want the AI to act as a judge of which studies were best; they wanted the AI to present the information in a way that allowed them to verify it quickly for themselves.
Through a series of interviews and detailed ratings, the researchers uncovered what the doctors actually needed. The most valued features were not about complex ranking systems, but about clarity and speed. Doctors preferred reports that were concise and could be scanned in seconds, with a clear summary at the very top. They wanted to see specific numbers from clinical trials, such as survival rates or side-effect percentages, rather than vague descriptions. Crucially, they demanded that every claim be linked directly to its source with a clickable citation, allowing them to verify the information instantly. They also wanted a clear separation between the raw evidence and any recommendations the AI might make, preferring the AI to summarize the data without sounding overly confident or directive. When the AI buried its sources, used a chatty tone, or presented a mix of high-quality and low-quality studies without distinction, trust evaporated.
The study also revealed that the standard OpenEvidence tool performed just as well as the general-purpose ChatGPT in terms of overall utility, but it was rated higher for the quality of its citations. This indicates that for medical professionals, the ability to trace a claim back to its original source is just as important as the answer itself. The researchers found that the way information is presented can change how a doctor perceives its value, even if the content is identical. A report that looks messy or feels like it is hiding its work will be rejected, while a clean, well-structured report that respects the doctor's time and need for verification will be embraced. The study concludes that for artificial intelligence to be truly helpful in a hospital, it must be designed not just for accuracy, but for the human workflow. The best tool is not necessarily the one that knows the most, but the one that presents its knowledge in a way that feels safe, transparent, and easy to use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.