← Latest papers
🧬 biology

Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

This study demonstrates that advanced large language models, particularly GPT-5 and GPT-5 Nano, can match domain experts in extracting and appraising evidence for microbial oncogenesis research, supporting their use for scalable systematic evidence synthesis despite persistent challenges in methodological appraisal and contradiction identification.

Original authors: Kaela Kokkas, Hairong Wang, Richard Klein, Nazir A. Ismail, Natalie Irwin, Mohammad Z. Moonsamy, Kubendran Naidoo, Jeremy Nel, Ekene E. Nweke, Raveen Parboosing, Emmanuel K. Sekyi, Rebecca T. van Dors
Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Kaela Kokkas, Hairong Wang, Richard Klein, Nazir A. Ismail, Natalie Irwin, Mohammad Z. Moonsamy, Kubendran Naidoo, Jeremy Nel, Ekene E. Nweke, Raveen Parboosing, Emmanuel K. Sekyi, Rebecca T. van Dorsten, Bruce A. Bassett, Robert F. Breiman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the world of medical research as a massive, chaotic library that never stops growing. Every single day, thousands of new books—scientific papers—are written about tiny germs (microbes) and how they might cause diseases like cancer. For a long time, the only way to find out if a specific germ is a villain was to have a team of super-smart human detectives (domain experts) read every single page, cross-reference facts, and decide if the evidence was strong enough. But with over a million new articles published every year, this "human-only" detective work has become impossible. It's like trying to drink from a firehose; the experts just can't keep up. This is where Artificial Intelligence (AI), specifically a type called Large Language Models (LLMs), steps in. Think of these AIs as super-fast, super-readers that can scan the entire library in seconds. But here's the big question: Can a robot really think like a human expert? Can it spot the subtle clues that prove a germ causes cancer, or will it just make up stories (a problem called "hallucinating") and get the facts wrong? Scientists needed to know if they could trust these digital detectives before letting them take over the job of sorting through the world's medical evidence.

This paper sets up a high-stakes showdown between human experts and AI to see who is better at investigating a specific mystery: whether a virus called the Mouse Mammary Tumor Virus-Like Virus (MMTV-LV) causes breast cancer in humans. The researchers gathered 24 real scientific papers about this virus and created a giant, 77-question quiz to test the detectives. They asked four different AI models (two from Google called Gemini, and two from OpenAI called GPT-5) to read the papers and answer the quiz, while a team of human experts did the same. The goal wasn't just to see who got the most answers right, but to see if the AI could think and judge evidence just like the humans.

The results were a mix of "wow" and "whoa." The study found that two of the AI models, GPT-5 and GPT-5 Nano, performed almost exactly like the human experts. When the researchers mixed the AI's answers in with the humans' answers, the group's overall agreement didn't change at all. It was as if the AI had joined the team of experts and was just another brilliant colleague. These AIs were great at reading the text, understanding complex medical language, and summarizing what the papers said. They rarely made up facts (hallucinations), and when they did, it was usually in the smaller, less powerful models.

However, the AI wasn't perfect. The study showed that the AI models, especially the Google Gemini ones, tended to be too nice. They were "lenient," meaning they were more likely to say a paper provided strong proof of a virus causing cancer than the humans were. They sometimes saw a connection where the humans saw none, like assuming a link between mouse studies and human disease without enough proof. The AI also struggled with the trickiest parts of the job: finding contradictions within a paper (like when the numbers in a chart didn't match the text) and critiquing the study's methods. It's like a student who can summarize a story perfectly but misses the fact that the main character's timeline doesn't make sense.

In the end, the paper suggests that we can trust advanced AI to help us sift through the mountain of medical research. They are fast, they are mostly accurate, and they can match human experts in extracting key facts. But they aren't ready to replace humans entirely just yet. The best approach, the authors suggest, is to use the AI as a powerful assistant to do the heavy lifting of reading and summarizing, while keeping human experts in the loop to double-check the tricky logic and catch the subtle mistakes. It's a partnership where the robot does the reading, and the human does the critical thinking.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →