← Latest papers
📄 bioengineering

Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models

The paper introduces FD-SCoPE, a feedback-driven language model framework that enhances clinicians' ability to query and derive insights from systematic review evidence tables by combining executable queries with expert corrections to provide auditable, high-accuracy answers.

Original authors: Roy Choudhury, M., Roy Chowdhury, S., Sahoo, S., Khan, M. A., Khakwani, K. Z. R., Sonbol, M. B., Riaz, I., Gupta, V.

Published 2026-10-05
📖 5 min read🧠 Deep dive

Original authors: Roy Choudhury, M., Roy Chowdhury, S., Sahoo, S., Khan, M. A., Khakwani, K. Z. R., Sonbol, M. B., Riaz, I., Gupta, V.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Medical progress relies on a steady stream of proof. When doctors decide which treatment to offer a patient, they look to systematic reviews, which are massive summaries of clinical trials. These reviews condense hundreds of studies into a single table, where each row represents a trial and the columns list details like the disease treated, the drugs used, and the results. This table is the foundation of modern care, yet it is often locked away from the very people who need it most. To find a specific piece of information, a doctor usually has to ask a data analyst to run a computer query, a process that can take days. Furthermore, the table only holds what was explicitly written down. If a doctor wants to know the "class" of a drug based on its name, or group follow-up times into meaningful categories, the table offers no direct answer. These missing details must be figured out by hand, a slow and error-prone task that varies from one expert to another.

Researchers at Arizona State University and the Mayo Clinic have built a new tool designed to unlock this locked table, allowing doctors to ask questions in plain English and get immediate, reliable answers. The system, called FD-SCoPE, acts as a bridge between human curiosity and structured data. It does not simply guess the answer; instead, it translates a doctor's question into a precise computer command to find the right trials, and then uses a separate, verified step to calculate any missing details. Crucially, the system learns from its mistakes. When an expert corrects an answer, the system saves that correction as a rule, ensuring that the same question gets the right answer next time, even if it involves a trial the system has never seen before.

The team tested this approach on a living evidence table containing 159 records of cancer trials involving immune checkpoint inhibitors, a type of drug that helps the body's immune system fight cancer. They asked the system 140 different questions that a clinician might actually ask, such as "Which phase 3 trials tested a specific drug with chemotherapy?" or "How many patients were enrolled in trials for a specific disease?" The system answered every single one of these tasks correctly. In comparison, other methods that relied on the same underlying language model but lacked the specific safeguards of FD-SCoPE got between 91% and 98% of the answers right. The errors in those other methods usually happened because they failed to match a drug name exactly or applied a condition to the wrong column of data. FD-SCoPE avoided these pitfalls by first checking the actual values stored in the table before generating its answer.

The real challenge, however, lies in questions that require the system to figure out something not explicitly recorded. For instance, a trial might list a drug by its generic name, but the doctor needs to know if it targets a specific protein. The researchers created 1,500 such questions to test this ability. FD-SCoPE successfully found the correct trials for 99.3% of these questions and derived the missing information with high accuracy. It outperformed four other approaches, including systems that tried to read the entire table at once or write their own code to solve the problem. The key to its success was a two-step process: first, it used a strict computer query to select the relevant trials, ensuring the right group of patients was being considered. Only after the correct trials were selected did the system calculate the missing attribute. This separation prevented the system from missing relevant studies, a common failure in other methods where about one in ten relevant trials were overlooked.

Perhaps the most significant finding was how the system improved over time through feedback. The researchers simulated a scenario where an expert reviewed a small set of 299 answers and corrected the ones that were wrong. The system took these corrections and turned them into reusable rules. When tested on a new set of 1,201 questions that the system had never seen before, the accuracy jumped significantly. For questions involving complex details like drug dosing schedules or specific types of treatment endpoints, the accuracy rose from roughly 60% to over 90%. This demonstrated that the system does not just memorize specific answers; it learns the underlying logic of how to derive new information. However, the researchers also found a limit to this learning. If a rule was applied to a type of question it was not designed for, the system could make confident errors. To prevent this, the design requires that any new rule learned by the system must be approved by an expert before it is used again.

The study also examined how well the system handled the messy reality of human language. Doctors often use abbreviations, brand names instead of generic names, or make small typing errors. The system remained robust against these variations, with only a tiny drop in accuracy when faced with typos or shorthand. It also included safety checks to ensure it could not be tricked into changing the data or providing medical advice outside its scope. The researchers were careful to note that while the system performed exceptionally well in these tests, it was evaluated on a single table of cancer trials and has not yet been used by doctors in a real-world setting. The results show that combining a language model with strict computer queries and expert feedback creates a powerful, auditable tool. It allows clinicians to access the full depth of trial evidence without needing a data analyst, turning a static table into a dynamic resource that can answer complex questions with speed and precision.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →