The Logical Structure of Clinical Trial Eligibility Criteria: A Corpus-Scale Measurement of Disjunction in Non-Small-Cell Lung Cancer Trials
This study demonstrates that nearly half of clinical trial eligibility criteria for non-small-cell lung cancer involve logical disjunctions rather than simple conjunctions, revealing that the common assumption of independent requirements leads to significant safety and efficiency errors in automated patient-trial matching.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every clinical trial begins with a list of rules. These are the eligibility criteria, the specific conditions a person must meet to join a study. They determine who gets the new treatment and who does not, acting as the gatekeepers of medical research. For decades, scientists and computer systems trying to match patients to these trials have operated on a simple, quiet assumption: that the rules work like a checklist where every single item must be checked off. In this view, a patient is eligible only if they satisfy condition A, and condition B, and condition C, all at the same time. It is a straightforward way to think about a complex document, but it treats the text as a flat list of requirements.
However, the language of medicine is rarely that simple. Doctors often write rules that offer choices. A patient might be allowed to join if they have one type of genetic marker or another. They might be excluded if they have taken drug X, drug Y, or drug Z. These are not just lists; they are logical alternatives. If a computer system ignores these choices and treats every option as a mandatory requirement, it makes a fundamental error. It might reject a patient who is actually eligible because they missed one option they didn't need, or it might accept a patient who should be excluded because it failed to see that having just one of several bad conditions was enough to disqualify them. Until now, no one had measured how often these choices actually appear in the massive archives of clinical trials.
A team of researchers set out to solve this puzzle by examining the eligibility criteria for non-small-cell lung cancer, a major form of the disease. They gathered the records of thousands of trials from a public registry, creating a massive collection of nearly 190,000 individual rules. To make sense of this mountain of text, they built a new kind of tool using advanced artificial intelligence. This tool was designed to read the logical structure of each sentence, looking specifically for the presence of "or" statements—places where the rule offers an alternative path. Before trusting the tool with the full dataset, the researchers tested it rigorously. They compared its work against a set of rules that had been carefully labeled by human experts, and then they had three independent medical professionals review a fresh batch of rules to create a new, blind standard. The tool proved to be remarkably accurate, matching the human experts almost perfectly and confirming that a cheaper version of the artificial intelligence could do the same job as a more expensive one.
When the researchers applied this tool to the full collection of lung cancer trials, the results were striking. They found that the simple checklist assumption was wrong for nearly half of all the rules. In the validated sample, 46.2 percent of the individual criteria expressed a logical choice rather than a single requirement. This means that in almost every single trial, there was at least one rule that offered an alternative. The pattern was even more pronounced in the rules that excluded patients. While about 36 percent of the rules that allowed patients in offered a choice, nearly 57 percent of the rules that kept patients out did so by listing multiple ways to be disqualified. If a computer system reads these exclusion rules as a list of things a patient must have to be rejected, rather than a list of things that could disqualify them, it creates a safety risk by letting in people who should be kept out.
The study also revealed how this complexity has changed over time. In the early 1990s, when the researchers looked at the oldest trials, only about 18 percent of the rules offered a choice. By the 2010s, that number had roughly doubled to around 50 percent, where it has stayed steady. This increase happened even though the length of the trial documents did not change dramatically during the period when the numbers were rising, suggesting that the rules themselves are becoming structurally more complex, not just longer. The researchers noted that this pattern held true regardless of who was running the trial or what stage of testing it was in, indicating that this is a general feature of how modern medical protocols are written.
The conclusion is that the way we currently try to match patients to trials needs a fundamental shift. For decades, the field has relied on methods that treat every rule as a mandatory condition. This study shows that approach is mis-specified for nearly half of the criteria and for almost every trial in the registry. The error is not just a minor statistical glitch; it has real consequences. When the system misses a choice in an inclusion rule, it turns away patients who could have helped the study. When it misses a choice in an exclusion rule, it risks the safety of participants by admitting people who should not be there. The researchers argue that before we can build better systems to match patients to trials, we must first measure and understand the logical shape of the rules themselves. The era of treating clinical trial criteria as a simple list of requirements is over; the reality is a landscape of choices that must be navigated with care.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.