← Latest papers
💬 NLP

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

This paper introduces the Lit2Test benchmark, a novel framework that evaluates language models' research ideation capabilities by requiring proposals to include falsifiable outcomes, thereby enabling a reliable, blind, and strict ranking of frontier models based on the quality of their proposed tests rather than surface fluency.

Original authors: Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (Stat
Published 2026-08-25
📖 6 min read🧠 Deep dive

Original authors: Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Hongyao Zuo (Tianjin University), Ziwen Gong (Hainan University), Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Tong Yang (Peking University), Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Xiaohui Li (Huawei Technologies), Haoli Bai (Huawei Technologies)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Science advances not just by having good ideas, but by having ideas that can be proven wrong. In the world of research, a proposal is only truly useful if the researcher can say, in advance, exactly what observation would force them to reject their own theory. Without this commitment, an idea can sound brilliant and convincing while remaining impossible to test. It floats in a space where success is measured by how well it matches a future outcome that has already happened, or by how smoothly it is written. This leaves a gap in how we evaluate artificial intelligence. When large language models are asked to suggest new research directions, current methods often reward style over substance, or they judge the model based on whether it accidentally guessed a paper that was later published. Neither approach tells us if the model can actually design a test that separates a real discovery from a lucky guess.

A team of researchers has built a new way to measure this specific ability. They created a system called Lit2Test, which asks models to do something much harder than simply brainstorming. Instead of generating a vague research topic, the models must fill out a strict six-part contract for every idea they propose. This contract requires them to identify a specific gap in existing knowledge, state a clear hypothesis, design a minimal experiment to test it, and crucially, define exactly what result would prove them wrong. By forcing the model to pre-commit to its own potential failure, the researchers turned the evaluation of research ideas from a matter of opinion into a matter of logic. The study found that when models are judged on their ability to create these falsifiable tests, a clear and consistent ranking emerges, driven not by how fluent the writing sounds, but by the quality of the experimental design.

The researchers began by gathering 200 real-world research scenarios, each built from a neighborhood of four recently published papers that shared a topic but disagreed on a key point. These were not made-up problems; they were drawn from actual scientific literature where experts were still debating the answer. For each scenario, four different advanced language models were asked to propose a next step. The models were not allowed to write a free-form essay. Instead, they had to fill in the six fields of the contract: the literature gap, the hypothesis, the minimal test, the decisive metric, the result that would support the idea, and the result that would falsify it. This structure ensured that every proposal was a complete, testable plan rather than a vague suggestion. The researchers then pitted these proposals against each other in blind comparisons, where a judge model, unaware of which AI generated which idea, decided which proposal was better based on the contract.

To ensure the results were not skewed by the order in which the ideas were presented, the researchers judged every pair of proposals twice: once in the original order and once with the order reversed. They found that about 20 percent of the time, the judge changed its mind when the order was swapped, indicating that the comparison was unstable. By setting aside these unstable cases and focusing only on the 80 percent where the judge agreed regardless of order, they established a strict ranking of the four models. The results showed a clear hierarchy: one model consistently produced the best proposals, followed by a second, then a third, and finally a fourth. This ranking held true across thousands of statistical resamplings, meaning the order was robust and not a fluke of the data.

The study also investigated what actually drove these rankings. It turned out that the models were not separated by how well they wrote or how confident they sounded. The decisive factor was the quality of the test design. The top-performing models were better at creating experiments that were small enough to be practical yet specific enough to isolate the mechanism they were testing. They also excelled at defining a metric that could clearly distinguish between success and failure. The requirement to state what would prove the idea wrong acted as a floor; all models had to meet this standard to be considered valid, but the best models went beyond the minimum to design tests that were genuinely insightful. The researchers confirmed this by checking if the judge was influenced by style. They found that when the content of a proposal was identical but the formatting was changed, the judge's decision did not change. The system was truly measuring the substance of the research plan, not the polish of the presentation.

To verify that their automated judge was not missing something obvious, the researchers brought in three human experts to review a smaller, representative sample of the proposals. These humans, who were senior students in computer science, were asked to judge the same proposals without knowing which model created them. The humans agreed with the automated judge's ranking in nearly 90 percent of the cases where the judge was confident. This high level of agreement suggested that the automated system was capturing something real about the quality of the ideas. However, the human judges also showed that the task was difficult; even among experts, there was some disagreement on the finer points, which reinforced the need for a strict, structured contract to make the evaluation possible.

The paper explicitly rules out the idea that a model's ability to predict future papers is a good measure of its research potential. The researchers showed that judging ideas based on whether they match a later-published paper is flawed because it rewards guessing the future rather than designing a test. It also penalizes valid ideas that take a different path than the one that eventually happened. By removing the "answer key" of a future paper and focusing entirely on the logic of the test itself, Lit2Test provides a different kind of insight. It shows that while large language models can generate fluent text, their true capability in scientific ideation lies in their ability to formulate a hypothesis that can be rigorously challenged. The study concludes that the most reliable way to evaluate these models is not to see if they can guess the future, but to see if they can clearly state what would prove them wrong today.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →