Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
This paper introduces a model-agnostic "historical backtesting" protocol that evaluates scientific question generators by freezing questions against a historical corpus and objectively scoring them against a temporally isolated future corpus, demonstrating that evidence-structure-first generation outperforms LLM-only prompting while revealing that human annotators struggle to agree on outcome taxonomies even more than AI judges do.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science has always been driven by questions. A researcher looks at the world, sees something puzzling, and asks, "What if?" or "Why?" For decades, artificial intelligence has been taught to read vast libraries of scientific papers and summarize them. Now, a new generation of AI is being asked to do something harder: to look at the same libraries and invent the next big questions that humans should be asking. But how do we know if an AI is actually good at this? If a computer suggests a new line of inquiry, is it a brilliant insight or just a lucky guess? Until now, the only way to judge these systems was to ask human experts for their opinions or to let other computers rate the ideas. Both methods are subjective. They rely on what people think sounds interesting at the moment, not on whether the idea actually leads to new discoveries.
This uncertainty matters because the scientific community is investing time and money into following these AI-generated leads. If the questions are poor, that investment is wasted. If the questions are brilliant, they could accelerate our understanding of the universe. The core challenge is that we cannot know the value of a scientific question until the future arrives and the scientific community answers it. A question asked today might be ignored for ten years, or it might be solved next month. Without a way to see the future, we have no way to tell if an AI is truly foresightful or just mimicking the style of recent papers.
A team of independent researchers has proposed a solution to this problem by turning the clock backward. Instead of trying to predict the future, they built a system to test the past. They call this "historical backtesting." The idea is simple but powerful: take a snapshot of all the scientific knowledge available up to a specific date in the past. Give an AI system only that old knowledge and ask it to generate new research questions. Then, freeze those questions so they cannot be changed. Finally, look at the scientific papers published after that date—papers the AI never saw—to see what actually happened. Did the community answer the question? Did they ignore it? Did they independently come up with the same question? Or did they prove the question's starting assumption was wrong? By comparing the AI's frozen questions against the real history that followed, the researchers created a way to measure the quality of scientific questions with hard data rather than opinions.
To test this method, the researchers focused on the study of exoplanet atmospheres, a fast-moving field where scientists analyze the light from distant worlds to determine what gases they contain. They set a cutoff date of December 31, 2020. They gathered every relevant paper published before that date and fed them into a system designed to find tensions and contradictions in the existing data. This system, which does not rely on a standard language model to "dream up" ideas but instead looks for specific structural conflicts in the evidence, generated ten frozen questions. The researchers then examined the literature published between 2021 and 2026. The results were striking. Every single one of the ten questions generated by the system was engaged with by the scientific community in the following years. Two were fully answered, seven were partially addressed, and one was independently posed by scientists working in the field but remained open. Most significantly, one question challenged a widely accepted conclusion about the water content of a specific planet, HD 209458 b. The community subsequently performed the exact tests the question had called for and proved the original conclusion wrong, showing that the low water levels were an artifact of the measurement method, not a feature of the planet.
The researchers did not stop there. They realized that a small sample of ten questions might not tell the whole story, so they expanded the test to include hundreds of questions generated by different methods. They compared their evidence-based system against other approaches, including a standard artificial intelligence that simply reads the old papers and tries to write new questions, and a system that just picks the most famous results from the past and asks if they are still true. When they looked at the larger group of 424 questions, they found that the standard AI was very good at asking broad, fashionable questions that attracted attention. However, it rarely led to a definitive answer or a refutation of old ideas. Its questions were like a wide net that caught everything but held nothing specific. In contrast, the system that looked for structural conflicts in the evidence was much better at finding questions that led to concrete answers or overturned established beliefs.
A critical part of their work involved testing whether the artificial intelligence was simply "recalling" the answers from its training data. Since modern AI models are trained on vast amounts of text, including papers published after the 2020 cutoff, there was a concern that they were just recalling the future rather than predicting it. The researchers designed a stress test to separate genuine foresight from memory. They ran the same systems against different time periods, including one where the future happened after the AI was even trained. They found that the standard AI performed well only when the future was close to its training data, suggesting it was relying on memory. But the evidence-based system continued to find valuable, specific questions even in time periods where the AI could not have known the answers. This proved that the ability to find good questions came from the structure of the evidence itself, not from the AI's memory of the future.
The study also revealed a surprising flaw in how we currently evaluate these systems. The researchers tried to use both human experts and different AI models to grade the questions. They found that human experts often disagreed with each other, and that different AI models agreed with each other much more than they agreed with humans. This suggests that the standard way of judging these questions—relying on a single AI or a small group of humans—is unreliable. The researchers concluded that the way we define a "good" question needs to be clearer and more consistent before we can trust any judge, human or machine, to score them.
Ultimately, this work offers a new way to measure the value of scientific ideas. It moves the field away from subjective opinions about what sounds interesting and toward a measurable record of what actually happened. The researchers have made their data, their code, and their frozen questions public, allowing anyone to run the same tests. They have also set up a new experiment where questions generated today will be frozen and scored against the scientific literature of the next four years. This prospective test will provide a contamination-free measure of whether AI can truly help scientists see the future. The findings suggest that while artificial intelligence is excellent at phrasing and summarizing, the real power to discover new scientific questions lies in systematically searching for contradictions and gaps in the evidence, a task that requires a structured approach rather than just a powerful language model.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.