← Latest papers
🤖 AI

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

This paper introduces FirstResearch, a framework that enhances the auditability of LLM-driven scientific discovery by generating structured "Research Question Certificates" that explicitly define mechanisms, assumptions, and falsifiable hypotheses, demonstrating superior performance over existing baselines in preliminary evaluations.

Original authors: Yufeng Wang

Published 2026-07-08
📖 4 min read☕ Coffee break read

Original authors: Yufeng Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you ask a very smart, well-read robot to come up with a new scientific idea. The robot might say, "Let's study how AI agents learn new skills!" It sounds great, but if you ask, "How exactly will we know if it works? What specific thing would prove you wrong?" the robot might fumble. It gives you a vague answer that sounds plausible but doesn't actually hold up to scientific scrutiny.

This paper introduces a system called FirstResearch to fix that problem. Think of it as a "Quality Control Inspector" for the very first step of scientific discovery.

Here is how it works, using some everyday analogies:

1. The Problem: The "Vague Idea" Trap

Current AI scientists are like enthusiastic interns who can write beautiful reports and run experiments. But sometimes, their starting idea is like a recipe that says, "Make a delicious cake." It doesn't say what kind of cake, how to mix it, or what happens if you forget the eggs. If the cake fails, you don't know if it was the flour, the oven, or the fact that you forgot the eggs.

In science, if you can't clearly state what would prove an idea wrong (a "falsifier"), you haven't really asked a good question.

2. The Solution: The "Research Question Certificate"

FirstResearch forces the AI to fill out a Research Question Certificate before it is allowed to do any real work. Think of this certificate like a building permit or a flight plan.

Before a plane takes off, the pilot must fill out a form that says:

  • The Basics: What are the rules of physics here? (Primitive definitions)
  • The Assumptions: What are we assuming is true?
  • The Engine: How does this actually work? (Mechanism model)
  • The Conflict: What is the specific problem or tension we are trying to solve?
  • The Test: What is the one simple experiment that could prove us wrong?
  • The "Oops" Plan: If the experiment fails, what specific part of our plan do we change?

If the AI can't fill out this form clearly, the system stops it. It doesn't let the AI run the experiment until the "flight plan" is solid.

3. The "Gatekeeper" (The Repair Shop)

The system has a smart gatekeeper. If the AI submits a certificate that is too vague (e.g., "We will see if it gets better"), the gatekeeper sends it back to a repair shop.

  • It says: "This is too generic. You need to find a specific 'tipping point' or a 'failure mode'."
  • It forces the AI to sharpen its question until it's specific enough to be tested.

4. The Results: Did It Work?

The authors tested this system against other AI methods (like "AI Scientist" or "Agent Laboratory") on 10 different topics about how AI agents learn.

  • The Judges: They used two different powerful AI "judges" (DeepSeek and Gemini) to grade the ideas.
  • The Score: FirstResearch scored 4.86 out of 5, while the next best system scored 4.38.
  • The "Certificate" Proof: To prove that the certificate was the secret sauce, they ran a test where they removed the certificate requirement. The score plummeted to below 1 out of 5. This shows that the structured form is what made the ideas good, not just the AI's general intelligence.

The Bottom Line

The paper doesn't claim that FirstResearch is a fully autonomous scientist that can cure diseases or discover new materials on its own yet. Instead, it claims that forcing AI to write a structured "permit" (the certificate) before starting makes its scientific questions much clearer, more testable, and easier for humans to check.

It turns a "maybe this is cool" idea into a "here is exactly how we will test this" plan.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →