← Latest papers
🤖 AI

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

This position paper argues that current agentic AI systems are ill-equipped for fully autonomous scientific discovery due to fundamental limitations in problem selection, training data gaps regarding tacit laboratory knowledge, output homogenization from preference optimization, and inadequate benchmarking, necessitating a redesign centered on simulation-based verification, persistent world models, and need-driven application.

Original authors: Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan

Published 2026-05-12
📖 6 min read🧠 Deep dive

Original authors: Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Co-Pilot" vs. The "Captain"

Imagine you are trying to navigate a ship through uncharted waters to find a new island.

  • The Current AI: It is an incredibly smart Co-Pilot. It can read the map, calculate the fastest route, and even steer the ship if you tell it exactly where to go. It is amazing at solving problems humans give it.
  • The Paper's Argument: The authors say this Co-Pilot is not ready to be the Captain. It cannot decide which island to look for, it doesn't know what to do when the map is wrong, and it tends to suggest the exact same islands that everyone else is already looking at.

The paper argues that while AI is a fantastic tool for helping scientists, it is fundamentally broken if we try to let it run the entire scientific process on its own.


The Four Major Roadblocks

The authors identify four specific reasons why AI cannot yet be an "autonomous scientist."

1. The "McNamara Fallacy": Measuring Only What's Easy

The Analogy: Imagine a police chief who only arrests people for jaywalking because it's easy to count the tickets, while ignoring actual crimes like bank robbery because they are harder to track.
The Paper's Claim: AI systems are trained to optimize for things they can measure (like data numbers). This makes them ignore the most important, difficult, and "unmeasurable" scientific questions.

  • Real-world example: In battery research, AI might focus only on "conductivity" because it's a number it can calculate. It ignores harder, unquantified factors like "is this material actually safe to manufacture?" or "will it last 10 years?" because those are messy and hard to score.

2. The "Silent Library": Missing the Mistakes

The Analogy: Imagine a student studying for a test using only a textbook that lists every correct answer but deletes every page where the author made a mistake, got confused, or tried a method that failed.
The Paper's Claim: AI learns from published scientific papers. But science has a "publication bias": journals only print success stories. They throw away the "failed experiments" and the "tacit knowledge" (the unwritten tricks scientists learn by doing, like "don't mix these chemicals on a rainy day").

  • The Result: The AI thinks science is a straight line of successes. It doesn't know how to handle failure, so when it hits a dead end in the real world, it doesn't know how to pivot. It's like a chef who has only read recipes but never actually burned a meal or learned how to adjust for a bad ingredient.

3. The "Hive Mind": Everyone Thinks Alike

The Analogy: Imagine asking five different geniuses to invent a new flavor of ice cream. If they all went to the same school, read the same books, and were graded by the same teacher who liked "strawberry," they would all invent strawberry ice cream.
The Paper's Claim: The paper calls this "Diversity Compression." AI models are trained on the same published literature and then "tuned" by humans to sound polite and agreeable.

  • The Experiment: The authors asked different AI models (from different companies) to come up with new scientific ideas. Even though they were different models, they all came up with the exact same ideas. They converged on the "popular" answers (like specific types of battery materials) and ignored weird, novel, or risky ideas. If you ask ten AI scientists for a hypothesis, you effectively get the answer from one scientist, repeated ten times.

4. The "Video Game" Problem: No Real-World Feedback

The Analogy: Imagine playing a video game where you get points for guessing the right answer, but you never actually have to build the thing you guessed. You can win the game without ever touching the real world.
The Paper's Claim: Current AI benchmarks (tests) are like video games. They ask the AI to predict a result once, and if it's right, it gets a gold star. Real science is a loop: You guess, you test, the test fails, you change your guess, and you test again.

  • The Gap: AI is great at the "guessing" part. But it is terrible at the "loop" part. If an experiment fails, the AI often treats it as "noise" rather than a clue to change its mind. It doesn't have a mechanism to say, "Okay, my simulation was wrong; I need to rewrite my understanding of physics."

The Solution: How to Fix the AI

The authors don't say we should stop using AI. They say we need to change how we build it.

  1. Teach it the "Hidden Rules": We need to train AI on videos of lab work and records of failed experiments, not just final papers. It needs to learn from the messiness of real science.
  2. Stop the "Hive Mind": We need to change how we train AI so it doesn't just try to please human reviewers. We need it to be encouraged to be weird and diverse.
  3. The "Pre-Registration" Rule: Before an AI runs an experiment, it must write down its plan and its hypothesis in a public log. This stops it from secretly trying a million things and only showing us the one that worked (a trick called "p-hacking").
  4. Simulators as Teachers: Instead of just reading books, AI should practice in high-fidelity computer simulations that act like a "flight simulator" for science, allowing it to crash and learn without wasting real materials.
  5. The "World Model": The AI needs a persistent memory that remembers what it learned yesterday and updates its beliefs today. It can't just be a chatbot that forgets the conversation as soon as the window closes.

The Bottom Line

The paper concludes that AI is a brilliant "Co-Scientist" that can help humans do their jobs faster. But it is not yet an "Autonomous Scientist" capable of leading the way.

If we try to let AI run the show right now, we will just get a faster version of the same old ideas, ignoring the hard, messy, and truly groundbreaking questions that require human judgment, intuition, and the ability to learn from failure.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →