Contemporary AI lacks the imagination to diverge or negate in science
This study, the largest scientist-in-the-loop evaluation to date, reveals that while AI can generate scientific ideas, it currently lacks the imagination to diverge or propose null hypotheses, often defaults to a narrow "hivemind" of conventional thoughts, and requires human grounding to overcome its limitations in novelty, context-awareness, and judgment compared to expert scientists.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A Reality Check on AI Scientists
Imagine a world where we ask a super-smart robot to help us solve the hardest mysteries of science. We hope it will come up with brilliant, brand-new ideas that humans haven't thought of. This paper is like a massive, real-world "stress test" to see if that robot is actually ready for the job.
The researchers didn't just ask the AI to write a story; they asked it to act like a scientist. They took over 120,000 real scientific papers (from biology, medicine, chemistry, and social sciences) and asked the AI to look at the "puzzle" in each paper and suggest the next step. Then, they sent those AI suggestions back to the actual authors of those papers to see what they thought.
Here is what they found, broken down into three main stories.
1. The "Hive Mind" vs. The "Wanderer"
The Finding: Most AI models (the "non-reasoning" ones) act like a giant, single-minded hive. If you ask ten of them to solve the same puzzle, they all come up with almost the exact same idea. They are very similar to each other, but they are also very similar to what humans have already said. They aren't really "thinking" outside the box; they are just rearranging the furniture inside the box.
The Analogy: Imagine asking ten students to write a story about a lost dog.
- Non-reasoning AI: All ten students write the exact same story: "The dog went to the park, found a bone, and came home." They are all copying the same popular movie plot.
- Reasoning AI: These models are a bit better. They wander around the park a bit more. Maybe one suggests the dog went to the beach, another suggests it went to a bakery. They have more variety.
The Missing Piece: Even the "Wanderers" (reasoning models) have a blind spot. They almost never suggest the idea that "nothing happened."
- The Analogy: If a scientist asks, "Does eating blueberries make you run faster?" the AI will suggest, "Yes, blueberries give you super speed!" or "Maybe they help with hydration."
- The Human Move: A human scientist might say, "Actually, maybe blueberries have zero effect on running speed."
- The AI Failure: The AI rarely suggests the "Null Hypothesis" (the idea that there is no connection). It's like a chef who only knows how to add ingredients but has forgotten how to say, "This dish needs nothing added." The paper suggests this is because the AI was trained on books that mostly report "success stories" (positive results), while "failures" (null results) are often hidden in the "file drawer" of science, unseen by the AI.
2. The "Echo Chamber" of Scientists
The Finding: When the real scientists reviewed the AI's ideas, they didn't act like objective judges. They acted like fans of their own work.
- The "Like Me" Bias: Scientists loved AI ideas that sounded just like their own previous work. They thought these ideas were more "feasible" and "likely to be true." But, ironically, they thought these ideas were less novel (less new).
- The "Seniority" Bias: Older, more famous scientists (the "senior" researchers) were much harsher critics than younger ones. They were less likely to adopt AI ideas, especially in social sciences.
- The "Field" Bias: Scientists in medicine were the most open to AI. Scientists in social sciences were the most skeptical. Why? Because social science is messy and depends heavily on context (like culture and history). The AI, which learns from patterns in data, struggles to understand these shifting, complex human contexts. It tries to copy the "average" opinion, but in social science, there often is no average opinion to copy.
The Analogy: Imagine a group of architects reviewing a new design.
- If the AI suggests a building that looks exactly like the architect's own previous skyscraper, the architect says, "Great! It's safe and doable!" (but also, "Boring, I've seen this before").
- If the AI suggests a wild, new design, the architect says, "That's risky."
- The older, famous architects are the strictest critics. They are like the "Grand Old Men" of the industry who have seen it all and are very skeptical of new gadgets.
3. The Broken Scorecard
The Finding: The scientific community currently uses other AI tools to grade scientific ideas automatically (to save time). This paper tested those tools and found they are terrible at their job.
- The Problem: These automated judges give everyone a "C-" grade. They are afraid to give high or low scores. They cluster everything in the middle. They cannot tell the difference between a brilliant idea and a bad one.
- The Solution: The researchers built a new, custom "reward model" (a specialized AI judge) trained specifically on what human scientists actually like. This new model learned the subtle "taste" of different scientific fields. It got much better at predicting what humans would think, closing the gap between a robot judge and a human judge.
The Analogy: Imagine a music contest where the judges are robots.
- Old Robot Judges: They listen to a rock song, a jazz song, and a country song, and they all give them a score of 5.5 out of 10. They can't tell the difference.
- New Custom Judge: The researchers taught a new robot by showing it thousands of examples of what human music critics liked. Now, this new robot can tell, "Ah, this jazz song has a cool twist that humans love," and gives it a 9. It finally understands the "vibe."
The Bottom Line
The paper concludes that today's AI is a helpful assistant, but not a creative partner.
- What it does well: It can read a lot of books, summarize them, and suggest the next logical step based on what has already been written. It expands the search for ideas.
- What it lacks: It lacks imagination and negation. It doesn't know how to say "What if this is wrong?" or "What if nothing happens?" It doesn't have curiosity. It predicts what has been said, rather than wondering what could be learned next.
The Final Metaphor:
Think of AI as a very fast, very well-read librarian.
- If you ask, "What books have been written about cats?" the librarian can instantly give you a list of every cat book ever written.
- But if you ask, "What is a cat not?" or "What if cats could fly?" the librarian gets confused. The librarian only knows the books that exist on the shelves. They don't know about the books that were never written, or the ideas that were tried and failed.
For science to truly advance, humans still need to be the ones asking the weird questions, imagining the impossible, and deciding what matters. The AI can help organize the library, but it can't write the next great chapter of human knowledge on its own.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.