← Latest papers
💬 NLP

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

This paper introduces ResearchBench, the first large-scale, contamination-free benchmark that evaluates LLMs on scientific discovery by decomposing the task into inspiration retrieval, hypothesis composition, and ranking, revealing that models excel at retrieving novel knowledge associations across 12 disciplines.

Original authors: Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, Dongzhan Zhou

Published 2026-04-21
📖 5 min read🧠 Deep dive

Original authors: Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, Dongzhan Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a scientist trying to invent something new, like a better battery or a cure for a disease. Usually, you start with a problem (the Research Question) and look at what's already known (the Background). The hardest part is the "spark of genius"—finding a weird, unrelated idea from a completely different field that, when mixed with your problem, creates a brilliant new solution.

This paper, ResearchBench, is like a giant, high-tech gym for Artificial Intelligence (AI) to test if it can be that "spark of genius."

Here is the breakdown of how they tested the AI, using simple analogies:

1. The Big Idea: Science as a Recipe

The researchers believe that almost every great scientific discovery is just a recipe with three ingredients:

  1. The Problem: What are we trying to fix?
  2. The Inspiration: A random, unrelated idea from somewhere else (like using a cooking technique to fix a computer chip).
  3. The Mix: Combining them to create a new hypothesis (a guess at a solution).

They broke down the job of "doing science" into three specific tasks to test the AI:

  • Task A: The Treasure Hunt (Inspiration Retrieval). Can the AI find that weird, unrelated idea that actually helps?
  • Task B: The Chef (Hypothesis Composition). Can the AI mix the problem and the idea into a coherent new recipe?
  • Task C: The Critic (Hypothesis Ranking). Can the AI look at ten different recipes and pick the one that is most likely to work?

2. The Test Kitchen: ResearchBench

To test this, they built a massive database called ResearchBench.

  • The Menu: They looked at 1,386 brand-new scientific papers (published in 2024 or later) across 12 different fields, from Physics to Law. They chose these dates so the AI couldn't have "cheated" by memorizing the answers during its training.
  • The Automation: Instead of humans reading every paper, they built a robot team (an "agentic framework") that reads the papers and automatically pulls out the problem, the background, the "spark" ideas, and the final solution.
  • The Quality Check: Real human experts (PhD students in science) tasted the robot's work. They confirmed the robot was about 92% accurate in finding the right ingredients.

3. The Results: How Did the AI Do?

🏆 The Treasure Hunt (Retrieval): The AI is a Genius Detective

  • The Challenge: The AI had to find the "hidden connection." For example, if the problem was about traffic jams, the AI needed to find a paper about ant colonies (which is a common real-world inspiration for traffic algorithms).
  • The Result: The AI was surprisingly good at this! Even when the connection was very weak or the ideas seemed totally unrelated, the AI could spot them.
  • The Analogy: Imagine you are in a library with millions of books. If you ask a human, "Find me a book about baking that might help me fix a broken engine," they might say, "That's impossible." The AI, however, said, "Oh, I found a book on heat distribution in ovens that explains how to cool an engine better." It found patterns humans missed.

🍳 The Chef (Composition): The AI is a Good Apprentice

  • The Challenge: Once the AI found the "spark" (the inspiration), it had to write a new scientific hypothesis combining the two.
  • The Result: The AI did okay, but it wasn't perfect. It could mix the ingredients, but sometimes the recipe was a bit messy or missed a key step. It's like a chef who knows the ingredients but hasn't mastered the cooking technique yet.

⚖️ The Critic (Ranking): The AI is a Confused Judge

  • The Challenge: The AI had to look at a "Good" hypothesis and a "Bad" one and pick the winner.
  • The Result: This was the hardest part. The AI often got confused.
  • The Glitch: The researchers found a funny flaw: The AI had a "position bias." If you showed it Hypothesis A first and Hypothesis B second, it often picked A. If you swapped them, it picked B. It was like a judge who just picks the first person they see, regardless of who is actually better.

4. The Big Takeaway: The "Research Mine"

The authors conclude that we shouldn't just think of AI as a calculator or a chatbot. They call it a "Research Hypothesis Mine."

  • The Mine: The AI is a deep mine full of hidden connections between different fields of knowledge.
  • The Miners: The more powerful the AI (the bigger the model), the more "miners" it has digging for these connections.
  • The Future: Even though the AI isn't perfect at writing the final solution yet, it is incredibly good at finding the raw materials (the inspirations) that humans might never think to look for.

In short: This paper shows that AI is becoming a fantastic "idea scout." It can wander through the library of human knowledge, find two things that seem unrelated, and say, "Hey, if you mix these, you might discover something new!" It's not quite ready to be the lead scientist yet, but it's the best research assistant we've ever had for finding the spark of inspiration.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →