← Latest papers
💬 NLP

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

The paper introduces NatureBench, a benchmark derived from Nature-family publications that reveals current AI coding agents struggle to surpass existing SOTA on real scientific discovery tasks, primarily succeeding through methodological translation rather than genuine invention and failing due to incorrect method selection and limited compute budgets.

Original authors: Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, Weifeng Huang, Kai Tian, Zhenzhao Yuan, Jincheng Zhong, Weizhi Wang, Ning Ding, Bowen Z
Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, Weifeng Huang, Kai Tian, Zhenzhao Yuan, Jincheng Zhong, Weizhi Wang, Ning Ding, Bowen Zhou, Kaiyan Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive, high-stakes cooking competition. The judges (the researchers) have a cookbook filled with 90 incredibly complex recipes from the world's most prestigious culinary magazines (the Nature family of scientific journals). These aren't just "how to make toast" recipes; they are intricate dishes like "deconstructing a protein structure" or "predicting how a new drug molecule will behave."

The contestants in this competition are AI Coding Agents—smart computer programs designed to write code and solve problems.

Here is the twist: The contestants are not allowed to see the original recipes. They are only given the raw ingredients (the data) and a description of what the final dish should taste like (the goal). Their job is to invent their own method to create a dish that is just as good as, or better than, the original award-winning version.

This is the story of NatureBench, a new test designed to see if AI can move from simply copying a recipe to actually inventing a new one.

The Problem: The "Copy-Paste" Trap

Before this test, many AI benchmarks were like giving a student a math problem and the answer key, then asking them to show their work. If the student just copied the answer key, they passed. But real science isn't about copying; it's about discovery.

The creators of NatureBench realized that previous tests were messy. Sometimes the "kitchen" (the computer environment) was broken, or the ingredients were missing, making it impossible to tell if the AI failed because it was dumb or because the test was rigged.

To fix this, they built NatureGym. Think of NatureGym as a super-organized, automated kitchen crew.

  1. The Filter: They scanned thousands of papers and picked only the 90 that were fair, had all the ingredients available, and could be graded by a machine.
  2. The Firewall: They built a "glass wall" around the AI. The AI can see the ingredients and the goal, but it cannot see the original chef's notes or the secret sauce used in the winning dish.
  3. The Container: Every task is put into its own sealed, perfect little kitchen (a container) so that the AI can't cheat by using outside tools or the internet.

The Competition: How Did the AI Do?

They invited 10 of the smartest AI "chefs" (like Claude Opus 4.7, GPT-5.5, and Gemini 3.5) to tackle these 90 dishes. They had a strict rule: No Googling. They had to figure it out from scratch.

The Results were sobering:

  • The "Good Enough" Club: Even the best AI only managed to match the original, award-winning recipe about 48% of the time.
  • The "Better Than Original" Club: The AI only managed to create a better dish than the original in just 18% of the cases.
  • The Reality Check: For the other 82% of the time, the AI either made a dish that was worse than the original or couldn't finish it at all.

How Did They Succeed (and Fail)?

The researchers looked closely at how the AI chefs worked to understand the results.

The Winning Strategy: "The Translation Trick"
When the AI succeeded, it usually didn't come up with a brilliant new scientific discovery. Instead, it used a trick called Methodological Translation.

  • The Analogy: Imagine the original recipe was for a complex French soufflé (a hard scientific problem). The AI didn't try to make a soufflé. Instead, it looked at the ingredients and said, "Hey, this looks like a pizza I know how to make!" It turned the difficult scientific problem into a simple, familiar math problem it had seen a million times before.
  • 45% of the successes happened because the AI just turned the science into a standard "guess the answer" game. It didn't invent new science; it just applied old tricks to new data.

The Losing Strategy: "Wrong Tools and Running Out of Time"
When the AI failed, it wasn't usually because it didn't understand the instructions.

  • Wrong Tool (45%): It picked the wrong "kitchen gadget." It tried to use a blender when it needed a whisk. It chose a method that was fundamentally wrong for the type of problem.
  • Running Out of Time/Money (24%): The AI started cooking but ran out of the 4-hour time limit or the computer power (budget) needed to finish the dish.
  • Misunderstanding (Rare): It was very rare for the AI to simply not understand what the dish was supposed to be.

The Difficulty Levels

Not all recipes were equally hard.

  • The Easy Dishes: Tasks involving "Relational Reasoning" (connecting dots) and "Protein Biology" were the easiest for the AI.
  • The Hard Dishes: "Physical Modeling" (simulating physics) and "Molecular Design" were the hardest.
  • The "Frankenstein" Dishes: The hardest tasks of all were the ones that mixed two different fields (like mixing biology and physics). The AI struggled to combine knowledge from different "cuisines."

The Big Takeaway

NatureBench tells us that while AI coding agents are getting very good at following instructions and fixing bugs, they are not yet true scientific inventors.

They are excellent at taking a complex scientific problem and saying, "I know a shortcut for this!" But they are not yet good at looking at a problem and saying, "I have a brand new idea that no human has thought of yet."

The paper concludes that for AI to truly help science, we need to stop just asking it to copy the winners and start teaching it how to choose the right tools and think like a scientist, not just a coder.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →