SciPaths: Forecasting Pathways to Scientific Discovery
This paper introduces SciPaths, a new benchmark and task for forecasting scientific discovery pathways that requires models to identify and ground the enabling contributions and prior work necessary for a target scientific contribution, revealing that current language models struggle significantly with this complex reasoning capability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a magnificent, finished skyscraper. It's the "target contribution"—a new scientific discovery.
Most current AI tools are like tour guides who can tell you, "This building was built in 2023," or "It looks like the building next door." They can list the books the architect read (citations) or guess what the architect might build next.
But they struggle to answer the most important question: What specific, invisible building blocks were absolutely necessary to make this skyscraper stand up in the first place?
This paper introduces SCIPATHS, a new way to test AI on exactly that question. It treats scientific discovery not as a list of references, but as a recipe or a construction blueprint.
The Core Idea: The "Backwards Recipe"
The authors created a benchmark called SCIPATHS (Scientific Paths). They asked a simple but hard question:
"If you wanted to build this specific scientific discovery today, what exact ingredients and tools would you need to have in your pantry before you started?"
They call these ingredients "enabling contributions."
For example, if the discovery is a new type of self-driving car, the "enabling contributions" aren't just "cars" or "computers." They are specific things like:
- A specific way of teaching computers to see (a method).
- A massive dataset of rainy-day photos (a resource).
- A specific mathematical formula for handling slippery roads (a concept).
The Two-Part Challenge
The paper tests AI models on a two-step "cooking challenge":
Task A (The Ingredient List): Given the final dish (the discovery), can the AI list the essential ingredients?
- The Trap: The AI shouldn't just list "flour" (a vague concept). It needs to list "sifted, high-protein flour" (the specific enabling contribution).
- The Result: The AI models are terrible at this. Even the smartest ones only get about 19% of the ingredients right. They are great at naming the "flour" (like "a dataset") but terrible at identifying the specific "sifted, high-protein" part (the core method) that actually made the dish work.
Task B (Finding the Source): Once the AI lists the ingredients, can it find the exact store where you can buy them?
- The Result: If you give the AI the correct list of ingredients (the "Gold" list written by human experts), it gets much better at finding the stores (the original papers).
- The Problem: If you let the AI make its own list first (which it usually gets wrong), it can't find the stores at all. It's like trying to find a store that sells "magic dust" because the AI guessed that was the ingredient, when the real ingredient was "yeast."
The Big Takeaway
The paper reveals a major gap in how AI understands science:
- Current AI is good at finding related books or guessing general ideas.
- What it lacks is the ability to reason backwards from a finished invention to the specific, functional building blocks that made it possible.
Think of it like this: An AI can tell you that a car needs "wheels." But it struggles to explain that this specific car needed "rubber tires with a specific tread pattern designed for snow," and that this specific pattern was invented in a different paper three years ago.
Why This Matters (According to the Paper)
The authors built a massive library of 262 expert-verified "blueprints" (and thousands of AI-generated ones) to prove that current AI models are missing this "dependency reasoning" skill.
They found that:
- AI is bad at the "Core Method": It can easily find the data sources (the "ingredients") but fails to identify the specific algorithms or methods (the "cooking techniques") that are the hardest to invent.
- Knowing what to search for is key: If you tell the AI exactly what building block it needs to find, it can find the paper. But if it has to guess what to look for, it fails.
In short, SCIPATHS is a test to see if AI can stop just "surfing" the internet of science and start actually understanding the construction logic behind how scientific breakthroughs are built. Right now, the AI is still learning how to hold the hammer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.