Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
The Science Edge Evaluation (SEE) benchmark reveals that current multimodal large language models struggle to reliably derive evidence-bounded insights from experimental data in chemistry, biology, and materials science, achieving at best 52.7% accuracy even with tool use, which highlights a critical gap between explaining established concepts and performing real scientific discovery.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of just reading a witness's statement, you have to look at a blurry photo, a fingerprint, a chemical test result, and a torn-up map all at the same time. This is what real scientific discovery feels like. Scientists don't just memorize facts from a textbook; they have to look at messy, real-world data—like images of tiny cells, squiggly lines from machines, and numbers from experiments—and figure out what is actually happening. They have to be careful not to guess based on what they think should happen, but only on what the evidence actually shows.
For a long time, we've been testing computer brains, called "Large Language Models" (LLMs), to see if they are smart enough to help with this detective work. These are the same types of AI that can write stories or answer trivia questions. But there's a big question: Can they handle the messy, visual, and complex reality of a real laboratory? If an AI is going to help scientists discover new medicines or materials, it can't just be a good talker; it has to be a good observer and a careful thinker who knows when to say, "I don't have enough clues yet."
The Great Lab Test: Can AI "See" the Science?
A team of researchers decided to put these AI detectives to the ultimate test. They created a new challenge called Science Edge Evaluation (SEE). Think of SEE as a giant, super-hard exam that doesn't ask, "What is the capital of France?" or "What is the formula for water?" Instead, it hands the AI a real scientific mystery: a photo of a microscope slide, a graph from a chemistry experiment, and a question that requires connecting all those dots to find the answer.
The exam covers three big fields: Chemistry (how stuff changes), Biology (how life works), and Materials Science (how we build new stuff). The questions come from real, published scientific papers and real lab experiments. There are 1,116 questions in total, and they are designed to be tricky. Some ask the AI to read a specific number off a graph, others ask it to spot a mistake in a biological diagram, and some require the AI to mix knowledge from biology and chemistry together.
The Results: A Reality Check
The researchers tested 19 different AI models on this exam. These included the biggest, most famous "general" AIs (the ones that can do everything) and a few "science-specialized" AIs (the ones trained specifically on science books).
Here is the big surprise: None of them passed with flying colors.
In fact, the best-performing model, a powerhouse called GPT-5.6-Sol (Max), only got 48.7% of the answers right. That's less than half! The other models did even worse, with some scoring as low as 15.9%. It's like a student taking a final exam and getting a D-minus.
Even more interesting, the "science-specialized" models (the ones that studied hard in science class) actually did worse on average than the general-purpose models. The general models averaged 34.9% correct, while the science experts only averaged 22.2%. This suggests that just memorizing a lot of science facts isn't enough. The AI needs to know how to use those facts when looking at a messy, real-world picture.
The "Tool" Problem: Does a Calculator Help?
The researchers wondered: "What if we give the AI a calculator and a search engine?" They let the top six models use tools like web search (to look up facts) and a code interpreter (to do math or analyze images).
Giving them tools helped, but not by much. The best model's score went up from 48.7% to 52.7%. While that's a small improvement, it's still not a passing grade. The tools gave the AI more information, but the AI still struggled to figure out which information was important and how to fit it together. Sometimes, the tools even made the AI confused, causing it to pick the wrong path or get stuck in a loop of overthinking.
The Real Problem: Why Are They Failing?
The paper digs deep to find out why the AI is failing. It turns out the problem isn't that the AI doesn't know enough facts. The problem is that the AI is bad at sticking to the evidence.
- The "Blind Guess" Habit: When the researchers took away the pictures and only gave the AI the text, the AI didn't say, "Hey, I can't solve this without the picture!" Instead, it just guessed based on what it had read before. It tried to solve the mystery using its memory instead of looking at the clues. In fact, the AI only admitted it was missing information in just 4.6% of the cases.
- Ignoring the Picture: If the picture showed something weird that didn't match what the AI expected, the AI would often ignore the picture and go with its "gut feeling" (what it learned from training data). It's like a detective ignoring a fingerprint because it doesn't match the suspect's story.
- Over-Confidence: The AI often tried to fill in the blanks. If an experiment was incomplete, the AI would invent a conclusion as if it had all the data. Real scientists know that sometimes you just don't have enough evidence to make a claim, but the AI is too eager to give an answer.
The Missing Step
The paper concludes that we are missing a crucial step in making AI useful for science. We have been trying to teach AI to be a library (storing facts) or a calculator (doing math). But real science requires the AI to be a detective who can look at messy evidence, admit when they don't know the answer, and only make conclusions that are strictly supported by what they can see.
The authors suggest that for AI to truly help with real scientific discovery, it needs to learn how to manage evidence better. It needs to learn when to stop guessing and when to ask for more data. Until AI can do that, it might be a great assistant for writing reports, but it's not quite ready to run the lab itself. The gap between "knowing facts" and "reasoning from evidence" is the missing step toward real scientific discovery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.