← Latest papers
💬 NLP

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

This paper evaluates large language models on scientific feasibility assessment, finding that providing experimental outcome evidence generally improves accuracy and robustness more effectively than providing experimental descriptions, which can degrade performance when context is incomplete.

Original authors: Seyedali Mohammadi, Manas Gaur, Francis Ferraro

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Seyedali Mohammadi, Manas Gaur, Francis Ferraro

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Is this scientific claim actually true, or is it just a wild guess?"

This paper is about testing how well modern AI detectives (Large Language Models, or LLMs) can solve these mysteries. The researchers wanted to know: Does giving the AI more "evidence" actually help it solve the case, or does it sometimes confuse it?

Here is the breakdown of their investigation using simple analogies.

1. The Setup: The "Scientific Feasibility" Test

The researchers gave the AI a specific claim to test, like:

"Drinking one extra glass of water every day will lower your blood pressure."

To figure out if this is Feasible (likely true and testable) or Infeasible (unlikely or impossible), the AI was allowed to look at different types of "evidence" from a real scientific paper. The researchers tested four scenarios:

  • Scenario A (The Guess): The AI gets only the claim. It has to guess based on what it already knows from its training (like a detective walking into a room with no clues).
  • Scenario B (The Blueprint): The AI gets the claim + a description of the Experiments (the "how-to" plan). Example: "We tested 100 people, gave them water, and measured their blood pressure." But the AI doesn't know the results yet.
  • Scenario C (The Result): The AI gets the claim + the Outcomes (the results). Example: "The study found that blood pressure went down slightly." But the AI doesn't know how they did the test.
  • Scenario D (The Full File): The AI gets the claim + the Experiments + the Outcomes. The whole story.

2. The Big Surprise: Results vs. Blueprints

The researchers expected that giving the AI the full file (Scenario D) would always be the best. They also thought the "Blueprint" (Scenario B) would be very helpful because it shows the scientific method.

They were wrong. Here is what they found:

  • The "Result" is King: Giving the AI the Outcomes (the results) was the most helpful thing. It was like giving the detective the final verdict. It made the AI much smarter and more accurate.
  • The "Blueprint" is a Trap: Surprisingly, giving the AI just the Experiment descriptions (without the results) often made the AI worse than if it had guessed on its own!
    • The Analogy: Imagine a detective is told, "We built a trap for a bear," but isn't told if the bear got caught. The detective might get confused, overthinking the trap design, and make a wrong guess. The AI got "distracted" by the details of the experiment design and lost track of the actual truth.

3. The "Half-Truth" Problem

The researchers also tested what happens if they only show the AI half of the evidence (e.g., only 50% of the experiments or results).

  • The Expectation: You'd think that if you have 50% of the clues, the AI would be 50% as good.
  • The Reality: The AI's performance was jumpy and unstable. Sometimes, having half the evidence made the AI perform worse than having no evidence at all.
    • The Analogy: Imagine trying to solve a puzzle with half the pieces. Instead of being "halfway" to the solution, you might force the wrong pieces together and create a picture that looks nothing like the real image. The AI got "anchored" on the few clues it had and made a confident but wrong conclusion.

4. Why Does This Happen?

The paper suggests three main reasons why the AI struggles:

  1. Surface Level Matching: The AI looks at the words. If the experiment description uses similar words to the claim (e.g., both say "water" and "blood pressure"), the AI thinks, "Aha! They match!" even if the experiment doesn't actually prove the claim.
  2. No "Gating" Mechanism: The AI doesn't know how to say, "Wait, this evidence doesn't actually test what I'm asking." It tries to force a connection even when the evidence is irrelevant or incomplete.
  3. Over-Confidence: When the AI sees some evidence, it gets too confident and stops being careful. When it sees no evidence, it's more humble and relies on its general knowledge, which sometimes works better than a bad piece of evidence.

5. The Takeaway for the Future

The main lesson is: More information doesn't always mean better answers for AI.

If you want an AI to check scientific facts, giving it the results of studies is crucial. But just dumping a bunch of experimental descriptions on it without the results can actually confuse it and make it less reliable than if you had left it alone.

In short:

  • Results = The truth (Helps the AI).
  • Methods/Blueprints = The process (Can confuse the AI if the results aren't there).
  • Half-evidence = A dangerous trap (Can make the AI confidently wrong).

The researchers conclude that we need to build AI that is better at knowing when to ignore bad or incomplete evidence, rather than just trying to process everything it is given.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →