← Latest papers
🤖 AI

AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

This paper demonstrates that in drug-asset valuation, while reasoning scaffolds improve an AI agent's calibration and discipline, the availability of proprietary evidence is the definitive limiting factor that sets the upper bound on factual accuracy and decision utility.

Original authors: Yinan Wang

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Yinan Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a team of detectives to solve a complex case: Is a new drug idea worth investing millions of dollars in?

To answer this, the detective needs to know three things:

  1. The Science: Does the biology make sense?
  2. The Competition: Who else is trying to do this?
  3. The History: Have similar drugs succeeded or failed before?

The paper you provided is a "stress test" of three different types of AI detectives to see which one is actually the best at solving this case. The researchers wanted to find out: Is the detective's intelligence (the AI model) the most important thing, or is the quality of the evidence they are allowed to read the real bottleneck?

Here is the breakdown of the experiment and what they found, using simple analogies.

The Three Detectives (The Experiment)

The researchers set up three versions of an AI agent, all using the exact same "brain" (the same underlying AI model). The only difference was what tools and information they were allowed to use.

  • Detective A (The "Web Surfer"):
    • Tools: Can only use a standard web search (like Google) and has no special training.
    • Analogy: This is a smart intern who goes to the public library and reads whatever books are on the open shelves. They are smart, but they can only find what anyone else can find.
  • Detective B (The "Trained Analyst"):
    • Tools: Has the same web access, but is given a strict playbook (a checklist of rules), a "red team" (a critic to check their work), and access to public databases (like clinical trial registries).
    • Analogy: This is the same intern, but now they have a senior mentor giving them a checklist, a rulebook on how to write reports, and access to public government records. They are more disciplined and organized, but they still can't see the "secret files."
  • Detective C (The "Insider"):
    • Tools: Has everything Detective B has, PLUS access to a massive, private, paid database (Noah AI) that contains curated records of drug trials, deals, and competitors that are hidden from the public web.
    • Analogy: This is the senior investigator who has a key to the confidential vault. They can see the private emails, the unpublished trial results, and the "long-tail" deals that never made the news.

The Results: What Did They Find?

The researchers tested these detectives on 13 different drug cases. Here is the verdict:

1. The "Playbook" Helps, But Only So Much (Detective B vs. A)

When Detective B (the trained analyst) was compared to Detective A (the web surfer), the results were interesting:

  • Better Discipline: Detective B wrote better reports. They were less likely to make wild guesses and followed the rules better.
  • The Problem: Detective B still missed 60% to 75% of the important competitors and deals.
  • The Analogy: Imagine Detective B is a very polite, well-organized librarian who follows every rule. But if the book they need to find is locked in a private vault, no amount of good organization will help them find it. They are "calibrated" (thinking clearly) but "blind" (missing facts).

2. The "Private Vault" Changes Everything (Detective C)

When Detective C (the insider) was brought in, the results changed dramatically:

  • Completeness: Detective C found 96% of the critical information. Detective A and B only found about 25% to 38%.
  • The "Long-Tail" Effect: The biggest gap was in finding "long-tail" data—small, obscure, or regional drug trials that never appear on the open web. Detective C found almost all of them; the others found almost none.
  • The Analogy: Detective C didn't just write a better report; they saw the whole board. While the others were playing chess with only half the pieces, Detective C saw the entire game.

3. The "Informed Decision" Score

This is the most important finding.

  • Raw Quality: If you just read the final report without knowing what facts were missing, Detective A and B sounded almost as good as Detective C. They sounded confident and logical.
  • Real Quality: But because A and B were missing so many facts, their decisions were based on incomplete information.
  • The Math: The researchers created a score called "Informed Decision Quality."
    • Detective A/B Score: ~2.0 (They sounded good, but were missing 75% of the facts).
    • Detective C Score: ~7.4 (They sounded good AND had all the facts).
  • The Ceiling: The paper argues that even if you gave Detective B a "perfect" brain and a perfect report, they would still be capped at a low score because they physically couldn't access the missing evidence. You cannot reason your way to a truth you cannot see.

The Big Takeaway

The paper concludes with a simple lesson for building AI scientists:

"Skills and data are both useful, but data is the limit."

  • Reasoning Skills (The Playbook): These are necessary. They stop the AI from being reckless, help it organize its thoughts, and ensure it follows rules. They make the AI a better writer and a more disciplined thinker.
  • Evidence Substrate (The Data): This is the limiting factor. If the AI doesn't have access to the private, curated, long-tail data, it literally cannot know the full truth. No amount of clever prompting or reasoning tricks can fill in the gaps of missing information.

In short: You can have the smartest detective in the world with the best rulebook, but if they aren't allowed into the evidence room, they will never solve the case correctly. For AI to be truly useful in high-stakes fields like drug discovery, it needs access to the private, curated evidence, not just the public internet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →