← Latest papers
🤖 machine learning

Self Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale

This paper introduces "Starling," a multi-agent deep research system that autonomously transforms the entire PubMed corpus into large-scale, high-accuracy, and nuance-rich structured biomedical datasets, outperforming traditional manually curated repositories in both scale and data quality.

Original authors: Haydn Jones, Yimeng Zeng, Alden Rose, Li S. Yifei, Yining Huang, Kaiwen Wu, Jiaming Liang, Maggie Ziyu Huan, Yoseph Barash, Cesar de la Fuente-Nunez, Osbert Bastani, Zachary Ives, Mark Yatskar, Jacob
Published 2026-05-11
📖 4 min read☕ Coffee break read

Original authors: Haydn Jones, Yimeng Zeng, Alden Rose, Li S. Yifei, Yining Huang, Kaiwen Wu, Jiaming Liang, Maggie Ziyu Huan, Yoseph Barash, Cesar de la Fuente-Nunez, Osbert Bastani, Zachary Ives, Mark Yatskar, Jacob R. Gardner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of biomedical research as a massive, chaotic library containing 22.5 million books (scientific papers). For decades, scientists trying to build AI to discover new drugs have been trying to learn from a tiny, expensive, and slightly outdated "cheat sheet" that a few librarians manually wrote. This cheat sheet (existing databases) is great, but it's missing pages, it's slow to update, and the librarians often left out the messy details of how the experiments were actually done.

This paper introduces Starling, a new system that acts like a super-powered, tireless librarian who can read the entire library of 22.5 million books, instantly understand the context, and rewrite the cheat sheet to be bigger, more accurate, and full of the missing details.

Here is how Starling works, broken down into simple steps:

1. The Problem: The "Cheat Sheet" is Flawed

Think of existing drug databases like a spreadsheet where someone wrote: "Drug X works."

  • It's incomplete: It misses thousands of new studies published since the spreadsheet was last updated.
  • It lacks nuance: It doesn't tell you when Drug X works. Maybe it only works if the patient hasn't eaten, or only if they are taking it with a specific other drug. The spreadsheet throws away this crucial context.
  • It has errors: The paper found that the old cheat sheets are wrong about 7% to 16% of the time.

2. The Solution: Starling (The Self-Driving Librarian)

Instead of hiring more human librarians to read the books (which takes years and costs a fortune), the authors built Starling, an AI system that does the work itself.

Step A: The Tagging System (The Index)
First, Starling reads every single paper and puts sticky notes on them. It tags 4.5 billion specific things (like drug names, genes, diseases, and proteins) across 19 different categories. It's like having a library where every book is instantly indexed by every character and plot point inside it.

Step B: The Search (The Detective)
When a researcher asks a question like, "Find me every time anyone ever talked about how a drug gets into the brain," Starling doesn't just guess. It uses a "hybrid search" (combining keyword matching with understanding the meaning of sentences) to find the exact pages in the 22.5 million books that are relevant.

Step C: The Extraction (The Scribe)
This is the magic part. Starling uses a team of AI agents to:

  1. Design the form: It figures out what information matters (e.g., "Was the patient fasting?").
  2. Read and Write: It reads the specific paragraphs, extracts the data, and fills out a structured record.
  3. The Judge: A final AI "judge" checks the work. It asks: "Did you actually find this in the paper? Is the drug name correct? Did you invent a number?" If the answer is no, it throws the record away.

3. The Results: A Better Library

The paper tested Starling on six different tasks, such as finding out how toxic a chemical is or how well a drug crosses the blood-brain barrier.

  • Scale: Starling produced about 6.3 million records. For some tasks, this is 20 to 30 times larger than the best existing manual databases.
  • Accuracy: Surprisingly, Starling's records were more accurate than the human-curated ones. While the old databases had error rates of 7% to 16%, Starling's error rate was only 0.6% to 7.7%.
  • Nuance: This is the biggest win. Starling didn't just say "Drug X has 60% bioavailability." It said, "Drug X has 60% bioavailability if taken on an empty stomach, but only 20% if taken with a high-fat meal." It captured the "story" behind the number, which previous databases threw away.

4. The Cost

The paper notes that this entire process cost about one cent per record. In contrast, manually curating these databases costs a fortune and takes years.

Summary

The paper claims that by using AI to autonomously read the primary scientific literature, we can build datasets that are larger, cheaper, more accurate, and richer in detail than anything humans have manually curated before. They have released the code and the datasets so others can use this "self-driving" method to build their own knowledge bases from the literature.

What the paper does NOT claim:

  • It does not claim that these datasets have already cured diseases or led to new drugs being approved.
  • It does not claim that the AI is perfect or that it understands biology like a human doctor (it still needs to be checked).
  • It does not claim that the system can read images or graphs in the papers (it currently only reads text and tables).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →